Back to jobs

MLOps / Serving Engineer

Hyderabad, Telangana, India

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Check the employer’s page ↗

Tools in this posting

  • AWS
  • Docker
  • Grafana
  • Kubernetes
Source — Tool mentions in context
Job Title: MLOps / Serving Engineer Experience : 5+ years Location : Hyderabad OR Pune Notice Period: 0-30 days Work mode - Hybrid We are seeking an experienced MLOps / Serving Engineer who can design and operate the production serving infrastructure for fine-tuned LLMs on AWS — optimised inference engines, shadow-mode and staged rollout pipelines, monitoring dashboards, and the path from experimental model to full production traffic. Key Responsibilities: Deploy fine-tuned LLMs using vLLM, TensorRT-LLM, or Triton with continuous batching on AWS GPU instances Build shadow-mode deployment: run fine-tuned model alongside production, log comparison data without impacting live traffic Execute staged rollout: canary (5%) → gradual ramp (25% → 50% → 100%) with automated rollback on quality degradation Optimize inference for input-heavy workloads (~17K token inputs, ~130 token outputs): prefill throughput, KV-cache, INT8 quantization Build monitoring dashboards: latency, throughput, accuracy metrics, cost per request Design auto-scaling; implement high-availability (2× instances); automated rollback triggers on end-to-end quality metrics Requirements 5+ years MLOps or ML infrastructure engineering Hands-on with vLLM, TensorRT-LLM, or Triton Inference Server Deep familiarity with g5, p4de, p5 instance families, EC2 auto-scaling Have worked on Deployment patterns like Shadow-mode, canary, A/B traffic routing, automated rollback Experience onto Continuous batching, INT8 quantization, KV-cache management Expertise on Docker, Kubernetes (EKS) for ML workloads Worked on CloudWatch, Prometheus, Grafana Benefits Comprehensive Medical Coverage: Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind. Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones. Retirement Benefits: PF and Gratuity provided as per standard government regulations. Flexible Work Options: Enjoy hybrid work arrangements & flexible working hours Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays. Employee Well-being Spaces: Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.

Job description

View original posting ↗

Job Title: MLOps / Serving Engineer Experience : 5+ years Location : Hyderabad OR Pune Notice Period: 0-30 days Work mode - Hybrid We are seeking an experienced MLOps / Serving Engineer who can design and operate the production serving infrastructure for fine-tuned LLMs on AWS — optimised inference engines, shadow-mode and staged rollout pipelines, monitoring dashboards, and the path from experimental model to full production traffic. Key Responsibilities: Deploy fine-tuned LLMs using vLLM, TensorRT-LLM, or Triton with continuous batching on AWS GPU instances Build shadow-mode deployment: run fine-tuned model alongside production, log comparison data without impacting live traffic Execute staged rollout: canary (5%) → gradual ramp (25% → 50% → 100%) with automated rollback on quality degradation Optimize inference for input-heavy workloads (~17K token inputs, ~130 token outputs): prefill throughput, KV-cache, INT8 quantization Build monitoring dashboards: latency, throughput, accuracy metrics, cost per request Design auto-scaling; implement high-availability (2× instances); automated rollback triggers on end-to-end quality metrics Requirements 5+ years MLOps or ML infrastructure engineering Hands-on with vLLM, TensorRT-LLM, or Triton Inference Server Deep familiarity with g5, p4de, p5 instance families, EC2 auto-scaling Have worked on Deployment patterns like Shadow-mode, canary, A/B traffic routing, automated rollback Experience onto Continuous batching, INT8 quantization, KV-cache management Expertise on Docker, Kubernetes (EKS) for ML workloads Worked on CloudWatch, Prometheus, Grafana Benefits Comprehensive Medical Coverage: Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind. Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones. Retirement Benefits: PF and Gratuity provided as per standard government regulations. Flexible Work Options: Enjoy hybrid work arrangements & flexible working hours Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays. Employee Well-being Spaces: Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.

Your next step

Check the employer’s posting for the current role and application details.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Hyderabad, Telangana, India

This passage needs a closer read in the full description.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 8, 2026
Recorded sightings
16
Last seen by us
Oct 8, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error
AI answers unavailable

We couldn’t identify enough role detail in this saved description to support an AI answer. Read the full posting