MLOps / Serving Engineer
Hyderabad, Telangana, India
- Pay
- Salary not listed in the saved posting
- Work setup
- Unconfirmed
- Employment
- Unconfirmed
Tools in this posting
- AWS
- Docker
- Grafana
- Kubernetes
Source — Tool mentions in context
Job Title: MLOps / Serving Engineer Experience : 5+ years Location : Hyderabad OR Pune Notice Period: 0-30 days Work mode - Hybrid We are seeking an experienced MLOps / Serving Engineer who can design and operate the production serving infrastructure for fine-tuned LLMs on AWS — optimised inference engines, shadow-mode and staged rollout pipelines, monitoring dashboards, and the path from experimental model to full production traffic. Key Responsibilities: Deploy fine-tuned LLMs using vLLM, TensorRT-LLM, or Triton with continuous batching on AWS GPU instances Build shadow-mode deployment: run fine-tuned model alongside production, log comparison data without impacting live traffic Execute staged rollout: canary (5%) → gradual ramp (25% → 50% → 100%) with automated rollback on quality degradation Optimize inference for input-heavy workloads (~17K token inputs, ~130 token outputs): prefill throughput, KV-cache, INT8 quantization Build monitoring dashboards: latency, throughput, accuracy metrics, cost per request Design auto-scaling; implement high-availability (2× instances); automated rollback triggers on end-to-end quality metrics Requirements 5+ years MLOps or ML infrastructure engineering Hands-on with vLLM, TensorRT-LLM, or Triton Inference Server Deep familiarity with g5, p4de, p5 instance families, EC2 auto-scaling Have worked on Deployment patterns like Shadow-mode, canary, A/B traffic routing, automated rollback Experience onto Continuous batching, INT8 quantization, KV-cache management Expertise on Docker, Kubernetes (EKS) for ML workloads Worked on CloudWatch, Prometheus, Grafana Benefits Comprehensive Medical Coverage: Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind. Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones. Retirement Benefits: PF and Gratuity provided as per standard government regulations. Flexible Work Options: Enjoy hybrid work arrangements & flexible working hours Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays. Employee Well-being Spaces: Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.
Job description
Job Title: MLOps / Serving Engineer Experience : 5+ years Location : Hyderabad OR Pune Notice Period: 0-30 days Work mode - Hybrid We are seeking an experienced MLOps / Serving Engineer who can design and operate the production serving infrastructure for fine-tuned LLMs on AWS — optimised inference engines, shadow-mode and staged rollout pipelines, monitoring dashboards, and the path from experimental model to full production traffic. Key Responsibilities: Deploy fine-tuned LLMs using vLLM, TensorRT-LLM, or Triton with continuous batching on AWS GPU instances Build shadow-mode deployment: run fine-tuned model alongside production, log comparison data without impacting live traffic Execute staged rollout: canary (5%) → gradual ramp (25% → 50% → 100%) with automated rollback on quality degradation Optimize inference for input-heavy workloads (~17K token inputs, ~130 token outputs): prefill throughput, KV-cache, INT8 quantization Build monitoring dashboards: latency, throughput, accuracy metrics, cost per request Design auto-scaling; implement high-availability (2× instances); automated rollback triggers on end-to-end quality metrics Requirements 5+ years MLOps or ML infrastructure engineering Hands-on with vLLM, TensorRT-LLM, or Triton Inference Server Deep familiarity with g5, p4de, p5 instance families, EC2 auto-scaling Have worked on Deployment patterns like Shadow-mode, canary, A/B traffic routing, automated rollback Experience onto Continuous batching, INT8 quantization, KV-cache management Expertise on Docker, Kubernetes (EKS) for ML workloads Worked on CloudWatch, Prometheus, Grafana Benefits Comprehensive Medical Coverage: Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind. Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones. Retirement Benefits: PF and Gratuity provided as per standard government regulations. Flexible Work Options: Enjoy hybrid work arrangements & flexible working hours Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays. Employee Well-being Spaces: Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.
Your next step
Check the employer’s posting for the current role and application details.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
No pay amount identified in the saved description.
- Location & working pattern
Hyderabad, Telangana, India
This passage needs a closer read in the full description.
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Sep 8, 2026
- Recorded sightings
- 16
- Last seen by us
- Oct 8, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorAI answers unavailable
We couldn’t identify enough role detail in this saved description to support an AI answer. Read the full posting