Senior MLOps Engineer
Abu Dhabi
- Pay
- Salary not listed in the saved posting
- Work setup
- Unconfirmed
- Employment
Full-time — employment source
Employment type Full-time
Read the full posting
What you’ll work on
Full postingDesign and manage scalable ML infrastructure on AWS using EKS, EC2, RDS, S3, and IAM-based access control.
Build and maintain Kubernetes deployments for LLM and TTS inference using Helm, ArgoCD, and Prometheus/Grafana monitoring.
Implement and optimize model serving pipelines using vLLM, SGLang, TensorRT, or similar frameworks for high-throughput inference.
From the employer’s posting
Key Responsibilities Design and manage scalable ML infrastructure on AWS using EKS, EC2, RDS, S3, and IAM-based access control. Build and maintain Kubernetes deployments for LLM and TTS inference using Helm, ArgoCD, and Prometheus/Grafana monitoring.
Design and manage scalable ML infrastructure on AWS using EKS, EC2, RDS, S3, and IAM-based access control. Build and maintain Kubernetes deployments for LLM and TTS inference using Helm, ArgoCD, and Prometheus/Grafana monitoring. Implement and optimize model serving pipelines using vLLM, SGLang, TensorRT, or similar frameworks for high-throughput inference.
Build and maintain Kubernetes deployments for LLM and TTS inference using Helm, ArgoCD, and Prometheus/Grafana monitoring. Implement and optimize model serving pipelines using vLLM, SGLang, TensorRT, or similar frameworks for high-throughput inference. Develop CI/CD and MLOps automation for data versioning, model validation, and deployment (GitHub Actions, Jenkins, or AWS CodePipeline).
Tools in this posting
- Python
- AWS
- Docker
- Grafana
- Kubernetes
- MLflow
- S3
- Terraform
Source — Tool mentions in context
- Proficiency with AWS services (EKS, EC2, S3, RDS, CloudWatch, IAM). - Solid experience with Python, Docker, Git, and CI/CD pipelines. - Strong understanding of model lifecycle management, data pipelines, and observability tools (Grafana, Prometheus, Loki).
As a Senior MLOps Engineer, you will design, build, and maintain robust ML(Machine Learning) infrastructure across training, inference, and deployment pipelines. You will take ownership of the model lifecycle — from data ingestion to real-time serving — and ensure our LLM and speech models are deployed efficiently, securely, and reproducibly in Kubernetes-based environments. This position requires deep hands-on experience with Kubernetes (EKS), Helm, AWS cloud infrastructure, and modern MLOps toolchains (e.g., vLLM, SGLang, OpenWebUI, Weights & Biases, MLflow). Familiarity with speech/voice AI frameworks like ElevenLabs, Whisper, and RVC is also valuable. Key Responsibilities
Key Responsibilities - Design and manage scalable ML infrastructure on AWS using EKS, EC2, RDS, S3, and IAM-based access control. - Build and maintain Kubernetes deployments for LLM and TTS inference using Helm, ArgoCD, and Prometheus/Grafana monitoring.
- Implement and optimize model serving pipelines using vLLM, SGLang, TensorRT, or similar frameworks for high-throughput inference. - Develop CI/CD and MLOps automation for data versioning, model validation, and deployment (GitHub Actions, Jenkins, or AWS CodePipeline). - Integrate OpenWebUI, Gradio, or similar UIs for user-facing model demos and internal evaluation tools.
- Experience deploying ML models via vLLM, SGLang, TensorRT, or Ray Serve. - Proficiency with AWS services (EKS, EC2, S3, RDS, CloudWatch, IAM). - Solid experience with Python, Docker, Git, and CI/CD pipelines.
- Design and manage scalable ML infrastructure on AWS using EKS, EC2, RDS, S3, and IAM-based access control. - Build and maintain Kubernetes deployments for LLM and TTS inference using Helm, ArgoCD, and Prometheus/Grafana monitoring. - Implement and optimize model serving pipelines using vLLM, SGLang, TensorRT, or similar frameworks for high-throughput inference.
- Solid experience with Python, Docker, Git, and CI/CD pipelines. - Strong understanding of model lifecycle management, data pipelines, and observability tools (Grafana, Prometheus, Loki). - Excellent collaboration skills with ML researchers and software engineers.
The Role As a Senior MLOps Engineer, you will design, build, and maintain robust ML(Machine Learning) infrastructure across training, inference, and deployment pipelines. You will take ownership of the model lifecycle — from data ingestion to real-time serving — and ensure our LLM and speech models are deployed efficiently, securely, and reproducibly in Kubernetes-based environments. This position requires deep hands-on experience with Kubernetes (EKS), Helm, AWS cloud infrastructure, and modern MLOps toolchains (e.g., vLLM, SGLang, OpenWebUI, Weights & Biases, MLflow). Familiarity with speech/voice AI frameworks like ElevenLabs, Whisper, and RVC is also valuable.
- 4+ years of experience in MLOps, DevOps, or Cloud Infrastructure Engineering for ML systems. - Strong proficiency in Kubernetes, Helm, and container orchestration. - Experience deploying ML models via vLLM, SGLang, TensorRT, or Ray Serve.
- Familiarity with multi-GPU scheduling, NCCL optimization, and HPC cluster integration. - Knowledge of security, cost management, and network policy in multi-tenant Kubernetes clusters and cloudflare systems. - Prior work in LLM deployment, fine-tuning pipelines, or foundation model research.
- Contribute to internal tools for dataset curation, model monitoring, and retraining pipelines. - Maintain infrastructure-as-code using Terraform and Helm charts for reproducibility and governance. - Support real-time multimodal workloads (voice, text, vision) across inference clusters.
Job description
Key Responsibilities
- Design and manage scalable ML infrastructure on AWS using EKS, EC2, RDS, S3, and IAM-based access control.
- Build and maintain Kubernetes deployments for LLM and TTS inference using Helm, ArgoCD, and Prometheus/Grafana monitoring.
- Implement and optimize model serving pipelines using vLLM, SGLang, TensorRT, or similar frameworks for high-throughput inference.
- Develop CI/CD and MLOps automation for data versioning, model validation, and deployment (GitHub Actions, Jenkins, or AWS CodePipeline).
- Integrate OpenWebUI, Gradio, or similar UIs for user-facing model demos and internal evaluation tools.
- Collaborate with ML researchers to productize models — including TTS (e.g., ElevenLabs API), ASR (Whisper), and LLM-based chat systems.
- Ensure observability, cost optimization, and reliability of cloud resources across multiple environments.
- Contribute to internal tools for dataset curation, model monitoring, and retraining pipelines.
- Maintain infrastructure-as-code using Terraform and Helm charts for reproducibility and governance.
- Support real-time multimodal workloads (voice, text, vision) across inference clusters.
Academic Qualifications
- 4+ years of experience in MLOps, DevOps, or Cloud Infrastructure Engineering for ML systems.
- Strong proficiency in Kubernetes, Helm, and container orchestration.
- Experience deploying ML models via vLLM, SGLang, TensorRT, or Ray Serve.
- Proficiency with AWS services (EKS, EC2, S3, RDS, CloudWatch, IAM).
- Solid experience with Python, Docker, Git, and CI/CD pipelines.
- Strong understanding of model lifecycle management, data pipelines, and observability tools (Grafana, Prometheus, Loki).
- Excellent collaboration skills with ML researchers and software engineers.
Professional Experience – Preferred
- Extensive Experience with vLLM, K8s, Elevenlabs, Whisper, Gradio/OpenWebUI, or custom TTS/ASR model hosting.
- Familiarity with multi-GPU scheduling, NCCL optimization, and HPC cluster integration.
- Knowledge of security, cost management, and network policy in multi-tenant Kubernetes clusters and cloudflare systems.
- Prior work in LLM deployment, fine-tuning pipelines, or foundation model research.
- Exposure to data governance and responsible AI operations in research or enterprise settings.
Employment type
Full-time
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
- Ask the employer about the salary range before committing time to the process.
Complete your application on jobs.lever.co. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
No pay amount identified in the saved description.
- Location & working pattern
Abu Dhabi
Working pattern and location restrictions need checking in the full posting.
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Apr 15, 2026
- Recorded sightings
- 121
- Last seen by us
- Oct 9, 2026
- Employer says posted
- Nov 3, 2025
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.