Back to jobs

AI MLOPS/LLMOps Engineer

Gurugram, Haryana, India

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at EXL Service

What you’ll work on

Full posting
  • Build and tune hybrid search pipelines combining GTE multilingual dense embeddings with GIN lexical search on Aurora PostgreSQL

  • Manage data flow across AWS services: S3, Athena, Glue, Fargate, SQS, Step Functions

  • Manage schema evolution and migrations using Liquibase on Aurora PostgreSQL

From the employer’s posting
Connect DS-coded AI modules into end-to-end production workflows via Airflow DAGs on AWS EKS Build and tune hybrid search pipelines combining GTE multilingual dense embeddings with GIN lexical search on Aurora PostgreSQL Integrate with OpenAI-based API platform for multilingual query expansion and LLM-driven trend detection
Author and maintain Airflow DAGs for batch processing (monthly entity refresh, trend detection, DEID pipeline) Manage data flow across AWS services: S3, Athena, Glue, Fargate, SQS, Step Functions Scale pipelines to handle 500K+ claims and hundreds of millions of text chunks
Deploy and version models using MLflow and Databricks Manage schema evolution and migrations using Liquibase on Aurora PostgreSQL Instrument pipelines with logging, monitoring, and evaluation scoring for retrieval quality

What you’ll bring

All qualifications

Core experience

  • Bachelor's degree in Computer Science, Information Technology, Data Science, Artificial Intelligence, Statistics, Mathematics, or a related field.
  • Master's degree in Data Science, AI/ML, Computer Science, or Analytics is preferred but not mandatory.
Qualification wording
Bachelor's degree in Computer Science, Information Technology, Data Science, Artificial Intelligence, Statistics, Mathematics, or a related field.
Master's degree in Data Science, AI/ML, Computer Science, or Analytics is preferred but not mandatory.

Tools in this posting

  • Python
  • SQL
  • AWS
  • Azure
  • Databricks
  • MLflow
  • S3
  • Airflow
  • PostgreSQL
Source — Tool mentions in context
Key strengths should include: - Proficiency in Python and SQL with experience deploying AI/NLP solutions such as document classification, entity extraction, NER, PII masking, de-identification, hybrid search, and LLM integrations. - Strong knowledge of Apache Airflow for orchestrating end-to-end data pipelines and automating batch processing workflows.
Area Skills Languages Python (primary), SQL AI / NLP LLM API integration, multilingual embeddings (e.g., GTE), hybrid search, text classification, entity extraction, NER, PII masking
- Strong knowledge of Apache Airflow for orchestrating end-to-end data pipelines and automating batch processing workflows. - Experience working with AWS services including S3, Athena, Glue, Fargate, EKS, SQS, and Step Functions. - Capability to design and maintain large-scale document processing systems handling complex JSON structures, embedded documents, and multilingual content.
- Integrate and adjust inference pipelines for NLP modules including document classification, entity extraction, de-identification (DEID), and LLM-based early trend detection - Connect DS-coded AI modules into end-to-end production workflows via Airflow DAGs on AWS EKS - Build and tune hybrid search pipelines combining GTE multilingual dense embeddings with GIN lexical search on Aurora PostgreSQL
- Author and maintain Airflow DAGs for batch processing (monthly entity refresh, trend detection, DEID pipeline) - Manage data flow across AWS services: S3, Athena, Glue, Fargate, SQS, Step Functions - Scale pipelines to handle 500K+ claims and hundreds of millions of text chunks
Data Pipelines Apache Airflow, batch orchestration, large-scale unstructured data processing Cloud & Infrastructure AWS (S3, Athena, Glue, Fargate, EKS, SQS, Step Functions) Databases PostgreSQL / Aurora, pgvector, GIN indexes, full-text search
- Familiarity with vector search and retrieval systems, including embeddings, pgvector, PostgreSQL/Aurora, GIN indexes, and full-text search. - Experience with ML lifecycle management using MLflow, Databricks/Azure Databricks, model deployment, monitoring, and evaluation frameworks. - Strong DevOps practices including GitHub-based development, CI/CD pipelines, schema management, and production support.
Databases PostgreSQL / Aurora, pgvector, GIN indexes, full-text search ML Platform MLflow, Databricks / Azure Databricks DevOps GitHub, CI/CD pipelines
Production Deployment & Quality - Deploy and version models using MLflow and Databricks - Manage schema evolution and migrations using Liquibase on Aurora PostgreSQL
Document Processing & Parsing - Design and maintain document preprocessing pipelines that parse deeply nested JSON structures (emails with attachments, embedded PDFs) from S3/DataLake - Handle multilingual unstructured text (English, Spanish, Portuguese, German, Dutch, French, Italian) across 300 GB of claim notes and documents
- Proficiency in Python and SQL with experience deploying AI/NLP solutions such as document classification, entity extraction, NER, PII masking, de-identification, hybrid search, and LLM integrations. - Strong knowledge of Apache Airflow for orchestrating end-to-end data pipelines and automating batch processing workflows. - Experience working with AWS services including S3, Athena, Glue, Fargate, EKS, SQS, and Step Functions.
Data Pipeline Engineering - Author and maintain Airflow DAGs for batch processing (monthly entity refresh, trend detection, DEID pipeline) - Manage data flow across AWS services: S3, Athena, Glue, Fargate, SQS, Step Functions
AI / NLP LLM API integration, multilingual embeddings (e.g., GTE), hybrid search, text classification, entity extraction, NER, PII masking Data Pipelines Apache Airflow, batch orchestration, large-scale unstructured data processing Cloud & Infrastructure AWS (S3, Athena, Glue, Fargate, EKS, SQS, Step Functions)
- Capability to design and maintain large-scale document processing systems handling complex JSON structures, embedded documents, and multilingual content. - Familiarity with vector search and retrieval systems, including embeddings, pgvector, PostgreSQL/Aurora, GIN indexes, and full-text search. - Experience with ML lifecycle management using MLflow, Databricks/Azure Databricks, model deployment, monitoring, and evaluation frameworks.
- Connect DS-coded AI modules into end-to-end production workflows via Airflow DAGs on AWS EKS - Build and tune hybrid search pipelines combining GTE multilingual dense embeddings with GIN lexical search on Aurora PostgreSQL - Integrate with OpenAI-based API platform for multilingual query expansion and LLM-driven trend detection
- Deploy and version models using MLflow and Databricks - Manage schema evolution and migrations using Liquibase on Aurora PostgreSQL - Instrument pipelines with logging, monitoring, and evaluation scoring for retrieval quality
Cloud & Infrastructure AWS (S3, Athena, Glue, Fargate, EKS, SQS, Step Functions) Databases PostgreSQL / Aurora, pgvector, GIN indexes, full-text search ML Platform MLflow, Databricks / Azure Databricks

Job description

View original posting ↗

Seeking a strong Data Engineer / AI Engineer with expertise in building and operationalizing large-scale AI and NLP solutions on cloud platforms. The ideal candidate should have hands-on experience integrating AI/LLM models into production workflows, developing scalable data pipelines, and processing large volumes of multilingual unstructured text.

Key strengths should include:

  • Proficiency in Python and SQL with experience deploying AI/NLP solutions such as document classification, entity extraction, NER, PII masking, de-identification, hybrid search, and LLM integrations.
  • Strong knowledge of Apache Airflow for orchestrating end-to-end data pipelines and automating batch processing workflows.
  • Experience working with AWS services including S3, Athena, Glue, Fargate, EKS, SQS, and Step Functions.
  • Capability to design and maintain large-scale document processing systems handling complex JSON structures, embedded documents, and multilingual content.
  • Familiarity with vector search and retrieval systems, including embeddings, pgvector, PostgreSQL/Aurora, GIN indexes, and full-text search.
  • Experience with ML lifecycle management using MLflow, Databricks/Azure Databricks, model deployment, monitoring, and evaluation frameworks.
  • Strong DevOps practices including GitHub-based development, CI/CD pipelines, schema management, and production support.

What You Will Do

AI Module Integration & Inference Pipelines

  • Integrate and adjust inference pipelines for NLP modules including document classification, entity extraction, de-identification (DEID), and LLM-based early trend detection
  • Connect DS-coded AI modules into end-to-end production workflows via Airflow DAGs on AWS EKS
  • Build and tune hybrid search pipelines combining GTE multilingual dense embeddings with GIN lexical search on Aurora PostgreSQL
  • Integrate with OpenAI-based API platform for multilingual query expansion and LLM-driven trend detection

Document Processing & Parsing

  • Design and maintain document preprocessing pipelines that parse deeply nested JSON structures (emails with attachments, embedded PDFs) from S3/DataLake
  • Handle multilingual unstructured text (English, Spanish, Portuguese, German, Dutch, French, Italian) across 300 GB of claim notes and documents
  • Build chunking strategies and metadata extraction for downstream embedding and retrieval workflows

Data Pipeline Engineering

  • Author and maintain Airflow DAGs for batch processing (monthly entity refresh, trend detection, DEID pipeline)
  • Manage data flow across AWS services: S3, Athena, Glue, Fargate, SQS, Step Functions
  • Scale pipelines to handle 500K+ claims and hundreds of millions of text chunks

Production Deployment & Quality

  • Deploy and version models using MLflow and Databricks
  • Manage schema evolution and migrations using Liquibase on Aurora PostgreSQL
  • Instrument pipelines with logging, monitoring, and evaluation scoring for retrieval quality
Area Skills Languages Python (primary), SQL AI / NLP LLM API integration, multilingual embeddings (e.g., GTE), hybrid search, text classification, entity extraction, NER, PII masking Data Pipelines Apache Airflow, batch orchestration, large-scale unstructured data processing Cloud & Infrastructure AWS (S3, Athena, Glue, Fargate, EKS, SQS, Step Functions) Databases PostgreSQL / Aurora, pgvector, GIN indexes, full-text search ML Platform MLflow, Databricks / Azure Databricks DevOps GitHub, CI/CD pipelines
Education
  • Bachelor's degree in Computer Science, Information Technology, Data Science, Artificial Intelligence, Statistics, Mathematics, or a related field.
  • Master's degree in Data Science, AI/ML, Computer Science, or Analytics is preferred but not mandatory.
  • Relevant cloud or data engineering certifications are advantageous.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on fa-ewjt-saasfaprod1.fa.ocs.oraclecloud.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Gurugram, Haryana, India

Key strengths should include: - Proficiency in Python and SQL with experience deploying AI/NLP solutions such as document classification, entity extraction, NER, PII masking, de-identification, hybrid search, and LLM integrations. - Strong knowledge of Apache Airflow for orchestrating end-to-end data pipelines and automating batch processing workflows.
More source context
- Connect DS-coded AI modules into end-to-end production workflows via Airflow DAGs on AWS EKS - Build and tune hybrid search pipelines combining GTE multilingual dense embeddings with GIN lexical search on Aurora PostgreSQL - Integrate with OpenAI-based API platform for multilingual query expansion and LLM-driven trend detection

More relevant text appears in the full description.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 23, 2026
Recorded sightings
24
Last seen by us
Oct 8, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.