Back to jobs

Data Engineer — ML Training Data Pipeline

Hyderabad, Telangana, India

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Check the employer’s page ↗

Tools in this posting

  • Python
  • AWS
  • S3
  • Huggingface
  • pandas
Source — Tool mentions in context
Job Title:Data Engineer - ML Training Data Pipeline Notice period: 0-30 Days Experience : 5+ Years Location: Hyderabad OR Pune We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS. What We Expect: Build end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets Process large-scale JSONL data on AWS S3 (tens of thousands of traces per batch) Convert between chat-completion formats (e.g., OpenAI → Llama 3.1 tool-calling format) Implement smart deduplication and sampling to balance training distribution Design identity-aware train/test splits that measure true generalization Build data validation gates to detect schema drift and format anomalies Create a continuous pipeline that auto-processes new production traces for retraining Requirements Experience: 6+ years data engineering focused on ML data pipelines Python: Strong — pandas, pyarrow, JSONL processing at scale ML Data Libraries: HuggingFace Datasets, Arrow-based storage Data Formats: Multi-turn conversation/chat data structures and tokenizer-specific formatting Deduplication: Content hashing, identity-based grouping strategies AWS: S3, EC2, batch processing workflows Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright. Benefits Comprehensive Medical Coverage: Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind. Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones. Retirement Benefits: PF and Gratuity provided as per standard government regulations. Flexible Work Options: Enjoy hybrid work arrangements & flexible working hours. Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays. Employee Well-being Spaces: Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.

Job description

View original posting ↗

Job Title:Data Engineer - ML Training Data Pipeline Notice period: 0-30 Days Experience : 5+ Years Location: Hyderabad OR Pune We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS. What We Expect: Build end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets Process large-scale JSONL data on AWS S3 (tens of thousands of traces per batch) Convert between chat-completion formats (e.g., OpenAI → Llama 3.1 tool-calling format) Implement smart deduplication and sampling to balance training distribution Design identity-aware train/test splits that measure true generalization Build data validation gates to detect schema drift and format anomalies Create a continuous pipeline that auto-processes new production traces for retraining Requirements Experience: 6+ years data engineering focused on ML data pipelines Python: Strong — pandas, pyarrow, JSONL processing at scale ML Data Libraries: HuggingFace Datasets, Arrow-based storage Data Formats: Multi-turn conversation/chat data structures and tokenizer-specific formatting Deduplication: Content hashing, identity-based grouping strategies AWS: S3, EC2, batch processing workflows Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright. Benefits Comprehensive Medical Coverage: Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind. Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones. Retirement Benefits: PF and Gratuity provided as per standard government regulations. Flexible Work Options: Enjoy hybrid work arrangements & flexible working hours. Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays. Employee Well-being Spaces: Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.

Your next step

Check the employer’s posting for the current role and application details.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Hyderabad, Telangana, India

This passage needs a closer read in the full description.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 8, 2026
Recorded sightings
16
Last seen by us
Oct 8, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error
AI answers unavailable

We couldn’t identify enough role detail in this saved description to support an AI answer. Read the full posting