Back to jobs

Big Data Engineer

Pune, Maharashtra, India

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at EXL Service

Tools in this posting

  • Python
  • Hadoop
  • Hive
  • Spark
  • SQL
Source — Tool mentions in context
- Mentor, guide, and support other data engineers through code reviews, design reviews, technical coaching, and hands-on problem-solving. - Build and optimize data workflows using Python, Spark, and the Hadoop ecosystem. - Work extensively with Hadoop ecosystem components (Hive, HDFS, Impala) to manage and query large-scale data.
- 7–10 years of overall experience in data engineering, with a proven track record in technical leadership (design ownership, mentoring, guiding development teams). - Python – strong hands-on development experience building production-grade data solutions. - Big Data / Hadoop ecosystem (Hadoop, Hive, Impala, HDFS) – deep, hands-on experience in on-premises environments.
- Build and optimize data workflows using Python, Spark, and the Hadoop ecosystem. - Work extensively with Hadoop ecosystem components (Hive, HDFS, Impala) to manage and query large-scale data. - Manage and optimize batch scheduling and job orchestration using enterprise schedulers such as CA7 or Control-M.
- Python – strong hands-on development experience building production-grade data solutions. - Big Data / Hadoop ecosystem (Hadoop, Hive, Impala, HDFS) – deep, hands-on experience in on-premises environments. - Apache Spark – solid experience developing and tuning large-scale distributed data processing jobs.
- Big Data / Hadoop ecosystem (Hadoop, Hive, Impala, HDFS) – deep, hands-on experience in on-premises environments. - Apache Spark – solid experience developing and tuning large-scale distributed data processing jobs. - Job scheduling / orchestration – hands-on experience with CA7 or Control-M (or comparable enterprise schedulers).
- Job scheduling / orchestration – hands-on experience with CA7 or Control-M (or comparable enterprise schedulers). - Strong understanding of data structures, ETL processes, and SQL. - Extensive experience with large-scale data processing and distributed systems.

Job description

View original posting ↗

Job Description: Lead Data Engineer Summary

We are looking for an experienced Lead Data Engineer with 7–10 years of hands-on experience to design, build, and maintain scalable data pipelines and processing systems in an on-premises Big Data environment. This is a technical leadership role: beyond strong individual contribution, the ideal candidate will own architecture and design decisions, set technical direction, and mentor and support other developers on the team. The role works closely with cross-functional teams to deliver reliable, high-quality data solutions that support business and analytics needs.

Roles & Responsibilities
  • Lead the design, development, and maintenance of robust, scalable data pipelines for ingestion, transformation, and processing of large datasets in an on-premises environment.
  • Own architectural and design decisions for data solutions, evaluating trade-offs and defining technical standards for the team.
  • Mentor, guide, and support other data engineers through code reviews, design reviews, technical coaching, and hands-on problem-solving.
  • Build and optimize data workflows using Python, Spark, and the Hadoop ecosystem.
  • Work extensively with Hadoop ecosystem components (Hive, HDFS, Impala) to manage and query large-scale data.
  • Manage and optimize batch scheduling and job orchestration using enterprise schedulers such as CA7 or Control-M.
  • Ensure data quality, integrity, and performance across data platforms.
  • Collaborate with data analysts, data scientists, and business stakeholders to translate data requirements into sound technical designs.
  • Troubleshoot and resolve complex issues in data pipelines and production environments, acting as an escalation point for the team.
  • Champion best practices for coding standards, version control, testing, and documentation.
  • Stay current with emerging technologies, particularly AI/ML capabilities, and identify opportunities to apply them to data engineering workflows.
Technical Skills Must Have
  • 7–10 years of overall experience in data engineering, with a proven track record in technical leadership (design ownership, mentoring, guiding development teams).
  • Python – strong hands-on development experience building production-grade data solutions.
  • Big Data / Hadoop ecosystem (Hadoop, Hive, Impala, HDFS) – deep, hands-on experience in on-premises environments.
  • Apache Spark – solid experience developing and tuning large-scale distributed data processing jobs.
  • Job scheduling / orchestration – hands-on experience with CA7 or Control-M (or comparable enterprise schedulers).
  • Strong understanding of data structures, ETL processes, and SQL.
  • Extensive experience with large-scale data processing and distributed systems.
  • Demonstrated ability to make sound architecture/design decisions and to mentor and support other developers.
  • Exposure to AI/ML concepts or tools, with a strong willingness to learn and grow in this space.
Roles & Responsibilities
  • Lead the design, development, and maintenance of robust, scalable data pipelines for ingestion, transformation, and processing of large datasets in an on-premises environment.
  • Own architectural and design decisions for data solutions, evaluating trade-offs and defining technical standards for the team.
  • Mentor, guide, and support other data engineers through code reviews, design reviews, technical coaching, and hands-on problem-solving.
  • Build and optimize data workflows using Python, Spark, and the Hadoop ecosystem.
  • Work extensively with Hadoop ecosystem components (Hive, HDFS, Impala) to manage and query large-scale data.
  • Manage and optimize batch scheduling and job orchestration using enterprise schedulers such as CA7 or Control-M.
  • Ensure data quality, integrity, and performance across data platforms.
  • Collaborate with data analysts, data scientists, and business stakeholders to translate data requirements into sound technical designs.
  • Troubleshoot and resolve complex issues in data pipelines and production environments, acting as an escalation point for the team.
  • Champion best practices for coding standards, version control, testing, and documentation.
  • Stay current with emerging technologies, particularly AI/ML capabilities, and identify opportunities to apply them to data engineering workflows.

7–10 years of overall experience in data engineering, with a proven track record in technical leadership (design ownership, mentoring, guiding development teams).

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on fa-ewjt-saasfaprod1.fa.ocs.oraclecloud.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Pune, Maharashtra, India

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Aug 12, 2026
Recorded sightings
147
Last seen by us
Sep 25, 2026
Employer says posted
Jul 21, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.