Back to jobs

Data Engineer

IND - Bengaluru

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Full-time — employment source
Time Type Full time
Read the full posting

What you’ll work on

Full posting
  • Design, develop, test, and maintain scalable data pipelines and integrations using Databricks, PySpark, and SQL.

  • Build datasets optimized for analytics, BI, and downstream consumption while ensuring data quality, reconciliation, and production reliability.

  • Work within established data frameworks, design patterns, and reusable components created by other engineering teams.

From the employer’s posting
RESPONSIBILITIES: Design, develop, test, and maintain scalable data pipelines and integrations using Databricks, PySpark, and SQL. Build datasets optimized for analytics, BI, and downstream consumption while ensuring data quality, reconciliation, and production reliability.
Design, develop, test, and maintain scalable data pipelines and integrations using Databricks, PySpark, and SQL. Build datasets optimized for analytics, BI, and downstream consumption while ensuring data quality, reconciliation, and production reliability. Work within established data frameworks, design patterns, and reusable components created by other engineering teams.
Build datasets optimized for analytics, BI, and downstream consumption while ensuring data quality, reconciliation, and production reliability. Work within established data frameworks, design patterns, and reusable components created by other engineering teams. Read, understand, troubleshoot, and extend existing codebases and pipeline logic in line with engineering standards.

What you’ll bring

All qualifications

Core experience

  • Bachelor’s or Master’s degree in Computer Science, Engineering, Information Systems, or related field.
  • Strong proficiency in PySpark and SQL; Python alone is not sufficient for this role.
  • 5+ years of experience in data engineering, data warehousing, or large-scale data platform development.
  • Strong understanding of distributed data processing, performance optimization, and scalable pipeline design.
  • Experience working with existing enterprise frameworks, shared libraries, and engineering standards.
  • Ability to work effectively within predefined patterns, frameworks, and architectural guardrails.

Preferred experience

  • Experience building and maintaining data pipelines for batch processing; exposure to streaming is a plus.
Qualification wording
Bachelor’s or Master’s degree in Computer Science, Engineering, Information Systems, or related field.
Strong proficiency in PySpark and SQL; Python alone is not sufficient for this role.
5+ years of experience in data engineering, data warehousing, or large-scale data platform development.
Strong understanding of distributed data processing, performance optimization, and scalable pipeline design.
Experience working with existing enterprise frameworks, shared libraries, and engineering standards.
Ability to work effectively within predefined patterns, frameworks, and architectural guardrails.
Experience building and maintaining data pipelines for batch processing; exposure to streaming is a plus.
Education & alternatives
EDUCATION AND EXPERIENCE: - Bachelor’s or Master’s degree in Computer Science, Engineering, Information Systems, or related field. - 5+ years of experience in data engineering, data warehousing, or large-scale data platform development.

Tools in this posting

  • AWS
  • Azure
  • Databricks
  • PySpark
  • SQL
  • Python
  • Spark
  • Kafka
  • Terraform
Source — Tool mentions in context
- Experience reading, understanding, debugging, and enhancing existing code developed by other teams. - Experience with cloud data platforms such as AWS or Azure. - Experience working in agile, cross-functional engineering environments.
- Preferred - Databricks Certified Data Engineer Associate / Professional - Preferred - AWS or Azure Data Engineering certification PHYSICAL DEMANDS:
- Build and maintain scalable data pipelines and datasets that support analytics, reporting, and downstream business systems. - Develop data solutions on Databricks using established engineering patterns, reusable frameworks, and enterprise standards. - Ensure reliable, high-quality, and performant data delivery across batch and, where relevant, streaming use cases.
RESPONSIBILITIES: - Design, develop, test, and maintain scalable data pipelines and integrations using Databricks, PySpark, and SQL. - Build datasets optimized for analytics, BI, and downstream consumption while ensuring data quality, reconciliation, and production reliability.
- 5+ years of experience in data engineering, data warehousing, or large-scale data platform development. - Strong hands-on experience with Databricks and distributed data processing. - Strong hands-on experience with PySpark for pipeline development and transformation of large datasets.
LICENSES/CERTIFICATIONS: - Preferred - Databricks Certified Data Engineer Associate / Professional - Preferred - AWS or Azure Data Engineering certification
- Strong hands-on experience with Databricks and distributed data processing. - Strong hands-on experience with PySpark for pipeline development and transformation of large datasets. - Strong hands-on experience with SQL, including joins, aggregations, optimization, and analytical data processing.
KEY SKILLS AND COMPETENCIES: - Strong proficiency in PySpark and SQL; Python alone is not sufficient for this role. - Strong understanding of distributed data processing, performance optimization, and scalable pipeline design.
- Strong hands-on experience with PySpark for pipeline development and transformation of large datasets. - Strong hands-on experience with SQL, including joins, aggregations, optimization, and analytical data processing. - Experience building and maintaining data pipelines for batch processing; exposure to streaming is a plus.
NICE TO HAVE: · Experience with streaming technologies such as Spark Structured Streaming or Kafka. · Experience with orchestration and workflow tools in enterprise data environments.
· Experience with orchestration and workflow tools in enterprise data environments. · Experience with Infrastructure as Code, preferably Terraform. · Experience designing and developing API-based integrations.

Job description

View original posting ↗

By clicking the “Apply” button, I understand that my employment application process with Takeda will commence and that the information I provide in my application will be processed in line with Takeda’s Privacy Notice and Terms of Use.  I further attest that all information I submit in my employment application is true to the best of my knowledge.

Job Description

PRIMARY OBJECTIVES: 


  • Build and maintain scalable data pipelines and datasets that support analytics, reporting, and downstream business systems.
  • Develop data solutions on Databricks using established engineering patterns, reusable frameworks, and enterprise standards.
  • Ensure reliable, high-quality, and performant data delivery across batch and, where relevant, streaming use cases.
  • Support Takeda’s data transformation journey through strong engineering practices, collaboration, and scalable platform-aligned development.

                                   

RESPONSIBILITIES: 

  • Design, develop, test, and maintain scalable data pipelines and integrations using Databricks, PySpark, and SQL.
  • Build datasets optimized for analytics, BI, and downstream consumption while ensuring data quality, reconciliation, and production reliability.
  • Work within established data frameworks, design patterns, and reusable components created by other engineering teams.
  • Read, understand, troubleshoot, and extend existing codebases and pipeline logic in line with engineering standards.
  • Collaborate with analytics, product, and business teams to support data models and data products for enterprise use cases.
  • Contribute to unit, integration, and performance testing, documentation, and engineering best practices.
  • Partner with platform, architecture, security, and DevOps teams to deploy and support pipeline solutions in cloud environments.
  • Troubleshoot data and pipeline issues and drive continuous improvement in performance, scalability, and maintainability.

 

SCOPE OF SUPERVISION:

NUMBER SUPERVISED WORKERS

Direct

Indirect

Employees

0-3

0-3

Non-Employees

0-3

0-3

 

 

 

EDUCATION AND EXPERIENCE:  

  • Bachelor’s or Master’s degree in Computer Science, Engineering, Information Systems, or related field.
  • 5+ years of experience in data engineering, data warehousing, or large-scale data platform development.
  • Strong hands-on experience with Databricks and distributed data processing.
  • Strong hands-on experience with PySpark for pipeline development and transformation of large datasets.
  • Strong hands-on experience with SQL, including joins, aggregations, optimization, and analytical data processing.
  • Experience building and maintaining data pipelines for batch processing; exposure to streaming is a plus.
  • Experience working with existing enterprise frameworks, shared libraries, and engineering standards.
  • Experience reading, understanding, debugging, and enhancing existing code developed by other teams.
  • Experience with cloud data platforms such as AWS or Azure.
  • Experience working in agile, cross-functional engineering environments.

 

KEY SKILLS AND COMPETENCIES:  

  • Strong proficiency in PySpark and SQL; Python alone is not sufficient for this role.
  • Strong understanding of distributed data processing, performance optimization, and scalable pipeline design.
  • Ability to work effectively within predefined patterns, frameworks, and architectural guardrails.
  • Strong code reading and code comprehension skills across shared enterprise codebases.
  • Good understanding of data modeling, schema design, and data quality controls.
  • Strong engineering discipline in testing, version control, documentation, and maintainable development.
  • Strong problem-solving skills and ability to troubleshoot production data issues.
  • Effective communication and collaboration with technical and non-technical stakeholders.

 

NICE TO HAVE:

·         Experience with streaming technologies such as Spark Structured Streaming or Kafka.

·         Experience with orchestration and workflow tools in enterprise data environments.

·         Experience with Infrastructure as Code, preferably Terraform.

·         Experience designing and developing API-based integrations.

 

 

LICENSES/CERTIFICATIONS:

  • Preferred - Databricks Certified Data Engineer Associate / Professional
  • Preferred - AWS or Azure Data Engineering certification

 

PHYSICAL DEMANDS: 

·         N/A

 

TRAVEL REQUIREMENTS:

·         Access to transportation to attend meetings.

·         Ability to fly to meetings regionally and globally.


Locations

IND - Bengaluru

Worker Type

Employee

Worker Sub-Type

Regular

Time Type

Full time

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on takeda.wd502.myworkdayjobs.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

IND - Bengaluru

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Oct 7, 2026
Recorded sightings
1
Employer says posted
Oct 5, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.