Back to jobs
Merck

Associate Specialist , Data Engineering

IND - Telangana - Hyderabad (Hitec City Raidurg)

Work setup
Hybrid statedwork setup source
Flexible Work Arrangements: Hybrid Shift:
Read the full posting

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Pay and employment type unconfirmed

Not confirmed in this saved copy: pay, employment type. Check the full posting

Tools in this posting

  • python
  • sql
  • aws
  • databricks
  • delta
  • docker
Source — Tool mentions in context
Role Overview We are looking for a Data Engineer with 2–4 years of hands-on experience in building and supporting data pipelines, ETL/ELT workflows, and analytics-ready datasets. The ideal candidate should have strong fundamentals in Python, PySpark, SQL, AWS, and Databricks, with practical exposure to data lakes, lakehouse patterns, data warehousing, data quality, and production support. This role is best suited for a hands-on engineer who can work from defined requirements, contribute to reliable data solutions, collaborate with cross-functional teams, and grow into larger ownership over time. What will you do in this role
- Build, enhance, and support batch and streaming data pipelines using defined technical designs and backlog requirements. - Develop and maintain ETL/ELT transformations using Python, PySpark, and SQL across data lake, lakehouse, and warehouse environments. - Work closely with Data Analysts, Data Scientists, senior engineers, tech leads, and product managers to understand requirements and deliver curated, analytics-ready datasets.
- Strong SQL skills for joins, window functions, data profiling, transformations, validations, and performance tuning. - Good working knowledge of Python and PySpark, including Spark fundamentals such as partitioning, shuffle, caching, file formats, debugging, and optimization. - Understanding of dimensional modeling concepts including facts, dimensions, star/snowflake schemas, and slowly changing dimensions (SCD).
Primary Skills: Python, PySpark, SQL, AWS, Databricks, GitHub, Data Lake, ETL/ELT and CI/CD Secondary Skills:
Required Skills: Amazon Web Services (AWS), CI/CD, Databricks Platform, Data ETL, Data Lake, GitHub, PySpark, Python (Programming Language), Structured Query Language (SQL) Preferred Skills:
- Support BI and analytics use cases by applying dimensional modeling concepts such as facts, dimensions, star/snowflake schemas, and slowly changing dimensions (SCD). - Write and tune SQL queries for data profiling, transformations, validations, debugging, and performance improvements. - Use AWS services such as S3, Glue, Lambda, Step Functions, EMR, and CloudWatch to support data engineering workloads while following security practices such as IAM, encryption, and least privilege.
- Hands-on experience with Databricks, Apache Spark, PySpark, and lakehouse concepts such as Delta Lake. - Strong SQL skills for joins, window functions, data profiling, transformations, validations, and performance tuning. - Good working knowledge of Python and PySpark, including Spark fundamentals such as partitioning, shuffle, caching, file formats, debugging, and optimization.
- Implement data quality checks, validations, reconciliations, and basic anomaly checks to improve trust and usability of data outputs. - Run, monitor, and troubleshoot pipelines using orchestration and observability tools such as Databricks Workflows, AWS Step Functions, scheduling, logging, monitoring, and alerting. - Follow engineering practices including unit testing, integration testing, automated data tests, code reviews, and quality gates within CI/CD.
- Write and tune SQL queries for data profiling, transformations, validations, debugging, and performance improvements. - Use AWS services such as S3, Glue, Lambda, Step Functions, EMR, and CloudWatch to support data engineering workloads while following security practices such as IAM, encryption, and least privilege. - Contribute to cloud resource provisioning and environment configuration using Terraform, with guidance from senior engineers.
- 2–4 years of hands-on experience in data engineering, including building or supporting production data pipelines and ETL/ELT workflows. - Practical experience with AWS services such as S3, Glue, Lambda, Step Functions, EMR, and CloudWatch; understanding of IAM, encryption, and cloud security basics. - Hands-on experience with Databricks, Apache Spark, PySpark, and lakehouse concepts such as Delta Lake.
- Exposure to orchestration tools such as Airflow or modern table formats such as Delta, Iceberg. - AWS certification such as Developer or Solutions Architect, or equivalent demonstrated cloud experience. Who we are
- Use GitHub for version control, branching, pull requests, code reviews, and contribution to CI/CD pipelines. - Develop scalable data processing logic on Databricks / Apache Spark using PySpark and lakehouse concepts such as Delta Lake, ACID transactions, and schema evolution. - Use Jupyter/Databricks notebooks for exploration, debugging, and PoCs; convert validated logic into reusable modules, tests, and deployment-ready pipelines.
- Develop scalable data processing logic on Databricks / Apache Spark using PySpark and lakehouse concepts such as Delta Lake, ACID transactions, and schema evolution. - Use Jupyter/Databricks notebooks for exploration, debugging, and PoCs; convert validated logic into reusable modules, tests, and deployment-ready pipelines. - Participate in Agile delivery ceremonies, provide task-level estimates, share progress updates, and raise risks or dependencies early.
- Practical experience with AWS services such as S3, Glue, Lambda, Step Functions, EMR, and CloudWatch; understanding of IAM, encryption, and cloud security basics. - Hands-on experience with Databricks, Apache Spark, PySpark, and lakehouse concepts such as Delta Lake. - Strong SQL skills for joins, window functions, data profiling, transformations, validations, and performance tuning.
- Contribute to cloud resource provisioning and environment configuration using Terraform, with guidance from senior engineers. - Package, deploy, and support workloads using Docker and related runtime configurations, including ECS/Fargate where applicable. - Use GitHub for version control, branching, pull requests, code reviews, and contribution to CI/CD pipelines.
- Exposure to GitHub, CI/CD, code reviews, branching, release practices, and engineering quality standards. - Working knowledge of Docker and Terraform for deployment, runtime configuration, and cloud environment support. - Ability to work in Agile teams, communicate clearly, collaborate with cross-functional stakeholders, and take ownership of assigned deliverables.
Secondary Skills: Dimensional modeling, Docker and Terraform Good to Have
Preferred Skills: Dimensional Modeling, Docker (Software), Terraform Current Employees apply HERE
- Use AWS services such as S3, Glue, Lambda, Step Functions, EMR, and CloudWatch to support data engineering workloads while following security practices such as IAM, encryption, and least privilege. - Contribute to cloud resource provisioning and environment configuration using Terraform, with guidance from senior engineers. - Package, deploy, and support workloads using Docker and related runtime configurations, including ECS/Fargate where applicable.
- Exposure to data quality or testing frameworks such as Collibra, Immuta, and basic awareness of data governance practices including catalog, lineage, and access controls. - Exposure to orchestration tools such as Airflow or modern table formats such as Delta, Iceberg. - AWS certification such as Developer or Solutions Architect, or equivalent demonstrated cloud experience.
All 12 tools

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.

Already applied? Track this application

About applying

Apply opens the employer’s site in a new tab. Add your outcome here after you submit.

Source details & eligibility

Before you apply

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

IND - Telangana - Hyderabad (Hitec City Raidurg)

Our Technology Centers are globally distributed hubs that enable our digital transformation and business outcomes across IT. They bring together diverse teams to collaborate, share best practices, and deliver solutions that save and improve lives. This role is based at our Hyderabad Tech Center and follows a hybrid working model (3 days onsite, 2 days remote). Candidates are expected to reside within commuting distance of the Hyderabad office. Role Overview
More source context
Flexible Work Arrangements: Hybrid Shift:
Work authorization
Domestic VISA Sponsorship: No
Posting history
Status in our records
Active
First seen by us
Sep 9, 2026
Recorded sightings
1

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

Skills in this posting

pythonsqlawsdatabricksdeltadockers3sparkterraformairflowpysparkiceberg

Job description

Job Description

Associate Specialist: Data Engineering

The Opportunity:

Join a global biopharma company with a 130-year legacy and mission to achieve new milestones in healthcare. Be part of a technology-driven, data-led organization supporting a diversified portfolio of medicines, vaccines, and animal health products. Work alongside passionate teams that use data, analytics, and insights to drive decisions and tackle some of the world’s greatest health threats.

Our Technology Centers are globally distributed hubs that enable our digital transformation and business outcomes across IT. They bring together diverse teams to collaborate, share best practices, and deliver solutions that save and improve lives.

This role is based at our Hyderabad Tech Center and follows a hybrid working model (3 days onsite, 2 days remote). Candidates are expected to reside within commuting distance of the Hyderabad office.

Role Overview 

We are looking for a Data Engineer with 2–4 years of hands-on experience in building and supporting data pipelines, ETL/ELT workflows, and analytics-ready datasets. The ideal candidate should have strong fundamentals in Python, PySpark, SQL, AWS, and Databricks, with practical exposure to data lakes, lakehouse patterns, data warehousing, data quality, and production support. This role is best suited for a hands-on engineer who can work from defined requirements, contribute to reliable data solutions, collaborate with cross-functional teams, and grow into larger ownership over time.

What will you do in this role 

  • Build, enhance, and support batch and streaming data pipelines using defined technical designs and backlog requirements.

  • Develop and maintain ETL/ELT transformations using Python, PySpark, and SQL across data lake, lakehouse, and warehouse environments.

  • Work closely with Data Analysts, Data Scientists, senior engineers, tech leads, and product managers to understand requirements and deliver curated, analytics-ready datasets.

  • Implement data quality checks, validations, reconciliations, and basic anomaly checks to improve trust and usability of data outputs.

  • Run, monitor, and troubleshoot pipelines using orchestration and observability tools such as Databricks Workflows, AWS Step Functions, scheduling, logging, monitoring, and alerting.

  • Follow engineering practices including unit testing, integration testing, automated data tests, code reviews, and quality gates within CI/CD.

  • Support BI and analytics use cases by applying dimensional modeling concepts such as facts, dimensions, star/snowflake schemas, and slowly changing dimensions (SCD).

  • Write and tune SQL queries for data profiling, transformations, validations, debugging, and performance improvements.

  • Use AWS services such as S3, Glue, Lambda, Step Functions, EMR, and CloudWatch to support data engineering workloads while following security practices such as IAM, encryption, and least privilege.

  • Contribute to cloud resource provisioning and environment configuration using Terraform, with guidance from senior engineers.

  • Package, deploy, and support workloads using Docker and related runtime configurations, including ECS/Fargate where applicable.

  • Use GitHub for version control, branching, pull requests, code reviews, and contribution to CI/CD pipelines.

  • Develop scalable data processing logic on Databricks / Apache Spark using PySpark and lakehouse concepts such as Delta Lake, ACID transactions, and schema evolution.

  • Use Jupyter/Databricks notebooks for exploration, debugging, and PoCs; convert validated logic into reusable modules, tests, and deployment-ready pipelines.

  • Participate in Agile delivery ceremonies, provide task-level estimates, share progress updates, and raise risks or dependencies early.

  • Create and maintain technical documentation such as pipeline specifications, data contracts, runbooks, and support notes.

What Should you have:

  • Bachelor’s degree in computer science, Engineering, or a related field, or equivalent practical experience.

  • 2–4 years of hands-on experience in data engineering, including building or supporting production data pipelines and ETL/ELT workflows.

  • Practical experience with AWS services such as S3, Glue, Lambda, Step Functions, EMR, and CloudWatch; understanding of IAM, encryption, and cloud security basics.

  • Hands-on experience with Databricks, Apache Spark, PySpark, and lakehouse concepts such as Delta Lake.

  • Strong SQL skills for joins, window functions, data profiling, transformations, validations, and performance tuning.

  • Good working knowledge of Python and PySpark, including Spark fundamentals such as partitioning, shuffle, caching, file formats, debugging, and optimization.

  • Understanding of dimensional modeling concepts including facts, dimensions, star/snowflake schemas, and slowly changing dimensions (SCD).

  • Exposure to GitHub, CI/CD, code reviews, branching, release practices, and engineering quality standards.

  • Working knowledge of Docker and Terraform for deployment, runtime configuration, and cloud environment support.

  • Ability to work in Agile teams, communicate clearly, collaborate with cross-functional stakeholders, and take ownership of assigned deliverables.

Primary Skills:

Python, PySpark, SQL, AWS, Databricks, GitHub, Data Lake, ETL/ELT and CI/CD

Secondary Skills:

Dimensional modeling, Docker and Terraform

Good to Have

  • Exposure to data quality or testing frameworks such as Collibra, Immuta, and basic awareness of data governance practices including catalog, lineage, and access controls.

  • Exposure to orchestration tools such as Airflow or modern table formats such as Delta, Iceberg.

  • AWS certification such as Developer or Solutions Architect, or equivalent demonstrated cloud experience.

Who we are

 We are known as well-known org Inc., Rahway, New Jersey, USA in the United States and Canada and MSD everywhere else. For more than a century, bringing forward medicines and vaccines for many of the world's most challenging diseases. Today, our company continues to be at the forefront of research to deliver innovative health solutions and advance the prevention and treatment of diseases that threaten people and animals around the world.

What we look for

Imagine getting up in the morning for a job as important as helping to save and improve lives around the world. Here, you have that opportunity. You can put your empathy, creativity, digital mastery, or scientific genius to work in collaboration with a diverse group of colleagues who pursue and bring hope to countless people who are battling some of the most challenging diseases of our time. Our team is constantly evolving, so if you are among the intellectually curious, join us—and start making your impact today.

Required Skills:

Amazon Web Services (AWS), CI/CD, Databricks Platform, Data ETL, Data Lake, GitHub, PySpark, Python (Programming Language), Structured Query Language (SQL)

Preferred Skills:

Dimensional Modeling, Docker (Software), Terraform

Current Employees apply HERE

Current Contingent Workers apply HERE

Secondary Language(s) Job Description:

#MSDHYDIT

Search Firm Representatives Please Read Carefully 
Merck & Co., Inc., Rahway, NJ, USA, also known as Merck Sharp & Dohme LLC, Rahway, NJ, USA, does not accept unsolicited assistance from search firms for employment opportunities. All CVs / resumes submitted by search firms to any employee at our company without a valid written search agreement in place for this position will be deemed the sole property of our company.  No fee will be paid in the event a candidate is hired by our company as a result of an agency referral where no pre-existing agreement is in place. Where agency agreements are in place, introductions are position specific. Please, no phone calls or emails. 

Employee Status:

Regular

Relocation:

Domestic

VISA Sponsorship:

No

Travel Requirements:

No Travel Required

Flexible Work Arrangements:

Hybrid

Shift:

Not Indicated

Valid Driving License:

No

Hazardous Material(s):

n/a

Job Posting End Date:

09/16/2026

*A job posting is effective until 11:59:59PM on the day BEFORE the listed job posting end date. Please ensure you apply to a job posting no later than the day BEFORE the job posting end date.