Back to jobs

Principal Data Engineer (AWS, Databricks, Ai, ML Flow, Data Architecture, Apache Airflow)

Pune, Mahārāshtra, India

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Mastercard

Tools in this posting

  • Python
  • SQL
  • AWS
  • Azure
  • Databricks
  • Docker
  • Kafka
  • MLflow
  • Spark
  • Terraform
  • Airflow
  • PySpark
  • Kubernetes
Source — Tool mentions in context
• Expert programming skills in Python, PySpark, and modern software engineering practices.
• Advanced SQL expertise, including data modeling, performance tuning, query optimization, and large-scale analytical processing.
Title and Summary Principal Data Engineer (AWS, Databricks, Ai, ML Flow, Data Architecture, Apache Airflow) Overview Mastercard Foundry R&D develops emerging technology solutions and transforms promising concepts into scalable products and platforms. We are seeking a Principal Data Engineer to design and build the data foundations that power advanced analytics, machine learning, generative AI, and agentic AI solutions. This role requires a hands-on technical leader who can turn ambiguous R&D objectives into secure, scalable, and production-ready data systems. You will collaborate with AI Engineers, Data Scientists, Researchers, Software Engineers, and Product Managers to create reusable data capabilities that accelerate experimentation and the delivery of innovative products.
• Experience building and operating cloud-native data platforms in AWS, or Azure, or other enterprise cloud environments.
• Experience with workflow orchestration platforms such as Apache Airflow, Databricks Workflows, AWS Step Functions, Azure Data Factory, or similar technologies.
Preferred Qualifications • Experience supporting machine learning, generative AI, or agentic AI systems. • Experience with MLflow or comparable ML lifecycle tooling. • Experience with Azure Machine Learning, Azure AI Foundry, or Microsoft Fabric. • Experience with Unity Catalog, Microsoft Purview, or comparable catalog and governance platforms. • Experience with Kafka or other event-streaming technologies. • Experience with containerized and cloud-native platforms, including Docker and Kubernetes. • Experience with Terraform or another Infrastructure as Code framework. • Experience integrating structured, semi-structured, unstructured, and streaming data. • Knowledge of responsible AI, model evaluation, and AI platform observability. • Experience in payments, financial services, or another regulated industry. Success in This Role Success requires someone who: • Remains strongly hands-on while providing technical direction. • Can build production-quality systems without introducing unnecessary platform complexity. • Converts experimentation into reusable engineering capabilities. • Designs for security, governance, reliability, and observability from the outset. • Challenges assumptions and validates architectural choices through evidence and prototypes. • Enables AI and product teams rather than taking ownership of data science or model research. • Influences across teams without depending on formal people-management authority.
• Strong hands-on experience with distributed data processing technologies such as Apache Spark and modern lakehouse platforms.

Job description

View original posting ↗

Our Purpose

Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.

Title and Summary

Principal Data Engineer (AWS, Databricks, Ai, ML Flow, Data Architecture, Apache Airflow)

Overview
Mastercard Foundry R&D develops emerging technology solutions and transforms promising concepts into scalable products and platforms.
We are seeking a Principal Data Engineer to design and build the data foundations that power advanced analytics, machine learning, generative AI, and agentic AI solutions. This role requires a hands-on technical leader who can turn ambiguous R&D objectives into secure, scalable, and production-ready data systems.
You will collaborate with AI Engineers, Data Scientists, Researchers, Software Engineers, and Product Managers to create reusable data capabilities that accelerate experimentation and the delivery of innovative products.

Role
What You'll Do
• Lead the architecture, design, and engineering of scalable batch, streaming, and event-driven data platforms.
• Design and build reliable data ingestion, transformation, enrichment, processing, and consumption pipelines.
• Develop governed lakehouse and modern data platform architectures using cloud-native technologies and distributed data processing frameworks.
• Build and operate platforms that support analytics, machine learning, generative AI, and agentic AI workloads.
• Create reusable data products, services, APIs, and self-service capabilities that accelerate experimentation and production delivery.
• Enable the full AI and machine learning lifecycle through robust feature engineering, training, evaluation, deployment, and inference pipelines.
• Optimize data workloads for performance, scalability, reliability, resiliency, and cost efficiency.
• Define engineering standards for data modeling, software development, testing, CI/CD, observability, and operational excellence.
• Implement data governance capabilities including quality controls, lineage, metadata management, cataloging, access control, retention, and compliance.
• Apply security and privacy-by-design principles for sensitive and regulated data environments.
• Evaluate emerging data and AI technologies through prototypes, proof-of-concepts, and technical assessments.
• Translate loosely defined research, innovation, or product requirements into practical architectures and incremental delivery plans.
• Provide hands-on technical leadership through architecture reviews, code reviews, troubleshooting, and engineering mentorship.
• Partner with cross-functional teams to mature successful R&D initiatives into enterprise-grade production solutions.
• Communicate technical decisions, trade-offs, risks, dependencies, and roadmap recommendations to both technical and business stakeholders.
• Drive adoption of modern data engineering practices, platform automation, and platform reliability disciplines.
All About You
Required Qualifications
• Typically 12-18 years of overall career relevant experience, including significant ownership of complex enterprise data engineering solutions.
• Extensive experience designing, building, and operating large-scale production data platforms.
• Expert programming skills in Python, PySpark, and modern software engineering practices.
• Advanced SQL expertise, including data modeling, performance tuning, query optimization, and large-scale analytical processing.
• Strong hands-on experience with distributed data processing technologies such as Apache Spark and modern lakehouse platforms.
• Experience building and operating cloud-native data platforms in AWS, or Azure, or other enterprise cloud environments.
• Strong knowledge of modern data architecture patterns, including data lakes, lakehouses, data meshes, data products, and event-driven architectures.
• Experience with cloud storage, data integration, streaming, serverless, observability, and security services across public cloud platforms.
• Experience building scalable batch and real-time data pipelines.
• Experience with workflow orchestration platforms such as Apache Airflow, Databricks Workflows, AWS Step Functions, Azure Data Factory, or similar technologies.
• Experience implementing CI/CD, automated testing, version control, Infrastructure as Code, and platform automation.
• Strong understanding of data governance, metadata management, lineage, quality frameworks, privacy controls, and access management.
• Experience diagnosing and resolving complex performance, reliability, scalability, and operational challenges.
• Ability to make sound architectural decisions while balancing delivery speed, innovation, maintainability, security, and cost.
• Proven ability to thrive in R&D and innovation-focused environments where priorities and requirements may evolve through experimentation.
• Strong communication, collaboration, technical leadership, and mentoring skills.
• Bachelor's degree in Computer Science, Engineering, Information Systems, or a related technical discipline, or equivalent practical experience.

Preferred Qualifications
• Experience supporting machine learning, generative AI, or agentic AI systems.
• Experience with MLflow or comparable ML lifecycle tooling.
• Experience with Azure Machine Learning, Azure AI Foundry, or Microsoft Fabric.
• Experience with Unity Catalog, Microsoft Purview, or comparable catalog and governance platforms.
• Experience with Kafka or other event-streaming technologies.
• Experience with containerized and cloud-native platforms, including Docker and Kubernetes.
• Experience with Terraform or another Infrastructure as Code framework.
• Experience integrating structured, semi-structured, unstructured, and streaming data.
• Knowledge of responsible AI, model evaluation, and AI platform observability.
• Experience in payments, financial services, or another regulated industry.
Success in This Role
Success requires someone who:
• Remains strongly hands-on while providing technical direction.
• Can build production-quality systems without introducing unnecessary platform complexity.
• Converts experimentation into reusable engineering capabilities.
• Designs for security, governance, reliability, and observability from the outset.
• Challenges assumptions and validates architectural choices through evidence and prototypes.
• Enables AI and product teams rather than taking ownership of data science or model research.
• Influences across teams without depending on formal people-management authority.

Corporate Security Responsibility


All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:

  • Abide by Mastercard’s security policies and practices;

  • Ensure the confidentiality and integrity of the information being accessed;

  • Report any suspected information security violation or breach, and

  • Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.




Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on mastercard.wd1.myworkdayjobs.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Pune, Mahārāshtra, India

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 23, 2026
Recorded sightings
7
Last seen by us
Sep 25, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.