Back to jobs

Principal Data Engineer

Boston, Massachusetts, United States

Pay
Salary not listed in the saved posting
Work setup
Hybrid stated — work setup source
Huge Plus: Prior experience working with healthcare data, such as medical claims, electronic health records (EHR), or billing code systems (ICD-10, CPT). Familiarity with healthcare data standards like HL7 or FHIR is highly desirable. Track record of influencing technical direction beyond immediate team — architecture reviews, establishing org-widestandards, technology selection. Strategic ability to make build-vs-buy decisions and evaluate emerging technologies for platform evolution Location: Boston, MA/Remote - Hybrid Job Type: Full-time, exempt , regular
Read the full posting
Employment
Full-time — employment source
Location: Boston, MA/Remote - Hybrid Job Type: Full-time, exempt , regular Compensation: 175,000-200,000
Read the full posting
Apply at codametrix

What you’ll work on

Full posting
  • Lead key technology decisions, disaster recovery design, architecture reviews, and data engineering standards across teams.

  • Own Terraform-based infrastructure across environments, including jobs, catalogs, schemas, permissions, compute policies, volumes, and external locations.

  • Ensure HIPAA and SOC 2 compliance through PHI masking, audit logging, environment-level data segmentation, user provisioning, compute policies, and cost attribution.

From the employer’s posting
Responsibilities Own the technical strategy and roadmap - Own the vision, architecture, and roadmap for the CodaMetrix data platform, ensuring scalability, reliability, regulatory alignment, and operational excellence. Lead key technology decisions, disaster recovery design, architecture reviews, and data engineering standards across teams. Platform Engineering at Scale - Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming. Own Terraform-based infrastructure across environments, including jobs, catalogs, schemas, permissions, compute policies, volumes, and external locations. Build Jenkins CI/CD workflows for automated testing, tagging, and deployments, while optimizing training pipelines, materialized views, retention policies, and production performance.
Own the technical strategy and roadmap - Own the vision, architecture, and roadmap for the CodaMetrix data platform, ensuring scalability, reliability, regulatory alignment, and operational excellence. Lead key technology decisions, disaster recovery design, architecture reviews, and data engineering standards across teams. Platform Engineering at Scale - Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming. Own Terraform-based infrastructure across environments, including jobs, catalogs, schemas, permissions, compute policies, volumes, and external locations. Build Jenkins CI/CD workflows for automated testing, tagging, and deployments, while optimizing training pipelines, materialized views, retention policies, and production performance. Data Governance & Security - Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions. Ensure HIPAA and SOC 2 compliance through PHI masking, audit logging, environment-level data segmentation, user provisioning, compute policies, and cost attribution.
Platform Engineering at Scale - Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming. Own Terraform-based infrastructure across environments, including jobs, catalogs, schemas, permissions, compute policies, volumes, and external locations. Build Jenkins CI/CD workflows for automated testing, tagging, and deployments, while optimizing training pipelines, materialized views, retention policies, and production performance. Data Governance & Security - Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions. Ensure HIPAA and SOC 2 compliance through PHI masking, audit logging, environment-level data segmentation, user provisioning, compute policies, and cost attribution. Cost Optimization & Operational Excellence - Drive platform cost reduction through compute policy tuning, serverless optimization, reserved pools, materialized view improvements, and remediation of underutilized resources. Monitor Databricks/AWS spend (using tools like CloudZero and AWS Cost Explorer), own cost attribution and budgeting, resolve production incidents, and maintain runbooks to ensure platform SLAs for uptime and performance.

What you’ll bring

All qualifications

Core experience

  • 8+ years of data engineering experience with progressive responsibility
  • 5+ years hands-on experience with the Databricks platform (Unity Catalog, Delta Lake, Structured Streaming, Spark SQL, cluster management, platform administration)
  • Deep understanding of medallion/lakehouse architecture patterns (bronze/silver/gold, SCD2, materialized views, slowly changing dimensions)
  • Demonstrated ability to operate in a HIPAA-regulated environment with PHI handling requirements
  • Experience with disaster recovery architecture — cross-region replication, failover procedures, read-only replicas
  • Experience using AI tools (e.g., Claude, Gemini, Codex) and agentic workflows to augment design, development, and testing processes.

Preferred experience

  • Familiarity with healthcare data standards like HL7 or FHIR is highly desirable.
Qualification wording
8+ years of data engineering experience with progressive responsibility
5+ years hands-on experience with the Databricks platform (Unity Catalog, Delta Lake, Structured Streaming, Spark SQL, cluster management, platform administration)
Deep understanding of medallion/lakehouse architecture patterns (bronze/silver/gold, SCD2, materialized views, slowly changing dimensions)
Demonstrated ability to operate in a HIPAA-regulated environment with PHI handling requirements
Experience with disaster recovery architecture — cross-region replication, failover procedures, read-only replicas
Experience using AI tools (e.g., Claude, Gemini, Codex) and agentic workflows to augment design, development, and testing processes.
Huge Plus: Prior experience working with healthcare data, such as medical claims, electronic health records (EHR), or billing code systems (ICD-10, CPT). Familiarity with healthcare data standards like HL7 or FHIR is highly desirable. Track record of influencing technical direction beyond immediate team — architecture reviews, establishing org-widestandards, technology selection. Strategic ability to make build-vs-buy decisions and evaluate emerging technologies for platform evolution
Education & alternatives
Must-Haves - A degree in Computer Science or a related field (Bachelor's, Master's, or Ph.D.), or an equivalent combination of education and demonstrable professional experience - 8+ years of data engineering experience with progressive responsibility

Tools in this posting

  • Python
  • SQL
  • AWS
  • Databricks
  • Delta
  • Kafka
  • S3
  • Spark
  • Tableau
  • Terraform
  • PySpark
  • Scala
Source — Tool mentions in context
- 5+ years hands-on experience with the Databricks platform (Unity Catalog, Delta Lake, Structured Streaming, Spark SQL, cluster management, platform administration) - Expert proficiency in PySpark and Python; working knowledge of Scala - Expert experience with Terraform for infrastructure-as-code (state management, modular structures, multi-environment deployments, YAML-driven configuration)
- 8+ years of data engineering experience with progressive responsibility - 5+ years hands-on experience with the Databricks platform (Unity Catalog, Delta Lake, Structured Streaming, Spark SQL, cluster management, platform administration) - Expert proficiency in PySpark and Python; working knowledge of Scala
- Proven track record building and maintaining CI/CD pipelines (Jenkins, GitHub Actions) for data platform deployments - Expert SQL skills — complex CTEs, window functions, performance optimization of large-scale queries across petabyte-scale datasets - Strong AWS experience (S3, IAM, Secrets Manager, VPC/Private Link, cross-region replication)
- Data Governance & Security - Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions. Ensure HIPAA and SOC 2 compliance through PHI masking, audit logging, environment-level data segmentation, user provisioning, compute policies, and cost attribution. - Cost Optimization & Operational Excellence - Drive platform cost reduction through compute policy tuning, serverless optimization, reserved pools, materialized view improvements, and remediation of underutilized resources. Monitor Databricks/AWS spend (using tools like CloudZero and AWS Cost Explorer), own cost attribution and budgeting, resolve production incidents, and maintain runbooks to ensure platform SLAs for uptime and performance. - Cross-Functional Enablement & Mentorship - Enable ML, BI, Analytics, DevOps, and customer success and implementation teams with training data pipelines, model-ready datasets, feature store architecture, optimized views, dashboards, Tableau refreshes, infrastructure changes, and tenant onboarding. Mentor engineers through code and design reviews, establish quality standards, and serve as a subject matter expert for data platform engineering across the organization.
- Expert experience with Terraform for infrastructure-as-code (state management, modular structures, multi-environment deployments, YAML-driven configuration) - Expert experience with Apache Kafka / AWS MSK (streaming ingestion, SASL/IAM auth, topic management, cluster migrations) - Deep understanding of medallion/lakehouse architecture patterns (bronze/silver/gold, SCD2, materialized views, slowly changing dimensions)
- Expert SQL skills — complex CTEs, window functions, performance optimization of large-scale queries across petabyte-scale datasets - Strong AWS experience (S3, IAM, Secrets Manager, VPC/Private Link, cross-region replication) - Demonstrated ability to operate in a HIPAA-regulated environment with PHI handling requirements
Overview The Principal Data Engineer is a member of the Data Platform team, reporting to the Director of Machine Learning Engineering and Data. The Data Platform team is responsible for executing the data strategy for the organization, ensuring high-quality external data is ingested into the Lakehouse and realized in powerful insights for internal and external customers, while ensuring ML/AI, BI and customer success teams have the data they need to develop and train their models, build insightful dashboards and design semantic layer. As a Principal Data Engineer (L4), you will serve as a key technical leader for CodaMetrix’s Databricks-based data platform, supporting streaming and batch workloads across 30+ healthcare customers. You will own the platform’s architecture, evolution, and operational excellence, including infrastructure-as-code, CI/CD automation, disaster recovery, cost optimization, data contracts, and build-vs-buy decisions, while enabling ML Engineering, Analytics, DevOps, and Product teams and influencing technical standards beyond the immediate team. This role operates with minimal oversight and is expected to influence technical standards beyond the immediate team. Responsibilities
- Own the technical strategy and roadmap - Own the vision, architecture, and roadmap for the CodaMetrix data platform, ensuring scalability, reliability, regulatory alignment, and operational excellence. Lead key technology decisions, disaster recovery design, architecture reviews, and data engineering standards across teams. - Platform Engineering at Scale - Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming. Own Terraform-based infrastructure across environments, including jobs, catalogs, schemas, permissions, compute policies, volumes, and external locations. Build Jenkins CI/CD workflows for automated testing, tagging, and deployments, while optimizing training pipelines, materialized views, retention policies, and production performance. - Data Governance & Security - Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions. Ensure HIPAA and SOC 2 compliance through PHI masking, audit logging, environment-level data segmentation, user provisioning, compute policies, and cost attribution.
- Platform Engineering at Scale - Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming. Own Terraform-based infrastructure across environments, including jobs, catalogs, schemas, permissions, compute policies, volumes, and external locations. Build Jenkins CI/CD workflows for automated testing, tagging, and deployments, while optimizing training pipelines, materialized views, retention policies, and production performance. - Data Governance & Security - Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions. Ensure HIPAA and SOC 2 compliance through PHI masking, audit logging, environment-level data segmentation, user provisioning, compute policies, and cost attribution. - Cost Optimization & Operational Excellence - Drive platform cost reduction through compute policy tuning, serverless optimization, reserved pools, materialized view improvements, and remediation of underutilized resources. Monitor Databricks/AWS spend (using tools like CloudZero and AWS Cost Explorer), own cost attribution and budgeting, resolve production incidents, and maintain runbooks to ensure platform SLAs for uptime and performance.
- Cost Optimization & Operational Excellence - Drive platform cost reduction through compute policy tuning, serverless optimization, reserved pools, materialized view improvements, and remediation of underutilized resources. Monitor Databricks/AWS spend (using tools like CloudZero and AWS Cost Explorer), own cost attribution and budgeting, resolve production incidents, and maintain runbooks to ensure platform SLAs for uptime and performance. - Cross-Functional Enablement & Mentorship - Enable ML, BI, Analytics, DevOps, and customer success and implementation teams with training data pipelines, model-ready datasets, feature store architecture, optimized views, dashboards, Tableau refreshes, infrastructure changes, and tenant onboarding. Mentor engineers through code and design reviews, establish quality standards, and serve as a subject matter expert for data platform engineering across the organization. Requirements
- Expert proficiency in PySpark and Python; working knowledge of Scala - Expert experience with Terraform for infrastructure-as-code (state management, modular structures, multi-environment deployments, YAML-driven configuration) - Expert experience with Apache Kafka / AWS MSK (streaming ingestion, SASL/IAM auth, topic management, cluster migrations)

Job description

View original posting ↗

CodaMetrix is revolutionizing Revenue Cycle Management with its AI-powered autonomous coding solution, a multi-specialty AI-platform that translates clinical information into accurate sets of medical codes. CodaMetrix’s autonomous coding drives efficiency under fee-for-service and value-based care models and supports improved patient care. We are passionate about getting physicians and healthcare providers away from the keyboard and back to clinical care.

Overview

The Principal Data Engineer is a member of the Data Platform team, reporting to the Director of Machine Learning Engineering and Data. The Data Platform team is responsible for executing the data strategy for the organization, ensuring high-quality external data is ingested into the Lakehouse and realized in powerful insights for internal and external customers, while ensuring ML/AI, BI and customer success teams have the data they need to develop and train their models, build insightful dashboards and design semantic layer. As a Principal Data Engineer (L4), you will serve as a key technical leader for CodaMetrix’s Databricks-based data platform, supporting streaming and batch workloads across 30+ healthcare customers. You will own the platform’s architecture, evolution, and operational excellence, including infrastructure-as-code, CI/CD automation, disaster recovery, cost optimization, data contracts, and build-vs-buy decisions, while enabling ML Engineering, Analytics, DevOps, and Product teams and influencing technical standards beyond the immediate team. This role operates with minimal oversight and is expected to influence technical standards beyond the immediate team.

Responsibilities
  • Own the technical strategy and roadmap - Own the vision, architecture, and roadmap for the CodaMetrix data platform, ensuring scalability, reliability, regulatory alignment, and operational excellence. Lead key technology decisions, disaster recovery design, architecture reviews, and data engineering standards across teams.

  • Platform Engineering at Scale - Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming. Own Terraform-based infrastructure across environments, including jobs, catalogs, schemas, permissions, compute policies, volumes, and external locations. Build Jenkins CI/CD workflows for automated testing, tagging, and deployments, while optimizing training pipelines, materialized views, retention policies, and production performance.

  • Data Governance & Security - Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions. Ensure HIPAA and SOC 2 compliance through PHI masking, audit logging, environment-level data segmentation, user provisioning, compute policies, and cost attribution.

  • Cost Optimization & Operational Excellence - Drive platform cost reduction through compute policy tuning, serverless optimization, reserved pools, materialized view improvements, and remediation of underutilized resources. Monitor Databricks/AWS spend (using tools like CloudZero and AWS Cost Explorer), own cost attribution and budgeting, resolve production incidents, and maintain runbooks to ensure platform SLAs for uptime and performance.

  • Cross-Functional Enablement & Mentorship - Enable ML, BI, Analytics, DevOps, and customer success and implementation teams with training data pipelines, model-ready datasets, feature store architecture, optimized views, dashboards, Tableau refreshes, infrastructure changes, and tenant onboarding. Mentor engineers through code and design reviews, establish quality standards, and serve as a subject matter expert for data platform engineering across the organization.

Requirements

Must-Haves

  • A degree in Computer Science or a related field (Bachelor's, Master's, or Ph.D.), or an equivalent combination of education and demonstrable professional experience

  • 8+ years of data engineering experience with progressive responsibility

  • 5+ years hands-on experience with the Databricks platform (Unity Catalog, Delta Lake, Structured Streaming, Spark SQL, cluster management, platform administration)

  • Expert proficiency in PySpark and Python; working knowledge of Scala

  • Expert experience with Terraform for infrastructure-as-code (state management, modular structures, multi-environment deployments, YAML-driven configuration)

  • Expert experience with Apache Kafka / AWS MSK (streaming ingestion, SASL/IAM auth, topic management, cluster migrations)

  • Deep understanding of medallion/lakehouse architecture patterns (bronze/silver/gold, SCD2, materialized views, slowly changing dimensions)

  • Proven track record building and maintaining CI/CD pipelines (Jenkins, GitHub Actions) for data platform deployments

  • Expert SQL skills — complex CTEs, window functions, performance optimization of large-scale queries across petabyte-scale datasets

  • Strong AWS experience (S3, IAM, Secrets Manager, VPC/Private Link, cross-region replication)

  • Demonstrated ability to operate in a HIPAA-regulated environment with PHI handling requirements

  • Experience with disaster recovery architecture — cross-region replication, failover procedures, read-only replicas

  • Experience using AI tools (e.g., Claude, Gemini, Codex) and agentic workflows to augment design, development, and testing processes.

  • Huge Plus: Prior experience working with healthcare data, such as medical claims, electronic health records (EHR), or billing code systems (ICD-10, CPT). Familiarity with healthcare data standards like HL7 or FHIR is highly desirable. Track record of influencing technical direction beyond immediate team — architecture reviews, establishing org-widestandards, technology selection. Strategic ability to make build-vs-buy decisions and evaluate emerging technologies for platform evolution

Location: Boston, MA/Remote - Hybrid

Job Type: Full-time, exempt , regular

Compensation: 175,000-200,000

What CodaMetrix can offer you:

Learn more about our full-time employee benefits and how we take care of our team.

  • Health Insurance: We cover 80% of the cost of medical and dental insurance and offer vision insurance

  • Retirement: We offer a 401(k) plan that eligible employees can contribute to one month after their first day

  • Flexibility: We have a generous Paid Time Off policy, which is managed but not limited, so you can take the time you need to relax and rejuvenate

  • Development: We provide annual performance evaluations and prioritize working with employees on what their individual growth looks like

  • Recognition: We recognize the outstanding achievements of our team through annual company awards where employees have the opportunity to nominate their peers

  • Office Location: A modern open plan workspace located in the bustling Back Bay neighborhood of Boston

  • Additional Employer Paid Benefits: We offer employer-paid life insurance and short-term and long-term disability insurance

Background Check Notice

All candidates will be required to complete a background check upon acceptance of a job offer.

Equal Employment Opportunity

Our company, as well as our products, are made better because we embrace diverse skills, perspectives, and ideas. CodaMetrix is an Equal Employment Opportunity Employer and all qualified applicants will receive consideration for employment.

Don’t meet every requirement? We invite you to apply anyway. Studies have shown that women, communities of color and historically underrepresented talent are less likely to apply to jobs unless they meet every single qualification. At CodaMetrix we are committed to building a diverse, inclusive and authentic workplace and encourage you to consider joining us.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Boston, Massachusetts, United States

- Huge Plus: Prior experience working with healthcare data, such as medical claims, electronic health records (EHR), or billing code systems (ICD-10, CPT). Familiarity with healthcare data standards like HL7 or FHIR is highly desirable. Track record of influencing technical direction beyond immediate team — architecture reviews, establishing org-widestandards, technology selection. Strategic ability to make build-vs-buy decisions and evaluate emerging technologies for platform evolution Location: Boston, MA/Remote - Hybrid Job Type: Full-time, exempt , regular
Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Jul 19, 2026
Recorded sightings
22
Last seen by us
Oct 5, 2026
Employer says posted
Jul 16, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.