Back to jobs

Senior Data Engineer

Chicago

Pay
$125,000–180,000/yearAnnual period assumed · Location-specific pay — pay source
Advanced modeling techniques. Experience with Data Vault 2.0, Master Data Management, or comparable enterprise modeling methodologies. CHI: $125,000-$180,000 The expected salary range may vary for other locations. Actual salary may vary based on qualifications and experience. Tempus offers a full range of benefits, which may include incentive compensation, restricted stock units, medical and other benefits depending on the position. We are an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
Read the full posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Tempus

What you’ll work on

Full posting
  • Build the pipelines that feed the agents.

  • Own the warehouse and its transformations.

  • Build and maintain our dbt models on BigQuery, along with the tests, documentation, and SQL standards that keep a growing model layer trustworthy.

From the employer’s posting
What You'll Do Build the pipelines that feed the agents. Develop the systems that fetch, parse, and serve both structured and unstructured data in formats optimized for real-time inference — spanning clinical EHR records, high-throughput genomics (NGS), and cardiovascular imaging (Echo, Cath, ECG). Own the warehouse and its transformations. Build and maintain our dbt models on BigQuery, along with the tests, documentation, and SQL standards that keep a growing model layer trustworthy.
Build the pipelines that feed the agents. Develop the systems that fetch, parse, and serve both structured and unstructured data in formats optimized for real-time inference — spanning clinical EHR records, high-throughput genomics (NGS), and cardiovascular imaging (Echo, Cath, ECG). Own the warehouse and its transformations. Build and maintain our dbt models on BigQuery, along with the tests, documentation, and SQL standards that keep a growing model layer trustworthy. Model the multi-modal patient record. Shape the data model across those domains, applying normalized and dimensional design as each one demands, and write the code that enforces it.

What you’ll bring

All qualifications

Core experience

  • Experience running containerized workloads on Kubernetes.
  • Experience with vector databases (Pinecone, Weaviate, pgvector) or graph databases to support RAG and agentic memory.
  • Experience with AWS services alongside GCP in a multi-cloud environment.
  • Experience with Data Vault 2.0, Master Data Management, or comparable enterprise modeling methodologies.
Qualification wording
Kubernetes. Experience running containerized workloads on Kubernetes.
AI data engineering. Experience with vector databases (Pinecone, Weaviate, pgvector) or graph databases to support RAG and agentic memory.
AWS. Experience with AWS services alongside GCP in a multi-cloud environment.
Advanced modeling techniques. Experience with Data Vault 2.0, Master Data Management, or comparable enterprise modeling methodologies.
Education & alternatives
- Primary Requirement: Bachelor's degree in Computer Science, Software Engineering, Data Science, Health Informatics, or a related technical field. - Preferred: Master's degree or Ph.D. in Computer Science (AI/ML or distributed systems focus) or Biomedical Informatics. - Alternative Background: Equivalent professional experience — including a portfolio of significant open-source contributions or industry-recognized technical writing — will be considered.

Tools in this posting

  • Python
  • SQL
  • TypeScript
  • AWS
  • BigQuery
  • dbt
  • Docker
  • Google Cloud (GCP)
  • NoSQL
  • PostgreSQL
  • Redis
  • Terraform
  • Kubernetes
Source — Tool mentions in context
- Healthcare data: Google Cloud Healthcare API FHIR stores, HL7/FHIR, DICOM, Avro - Languages: Python for data pipelines and transforms; TypeScript on Node for platform services and APIs - Application frameworks: NestJS, TypeORM
- Software engineering ability. You write production-quality application and service code, not just pipeline glue — including APIs, tests, and the design work that goes with them. - Python and TypeScript. Python strong enough for production pipelines as well as hands-on data profiling and debugging, plus enough TypeScript or another statically typed language to work confidently in our service and application code. - Event-driven systems. Experience with pub/sub or queue-based architectures and the failure modes that come with them — retries, ordering, idempotency, and dead-letter handling.
- Build the pipelines that feed the agents. Develop the systems that fetch, parse, and serve both structured and unstructured data in formats optimized for real-time inference — spanning clinical EHR records, high-throughput genomics (NGS), and cardiovascular imaging (Echo, Cath, ECG). - Own the warehouse and its transformations. Build and maintain our dbt models on BigQuery, along with the tests, documentation, and SQL standards that keep a growing model layer trustworthy. - Model the multi-modal patient record. Shape the data model across those domains, applying normalized and dimensional design as each one demands, and write the code that enforces it.
- Write the software on top. Build the TypeScript services and APIs that handle agent input and output and coordinate specialized agents, meeting the platform's performance and scalability demands. You are expected to be comfortable in the application codebase, not only in the data layer. - Own the infrastructure your platform runs on. Extend and operate the platform's Google Cloud footprint in Terraform — BigQuery datasets, Pub/Sub, Cloud SQL, Cloud Storage, Memorystore, Secret Manager, and the service-account and IAM model that governs access to clinical data. - Scale across hospital networks. Build for federated networks of hospitals: multi-tenancy, high availability, and performance across hybrid on-prem and cloud environments built for sensitive health-system integrations.
- Warehouse and transformation: BigQuery, dbt - Operational data stores: Cloud SQL (PostgreSQL), Memorystore (Redis), Cloud Storage - Messaging: Pub/Sub with dead-letter queues
- Data engineering depth. Proven track record building and operating production data pipelines that handle structured and unstructured data at scale, with real ownership of reliability and correctness. - Google Cloud fluency. Hands-on experience designing and running workloads on GCP — BigQuery, Pub/Sub, Cloud Storage, Cloud SQL, and Secret Manager — including the IAM and service-account model that controls access to sensitive data. - Analytics engineering. Strong SQL and practical experience with dbt or an equivalent transformation framework, including testing, documentation, and managing a model layer as it grows.
- Google Cloud fluency. Hands-on experience designing and running workloads on GCP — BigQuery, Pub/Sub, Cloud Storage, Cloud SQL, and Secret Manager — including the IAM and service-account model that controls access to sensitive data. - Analytics engineering. Strong SQL and practical experience with dbt or an equivalent transformation framework, including testing, documentation, and managing a model layer as it grows. - Infrastructure practice. Comfort owning infrastructure as code in Terraform, working in containers, and taking responsibility for the operational characteristics of what you deploy.
- AI/ML Orchestration: 1+ years hands-on building with Large Language Models — agentic workflows, RAG, or autonomous tool use. - Data at Scale: Demonstrated experience managing structured (SQL, NoSQL) and unstructured data at a scale of millions of records, ensuring data integrity for downstream AI consumption. Education
We are building the Patient Evaluation Engine: a high-scale, multi-modal healthcare platform where autonomous AI agents reason over clinical data to drive real-time clinical evaluation across federated networks of hospitals. We are looking for a Senior Data Engineer to build and own the data platform underneath it — the pipelines, models, and services that make EHR records, genomic results, and cardiovascular imaging discoverable, trustworthy, and usable by agents. This is a data engineering role at its core, and it asks for two things beyond the usual scope. First, you should be a capable software engineer: the person who builds the pipeline here is the person who writes the service that exposes it, and you will regularly work in our TypeScript application and service code rather than handing that off. Second, you should know cloud infrastructure well, specifically Google Cloud — you will make real decisions about how this platform is deployed, scaled, secured, and paid for, not just what runs on it. Our goal is to move beyond static data warehousing toward a dynamic, "agent-ready" data fabric that supports real-time clinical evaluation at enterprise scale, in a HIPAA-regulated environment. The platform is early and much of it is still being built, which is why we are looking for someone with high ownership and a strong self-starting instinct rather than someone waiting for a fully specified backlog.
- Make the data agent-ready. Build the data access patterns and metadata layers that let AI agents autonomously discover, query, and reason over structured and unstructured datasets, and the retrieval services those agents call. - Write the software on top. Build the TypeScript services and APIs that handle agent input and output and coordinate specialized agents, meeting the platform's performance and scalability demands. You are expected to be comfortable in the application codebase, not only in the data layer. - Own the infrastructure your platform runs on. Extend and operate the platform's Google Cloud footprint in Terraform — BigQuery datasets, Pub/Sub, Cloud SQL, Cloud Storage, Memorystore, Secret Manager, and the service-account and IAM model that governs access to clinical data.
- Decisioning: GoRules ZEN engine for versioned decision models - Cloud: primarily Google Cloud, with some AWS at the edges What We're Looking For
- AI data engineering. Experience with vector databases (Pinecone, Weaviate, pgvector) or graph databases to support RAG and agentic memory. - AWS. Experience with AWS services alongside GCP in a multi-cloud environment. - Advanced modeling techniques. Experience with Data Vault 2.0, Master Data Management, or comparable enterprise modeling methodologies.
You will not have used all of this, and we do not expect you to have. It is here so you know what you would be working in. - Warehouse and transformation: BigQuery, dbt - Operational data stores: Cloud SQL (PostgreSQL), Memorystore (Redis), Cloud Storage
- Application frameworks: NestJS, TypeORM - Infrastructure: Terraform, Docker, Secret Manager, service-account and IAM-based access control - Decisioning: GoRules ZEN engine for versioned decision models
- Messaging: Pub/Sub with dead-letter queues - Healthcare data: Google Cloud Healthcare API FHIR stores, HL7/FHIR, DICOM, Avro - Languages: Python for data pipelines and transforms; TypeScript on Node for platform services and APIs
- Data Engineering: 3+ years focused on data engineering, pipeline ownership, or data modeling, ideally in the healthcare or life sciences domain. - Cloud Infrastructure: 2+ years hands-on building and operating on Google Cloud, with demonstrated ownership of infrastructure decisions rather than consuming someone else's. - Healthcare Domain: 2+ years in HIPAA-regulated environments, with hands-on exposure to EMR integrations (Epic, Cerner) and healthcare data standards.
Bonus Points - Google Cloud Healthcare API. Direct experience with managed FHIR or DICOM stores. - Specialized clinical data. Direct experience with OMOP, DICOM, genomic data models, or longitudinal patient records.
- Analytics engineering. Strong SQL and practical experience with dbt or an equivalent transformation framework, including testing, documentation, and managing a model layer as it grows. - Infrastructure practice. Comfort owning infrastructure as code in Terraform, working in containers, and taking responsibility for the operational characteristics of what you deploy. - Software engineering ability. You write production-quality application and service code, not just pipeline glue — including APIs, tests, and the design work that goes with them.
- Specialized clinical data. Direct experience with OMOP, DICOM, genomic data models, or longitudinal patient records. - Kubernetes. Experience running containerized workloads on Kubernetes. - AI data engineering. Experience with vector databases (Pinecone, Weaviate, pgvector) or graph databases to support RAG and agentic memory.

Job description

View original posting ↗

Passionate about precision medicine and advancing the healthcare industry?

Recent advancements in underlying technology have finally made it possible for AI to impact clinical care in a meaningful way. Tempus' proprietary platform connects an entire ecosystem of real-world evidence to deliver real-time, actionable insights to physicians, providing critical information about the right treatments for the right patients, at the right time.

We are building the Patient Evaluation Engine: a high-scale, multi-modal healthcare platform where autonomous AI agents reason over clinical data to drive real-time clinical evaluation across federated networks of hospitals. We are looking for a Senior Data Engineer to build and own the data platform underneath it — the pipelines, models, and services that make EHR records, genomic results, and cardiovascular imaging discoverable, trustworthy, and usable by agents.


This is a data engineering role at its core, and it asks for two things beyond the usual scope. First, you should be a capable software engineer: the person who builds the pipeline here is the person who writes the service that exposes it, and you will regularly work in our TypeScript application and service code rather than handing that off. Second, you should know cloud infrastructure well, specifically Google Cloud — you will make real decisions about how this platform is deployed, scaled, secured, and paid for, not just what runs on it.


Our goal is to move beyond static data warehousing toward a dynamic, "agent-ready" data fabric that supports real-time clinical evaluation at enterprise scale, in a HIPAA-regulated environment. The platform is early and much of it is still being built, which is why we are looking for someone with high ownership and a strong self-starting instinct rather than someone waiting for a fully specified backlog.

What You'll Do

  • Build the pipelines that feed the agents. Develop the systems that fetch, parse, and serve both structured and unstructured data in formats optimized for real-time inference — spanning clinical EHR records, high-throughput genomics (NGS), and cardiovascular imaging (Echo, Cath, ECG).

  • Own the warehouse and its transformations. Build and maintain our dbt models on BigQuery, along with the tests, documentation, and SQL standards that keep a growing model layer trustworthy.

  • Model the multi-modal patient record. Shape the data model across those domains, applying normalized and dimensional design as each one demands, and write the code that enforces it.

  • Move data through event-driven services. Build and operate the Pub/Sub topics, subscriptions, and dead-letter handling that connect ingestion, evaluation, and result delivery, with the retry and idempotency behavior that reliability at scale requires.

  • Make the data agent-ready. Build the data access patterns and metadata layers that let AI agents autonomously discover, query, and reason over structured and unstructured datasets, and the retrieval services those agents call.

  • Write the software on top. Build the TypeScript services and APIs that handle agent input and output and coordinate specialized agents, meeting the platform's performance and scalability demands. You are expected to be comfortable in the application codebase, not only in the data layer.

  • Own the infrastructure your platform runs on. Extend and operate the platform's Google Cloud footprint in Terraform — BigQuery datasets, Pub/Sub, Cloud SQL, Cloud Storage, Memorystore, Secret Manager, and the service-account and IAM model that governs access to clinical data.

  • Scale across hospital networks. Build for federated networks of hospitals: multi-tenancy, high availability, and performance across hybrid on-prem and cloud environments built for sensitive health-system integrations.

  • Guarantee ground truth. Implement automated solutions to monitor data quality and lineage with strict traceability back to source systems, ensuring "ground truth" for agentic evaluations.

  • Instrument for trust. Build the observability, error tracking, and human-in-the-loop checkpoints that make automated clinical evaluation transparent and debuggable.

  • Raise the standard around you. Partner with clinical, analytics, and platform engineering teams on data modeling standards, governance, and practices for maintaining data integrity in a HIPAA-regulated environment.

How You Work

We care about these as much as the technical checklist.

  • High ownership. You own what you build all the way into production — you care whether it stays up, you chase root causes instead of symptoms, and you do not treat the deploy boundary as the end of your responsibility.

  • Self-starter. The problem space is genuinely open. You are comfortable identifying the most valuable next thing and starting on it without a fully specified ticket, and you surface ambiguity early rather than stalling on it.

  • Collaborative. You work directly with clinical, analytics, and platform engineering partners. You write things down, you explain trade-offs to non-specialists, and you make the people around you faster.

  • Quick to add impact and value. You bias toward shipping something real and incremental early over long design cycles, and you look for the change that moves the platform now.

Our Stack

You will not have used all of this, and we do not expect you to have. It is here so you know what you would be working in.

  • Warehouse and transformation: BigQuery, dbt

  • Operational data stores: Cloud SQL (PostgreSQL), Memorystore (Redis), Cloud Storage

  • Messaging: Pub/Sub with dead-letter queues

  • Healthcare data: Google Cloud Healthcare API FHIR stores, HL7/FHIR, DICOM, Avro

  • Languages: Python for data pipelines and transforms; TypeScript on Node for platform services and APIs

  • Application frameworks: NestJS, TypeORM

  • Infrastructure: Terraform, Docker, Secret Manager, service-account and IAM-based access control

  • Decisioning: GoRules ZEN engine for versioned decision models

  • Cloud: primarily Google Cloud, with some AWS at the edges

What We're Looking For

  • Data engineering depth. Proven track record building and operating production data pipelines that handle structured and unstructured data at scale, with real ownership of reliability and correctness.

  • Google Cloud fluency. Hands-on experience designing and running workloads on GCP — BigQuery, Pub/Sub, Cloud Storage, Cloud SQL, and Secret Manager — including the IAM and service-account model that controls access to sensitive data.

  • Analytics engineering. Strong SQL and practical experience with dbt or an equivalent transformation framework, including testing, documentation, and managing a model layer as it grows.

  • Infrastructure practice. Comfort owning infrastructure as code in Terraform, working in containers, and taking responsibility for the operational characteristics of what you deploy.

  • Software engineering ability. You write production-quality application and service code, not just pipeline glue — including APIs, tests, and the design work that goes with them.

  • Python and TypeScript. Python strong enough for production pipelines as well as hands-on data profiling and debugging, plus enough TypeScript or another statically typed language to work confidently in our service and application code.

  • Event-driven systems. Experience with pub/sub or queue-based architectures and the failure modes that come with them — retries, ordering, idempotency, and dead-letter handling.

  • Interoperability standards. Working knowledge of HL7, FHIR, and Epic/Cerner data structures, along with DICOM and genomic data formats.

  • Regulatory fluency. Familiarity with building secure, resilient systems under HIPAA and SOC 2.

Experience Requirements

  • Total Professional Experience: 5+ years building data-intensive software systems in production.

  • Data Engineering: 3+ years focused on data engineering, pipeline ownership, or data modeling, ideally in the healthcare or life sciences domain.

  • Cloud Infrastructure: 2+ years hands-on building and operating on Google Cloud, with demonstrated ownership of infrastructure decisions rather than consuming someone else's.

  • Healthcare Domain: 2+ years in HIPAA-regulated environments, with hands-on exposure to EMR integrations (Epic, Cerner) and healthcare data standards.

  • AI/ML Orchestration: 1+ years hands-on building with Large Language Models — agentic workflows, RAG, or autonomous tool use.

  • Data at Scale: Demonstrated experience managing structured (SQL, NoSQL) and unstructured data at a scale of millions of records, ensuring data integrity for downstream AI consumption.

Education

  • Primary Requirement: Bachelor's degree in Computer Science, Software Engineering, Data Science, Health Informatics, or a related technical field.

  • Preferred: Master's degree or Ph.D. in Computer Science (AI/ML or distributed systems focus) or Biomedical Informatics.

  • Alternative Background: Equivalent professional experience — including a portfolio of significant open-source contributions or industry-recognized technical writing — will be considered.

Bonus Points

  • Google Cloud Healthcare API. Direct experience with managed FHIR or DICOM stores.

  • Specialized clinical data. Direct experience with OMOP, DICOM, genomic data models, or longitudinal patient records.

  • Kubernetes. Experience running containerized workloads on Kubernetes.

  • AI data engineering. Experience with vector databases (Pinecone, Weaviate, pgvector) or graph databases to support RAG and agentic memory.

  • AWS. Experience with AWS services alongside GCP in a multi-cloud environment.

  • Advanced modeling techniques. Experience with Data Vault 2.0, Master Data Management, or comparable enterprise modeling methodologies.

CHI: $125,000-$180,000

The expected salary range may vary for other locations. Actual salary may vary based on qualifications and experience. Tempus offers a full range of benefits, which may include incentive compensation, restricted stock units, medical and other benefits depending on the position.

We are an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. 

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.

Complete your application on tempus.wd5.myworkdayjobs.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay
Advanced modeling techniques. Experience with Data Vault 2.0, Master Data Management, or comparable enterprise modeling methodologies. CHI: $125,000-$180,000 The expected salary range may vary for other locations. Actual salary may vary based on qualifications and experience. Tempus offers a full range of benefits, which may include incentive compensation, restricted stock units, medical and other benefits depending on the position. We are an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
Location & working pattern

Chicago

- Own the infrastructure your platform runs on. Extend and operate the platform's Google Cloud footprint in Terraform — BigQuery datasets, Pub/Sub, Cloud SQL, Cloud Storage, Memorystore, Secret Manager, and the service-account and IAM model that governs access to clinical data. - Scale across hospital networks. Build for federated networks of hospitals: multi-tenancy, high availability, and performance across hybrid on-prem and cloud environments built for sensitive health-system integrations. - Guarantee ground truth. Implement automated solutions to monitor data quality and lineage with strict traceability back to source systems, ensuring "ground truth" for agentic evaluations.
Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 3, 2026
Recorded sightings
70
Last seen by us
Oct 8, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.