Back to jobs

Data Engineer (Temporal & Apache Kafka required)

McLean, VA

Pay
Pay amount needs review — pay source
Advanced SQL & Data Modeling: Strong experience with relational databases, dimensional data modeling, and query performance tuning. Infinitive is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. The determination of this range includes various factors not limited to skill set, level, experience, relevant training, and licensure and certifications. Compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range for this role in the U.S. is $90,000 - $154,00.00. Infinitive is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic protected by applicable federal, state, or local law.
Read the full posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Infinitive

What you’ll work on

Full posting

We are seeking an experienced Data Engineer to help design, build, and scale our next-generation event-driven data platforms.

You will work extensively with Apache Kafka for real-time event streaming and Temporal for durable workflow orchestration, building resilient, distributed, stateful workflows and data pipelines.

What you’ll bring

All qualifications

Core experience

  • 4+ years of professional experience in data engineering, backend distributed systems, or software engineering.
  • Demonstrated proficiency with schema definition frameworks (Apache Avro, Protocol Buffers/gRPC, or JSON Schema).
  • Hands-on experience with Temporal (or Cadence): Proven understanding of durable workflows, activities, retries, signals, queries, and long-running distributed task orchestration.
  • Practical experience managing schema evolution, compatibility modes (backward/forward/full), and schema registries (e.g., Confluent Schema Registry, AWS Glue Schema Registry).
  • Deep expertise with Apache Kafka: Practical experience with message partitioning, consumer groups, offset management, and topic design.
  • Experience enforcing data validation rules, contract testing, and data quality checks (e.g., Great Expectations, Pandera, Pydantic, dbt tests).
Qualification wording
4+ years of professional experience in data engineering, backend distributed systems, or software engineering.
Demonstrated proficiency with schema definition frameworks (Apache Avro, Protocol Buffers/gRPC, or JSON Schema).
Hands-on experience with Temporal (or Cadence): Proven understanding of durable workflows, activities, retries, signals, queries, and long-running distributed task orchestration.
Practical experience managing schema evolution, compatibility modes (backward/forward/full), and schema registries (e.g., Confluent Schema Registry, AWS Glue Schema Registry).
Deep expertise with Apache Kafka: Practical experience with message partitioning, consumer groups, offset management, and topic design.
Experience enforcing data validation rules, contract testing, and data quality checks (e.g., Great Expectations, Pandera, Pydantic, dbt tests).

Tools in this posting

  • Go
  • Java
  • Python
  • SQL
  • AWS
  • BigQuery
  • Databricks
  • dbt
  • Kafka
  • Snowflake
  • Spark
  • Redshift
  • PySpark
  • Great_expectations
Source — Tool mentions in context
- Experience enforcing data validation rules, contract testing, and data quality checks (e.g., Great Expectations, Pandera, Pydantic, dbt tests). - Strong programming proficiency in Python (Go or Java is a plus) with clean code, design patterns, and unit/integration testing standards. - Distributed computing experience: Hands-on development with Apache Spark (PySpark/Spark SQL) processing large-scale datasets.
We are seeking an experienced Data Engineer to help design, build, and scale our next-generation event-driven data platforms. In this role, you will be instrumental in bridging high-throughput distributed streaming with complex, fault-tolerant workflow orchestration and strict data governance. You will work extensively with Apache Kafka for real-time event streaming and Temporal for durable workflow orchestration, building resilient, distributed, stateful workflows and data pipelines. A key focus will be designing data schemas and implementing automated validation to enforce reliable data contracts across distributed systems. You will also develop batch and streaming ETL/ELT pipelines using Python, Apache Spark, and modern cloud data warehouses and Lakehouse platforms. You will work extensively with Apache Kafka for real-time event streaming and Temporal for durable workflow orchestration, building resilient, distributed, stateful workflows and data pipelines. A key focus will be designing data schemas and implementing automated validation to enforce reliable data contracts across distributed systems. You will also develop batch and streaming ETL/ELT pipelines using Python, Apache Spark, and modern cloud data warehouses and Lakehouse platforms.
You will work extensively with Apache Kafka for real-time event streaming and Temporal for durable workflow orchestration, building resilient, distributed, stateful workflows and data pipelines. A key focus will be designing data schemas and implementing automated validation to enforce reliable data contracts across distributed systems. You will also develop batch and streaming ETL/ELT pipelines using Python, Apache Spark, and modern cloud data warehouses and Lakehouse platforms. You will work extensively with Apache Kafka for real-time event streaming and Temporal for durable workflow orchestration, building resilient, distributed, stateful workflows and data pipelines. A key focus will be designing data schemas and implementing automated validation to enforce reliable data contracts across distributed systems. You will also develop batch and streaming ETL/ELT pipelines using Python, Apache Spark, and modern cloud data warehouses and Lakehouse platforms. Key Responsibilities
- Resilient Workflow Orchestration: Design and implement durable execution workflows using Temporal to coordinate long-running distributed pipelines, compensate transactions (Saga pattern), and manage cross-system ETL tasks. - Pipeline Development: Build end-to-end batch and near-real-time pipelines using Python, SQL, and Apache Spark / PySpark. - Data Modeling & Warehousing: Design and optimize analytical data models (dimensional/star schema) in modern cloud data warehouses / Lakehouse (e.g., Snowflake, BigQuery, Databricks, Redshift).
- Strong programming proficiency in Python (Go or Java is a plus) with clean code, design patterns, and unit/integration testing standards. - Distributed computing experience: Hands-on development with Apache Spark (PySpark/Spark SQL) processing large-scale datasets. - Advanced SQL & Data Modeling: Strong experience with relational databases, dimensional data modeling, and query performance tuning.
- Distributed computing experience: Hands-on development with Apache Spark (PySpark/Spark SQL) processing large-scale datasets. - Advanced SQL & Data Modeling: Strong experience with relational databases, dimensional data modeling, and query performance tuning. Infinitive is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. The determination of this range includes various factors not limited to skill set, level, experience, relevant training, and licensure and certifications. Compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range for this role in the U.S. is $90,000 - $154,00.00.
- Demonstrated proficiency with schema definition frameworks (Apache Avro, Protocol Buffers/gRPC, or JSON Schema). - Practical experience managing schema evolution, compatibility modes (backward/forward/full), and schema registries (e.g., Confluent Schema Registry, AWS Glue Schema Registry). - Experience enforcing data validation rules, contract testing, and data quality checks (e.g., Great Expectations, Pandera, Pydantic, dbt tests).
- Pipeline Development: Build end-to-end batch and near-real-time pipelines using Python, SQL, and Apache Spark / PySpark. - Data Modeling & Warehousing: Design and optimize analytical data models (dimensional/star schema) in modern cloud data warehouses / Lakehouse (e.g., Snowflake, BigQuery, Databricks, Redshift). - Reliability & Data Quality: Implement automated testing, continuous schema validation, data drift detection, and observability across streaming and batch workflows.
- Practical experience managing schema evolution, compatibility modes (backward/forward/full), and schema registries (e.g., Confluent Schema Registry, AWS Glue Schema Registry). - Experience enforcing data validation rules, contract testing, and data quality checks (e.g., Great Expectations, Pandera, Pydantic, dbt tests). - Strong programming proficiency in Python (Go or Java is a plus) with clean code, design patterns, and unit/integration testing standards.
Key Responsibilities - Stream Processing & Messaging: Architect, deploy, and maintain high-volume distributed data streams using Apache Kafka (producers, consumers, Kafka Connect, Schema Registry). - Data Schema Design & Validation: Establish and enforce schema design standards, versioning strategies, and automated schema validation (e.g., Avro, Protobuf, JSON Schema) to maintain strict data contracts across microservices, streaming consumers, and Lakehouse storage.
- Hands-on experience with Temporal (or Cadence): Proven understanding of durable workflows, activities, retries, signals, queries, and long-running distributed task orchestration. - Deep expertise with Apache Kafka: Practical experience with message partitioning, consumer groups, offset management, and topic design. - Strong background in Data Schema Design & Validation:

Job description

View original posting ↗

About Infinitive
Infinitive is a data and AI consultancy that helps clients modernize, monetize, and operationalize their data to generate lasting value. They pride themselves on their deep industry and technology expertise, ensuring that they drive and sustain the adoption of new capabilities. Infinitive is committed to aligning their team with their clients' culture, ensuring a successful partnership by bringing the right mix of talent and skills for high return on investment.

Infinitive has earned recognition as one of the "Best Small Firms to Work For" by Consulting Magazine, receiving this accolade nine times, most recently in 2026. They have also been honored as a “Top Workplace” by the Washington Post, “Best Places to Work” by the Washington Business Journal, and “Best Places to Work” by Virginia Business.

About the Role

We are seeking an experienced Data Engineer to help design, build, and scale our next-generation event-driven data platforms. In this role, you will be instrumental in bridging high-throughput distributed streaming with complex, fault-tolerant workflow orchestration and strict data governance.

You will work extensively with Apache Kafka for real-time event streaming and Temporal for durable workflow orchestration, building resilient, distributed, stateful workflows and data pipelines. A key focus will be designing data schemas and implementing automated validation to enforce reliable data contracts across distributed systems. You will also develop batch and streaming ETL/ELT pipelines using Python, Apache Spark, and modern cloud data warehouses and Lakehouse platforms.

You will work extensively with Apache Kafka for real-time event streaming and Temporal for durable workflow orchestration, building resilient, distributed, stateful workflows and data pipelines. A key focus will be designing data schemas and implementing automated validation to enforce reliable data contracts across distributed systems. You will also develop batch and streaming ETL/ELT pipelines using Python, Apache Spark, and modern cloud data warehouses and Lakehouse platforms.

Key Responsibilities

  • Stream Processing & Messaging: Architect, deploy, and maintain high-volume distributed data streams using Apache Kafka (producers, consumers, Kafka Connect, Schema Registry).

  • Data Schema Design & Validation: Establish and enforce schema design standards, versioning strategies, and automated schema validation (e.g., Avro, Protobuf, JSON Schema) to maintain strict data contracts across microservices, streaming consumers, and Lakehouse storage.

  • Resilient Workflow Orchestration: Design and implement durable execution workflows using Temporal to coordinate long-running distributed pipelines, compensate transactions (Saga pattern), and manage cross-system ETL tasks.

  • Pipeline Development: Build end-to-end batch and near-real-time pipelines using Python, SQL, and Apache Spark / PySpark.

  • Data Modeling & Warehousing: Design and optimize analytical data models (dimensional/star schema) in modern cloud data warehouses / Lakehouse (e.g., Snowflake, BigQuery, Databricks, Redshift).

  • Reliability & Data Quality: Implement automated testing, continuous schema validation, data drift detection, and observability across streaming and batch workflows.

  • Cross-Functional Collaboration: Partner with software engineers, machine learning engineers, and analysts to define standard schema definitions, data contracts, and production-grade CI/CD release patterns.

Required Experience

  • 4+ years of professional experience in data engineering, backend distributed systems, or software engineering.

  • Hands-on experience with Temporal (or Cadence): Proven understanding of durable workflows, activities, retries, signals, queries, and long-running distributed task orchestration.

  • Deep expertise with Apache Kafka: Practical experience with message partitioning, consumer groups, offset management, and topic design.

  • Strong background in Data Schema Design & Validation:

    • Demonstrated proficiency with schema definition frameworks (Apache Avro, Protocol Buffers/gRPC, or JSON Schema).

    • Practical experience managing schema evolution, compatibility modes (backward/forward/full), and schema registries (e.g., Confluent Schema Registry, AWS Glue Schema Registry).

    • Experience enforcing data validation rules, contract testing, and data quality checks (e.g., Great Expectations, Pandera, Pydantic, dbt tests).

  • Strong programming proficiency in Python (Go or Java is a plus) with clean code, design patterns, and unit/integration testing standards.

  • Distributed computing experience: Hands-on development with Apache Spark (PySpark/Spark SQL) processing large-scale datasets.

  • Advanced SQL & Data Modeling: Strong experience with relational databases, dimensional data modeling, and query performance tuning.


Infinitive is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. The determination of this range includes various factors not limited to skill set, level, experience, relevant training, and licensure and certifications. Compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range for this role in the U.S. is $90,000 - $154,00.00.

Infinitive is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic protected by applicable federal, state, or local law.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.

Complete your application on infinitive.applytojob.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay
Advanced SQL & Data Modeling: Strong experience with relational databases, dimensional data modeling, and query performance tuning. Infinitive is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. The determination of this range includes various factors not limited to skill set, level, experience, relevant training, and licensure and certifications. Compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range for this role in the U.S. is $90,000 - $154,00.00. Infinitive is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic protected by applicable federal, state, or local law.
Location & working pattern

McLean, VA

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 18, 2026
Recorded sightings
17
Last seen by us
Oct 7, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.