Back to jobs

AWS Lakehouse Data Engineer

Atlanta, GA

Pay
Salary not listed in the saved posting
Work setup
Remote stated — work setup source
Clearance: Ability to Obtain Public Trust Location: 100% Remote (prefer DMV) Ability to Obtain Public Trust
Read the full posting
Employment
Unconfirmed
Apply at Delan Associates, Inc

What you’ll work on

Full posting
  • Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning.

  • Work closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission needs and delivery timelines.

  • Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.

From the employer’s posting
Build and Operate Data Pipelines (Batch and Streaming)Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners. Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning. Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
Cross-Team Collaboration and Documentation Work closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission needs and delivery timelines. Maintain high-quality engineering documentation, including architecture diagrams, data models, SOPs, interface specifications, operational runbooks, and secure configuration baselines.
Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning. Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks. Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.

What you’ll bring

All qualifications

Core experience

  • Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years equivalent practical experience in leu of degree.
  • Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
  • Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
  • Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
  • Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
  • Experience with AWS security fundamentals, including IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.
Qualification wording
Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years equivalent practical experience in leu of degree.
Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
Experience with AWS security fundamentals, including IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.

Tools in this posting

  • Python
  • SQL
  • AWS
  • Databricks
  • Delta
  • Iceberg
  • Redshift
  • S3
  • Tableau
  • Terraform
  • Airflow
  • PySpark
  • Power BI
  • Docker
Source — Tool mentions in context
AWS Lakehouse Data Engineer We are seeking an AWS Lakehouse Data Engineer to design, implement, and operate the cloud-native data platform that powers AI/ML, analytics, reporting, and data visualization. You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, providing Databricks-like capabilities while maintaining portability, strong governance, cost efficiency, and operational control. You will also develop scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated cloud provisioning and CI/CD across environments. This role is ideal for an engineer who enjoys platform building, automation, performance optimization, and enabling advanced analytics through trusted, secure, and well-governed data.
- Build and Operate Data Pipelines (Batch and Streaming)Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners. - Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning. - Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift. Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling. Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
- Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet. - Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg. - Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.
Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization. Advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads. Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
Position: AWS Lakehouse Data Engineer Clearance: Ability to Obtain Public Trust
Ability to Obtain Public Trust AWS Lakehouse Data Engineer We are seeking an AWS Lakehouse Data Engineer to design, implement, and operate the cloud-native data platform that powers AI/ML, analytics, reporting, and data visualization. You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, providing Databricks-like capabilities while maintaining portability, strong governance, cost efficiency, and operational control. You will also develop scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated cloud provisioning and CI/CD across environments.
- Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks. - Deliver an AWS-Native Lakehouse Data Platform - Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.
- Deliver an AWS-Native Lakehouse Data Platform - Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization. - Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
- Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg. - Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate. - Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and efficient separation of compute and storage.
- Metadata, Governance, Access Control, Lineage, and Quality - Implement data governance and fine-grained access control using AWS-native services, including AWS Lake Formation, AWS Glue Data Catalog, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), and related security services. - Implement a managed metadata repository for dataset cataloging, ownership, business definitions, tagging, classification, and discoverability.
- Build operational data quality checks for freshness, completeness, uniqueness, validity, consistency, and anomaly detection, and publish measurable SLAs/SLOs. - AWS Automation, CI/CD, and Operations - Implement automated AWS provisioning using Infrastructure as Code (IaC) to create consistent environments and secure-by-default baselines.
- AWS Automation, CI/CD, and Operations - Implement automated AWS provisioning using Infrastructure as Code (IaC) to create consistent environments and secure-by-default baselines. - Build and enhance CI/CD for data pipelines and lakehouse components, including automated tests, security checks, validation gates, packaging, deployment, promotion, and rollback strategies.
SIX (6) years of relevant experience. Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift. Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls. Experience with AWS security fundamentals, including IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices. Experience provisioning AWS resources using IaC and operating data platforms across multiple environments.
Experience with AWS security fundamentals, including IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices. Experience provisioning AWS resources using IaC and operating data platforms across multiple environments. Experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.
Ability to troubleshoot distributed data-processing workloads and optimize performance, reliability, and cost. What Would Be Nice to Have Hands-on experience with Databricks, Delta Lake, or migrating Databricks workloads to AWS-native services and Apache Iceberg. - Experience with AWS Step Functions, Amazon Managed Workflows for Apache Airflow (MWAA), Amazon Kinesis, AWS Database Migration Service (DMS), AWS Lambda, Amazon MSK, or similar ingestion and orchestration services.
What Would Be Nice to Have Hands-on experience with Databricks, Delta Lake, or migrating Databricks workloads to AWS-native services and Apache Iceberg. - Experience with AWS Step Functions, Amazon Managed Workflows for Apache Airflow (MWAA), Amazon Kinesis, AWS Database Migration Service (DMS), AWS Lambda, Amazon MSK, or similar ingestion and orchestration services. - Experience with modern DevOps practices and tools such as Git, Terraform, AWS CloudFormation or AWS CDK, Jenkins, AWS CodePipeline, GitHub Actions, and Docker.
- Experience with AWS Step Functions, Amazon Managed Workflows for Apache Airflow (MWAA), Amazon Kinesis, AWS Database Migration Service (DMS), AWS Lambda, Amazon MSK, or similar ingestion and orchestration services. - Experience with modern DevOps practices and tools such as Git, Terraform, AWS CloudFormation or AWS CDK, Jenkins, AWS CodePipeline, GitHub Actions, and Docker. - Experience integrating lakehouse data with business intelligence and visualization tools such as Amazon QuickSight, Tableau, or Power BI.
- Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization. - Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet. - Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg.
Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling. Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization. Advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.
- Experience with modern DevOps practices and tools such as Git, Terraform, AWS CloudFormation or AWS CDK, Jenkins, AWS CodePipeline, GitHub Actions, and Docker. - Experience integrating lakehouse data with business intelligence and visualization tools such as Amazon QuickSight, Tableau, or Power BI. - Experience using AI-assisted coding tools, such as GitHub Copilot, ChatGPT, Cursor, or Kiro, to accelerate implementation while maintaining code quality, testing, review, privacy, and security controls.

Job description

View original posting ↗

Position: AWS Lakehouse Data Engineer

Clearance: Ability to Obtain Public Trust

Location: 100% Remote (prefer DMV)


Ability to Obtain Public Trust

AWS Lakehouse Data Engineer

We are seeking an AWS Lakehouse Data Engineer to design, implement, and operate the cloud-native data platform that powers AI/ML, analytics, reporting, and data visualization. You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, providing Databricks-like capabilities while maintaining portability, strong governance, cost efficiency, and operational control. You will also develop scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated cloud provisioning and CI/CD across environments.

This role is ideal for an engineer who enjoys platform building, automation, performance optimization, and enabling advanced analytics through trusted, secure, and well-governed data.

What You Will Do

  • Build and Operate Data Pipelines (Batch and Streaming)Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.
  • Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning.
  • Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
  • Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.
  • Deliver an AWS-Native Lakehouse Data Platform
  • Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.
  • Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
  • Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg.
  • Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.
  • Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and efficient separation of compute and storage.
  • Establish standardized development, test, and production environments with consistent configuration and controlled promotion across stages.
  • Metadata, Governance, Access Control, Lineage, and Quality
  • Implement data governance and fine-grained access control using AWS-native services, including AWS Lake Formation, AWS Glue Data Catalog, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), and related security services.
  • Implement a managed metadata repository for dataset cataloging, ownership, business definitions, tagging, classification, and discoverability.
  • Enable end-to-end lineage from source through transformation and consumption to support auditability, impact analysis, and regulatory requirements.
  • Apply policy-based access, least-privilege permissions, row-, column-, and cell-level controls where required, data classification, retention, encryption, and secure data handling.
  • Build operational data quality checks for freshness, completeness, uniqueness, validity, consistency, and anomaly detection, and publish measurable SLAs/SLOs.
  • AWS Automation, CI/CD, and Operations
  • Implement automated AWS provisioning using Infrastructure as Code (IaC) to create consistent environments and secure-by-default baselines.
  • Build and enhance CI/CD for data pipelines and lakehouse components, including automated tests, security checks, validation gates, packaging, deployment, promotion, and rollback strategies.
  • Implement observability with centralized metrics, logs, traces, alerts, dashboards, runbooks, and incident-response procedures.
  • Continuously evaluate platform performance, scalability, reliability, security, and cost, and implement measurable improvements.
  • Cross-Team Collaboration and Documentation
  • Work closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission needs and delivery timelines.
  • Maintain high-quality engineering documentation, including architecture diagrams, data models, SOPs, interface specifications, operational runbooks, and secure configuration baselines.
  • Present technical findings, trade-offs, risks, and recommendations clearly to technical and non-technical stakeholders.

What You Will Need:

Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years equivalent practical experience in leu of degree.

SIX (6) years of relevant experience.

Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.

Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.

Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.

Advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.

Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.

Experience with AWS security fundamentals, including IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.

Experience provisioning AWS resources using IaC and operating data platforms across multiple environments.

Experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.

Ability to troubleshoot distributed data-processing workloads and optimize performance, reliability, and cost.


What Would Be Nice to Have
Hands-on experience with Databricks, Delta Lake, or migrating Databricks workloads to AWS-native services and Apache Iceberg.

  • Experience with AWS Step Functions, Amazon Managed Workflows for Apache Airflow (MWAA), Amazon Kinesis, AWS Database Migration Service (DMS), AWS Lambda, Amazon MSK, or similar ingestion and orchestration services.
  • Experience with modern DevOps practices and tools such as Git, Terraform, AWS CloudFormation or AWS CDK, Jenkins, AWS CodePipeline, GitHub Actions, and Docker.
  • Experience integrating lakehouse data with business intelligence and visualization tools such as Amazon QuickSight, Tableau, or Power BI.
  • Experience using AI-assisted coding tools, such as GitHub Copilot, ChatGPT, Cursor, or Kiro, to accelerate implementation while maintaining code quality, testing, review, privacy, and security controls.
  • Knowledge graph and Graph RAG experience, including graph modeling, ontology and taxonomy alignment, entity resolution, relationship extraction, and hybrid retrieval that combines graph traversal with semantic or vector search.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on delan-associates-inc.breezy.hr. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Atlanta, GA

Clearance: Ability to Obtain Public Trust Location: 100% Remote (prefer DMV) Ability to Obtain Public Trust
More source context
- Experience using AI-assisted coding tools, such as GitHub Copilot, ChatGPT, Cursor, or Kiro, to accelerate implementation while maintaining code quality, testing, review, privacy, and security controls. - Knowledge graph and Graph RAG experience, including graph modeling, ontology and taxonomy alignment, entity resolution, relationship extraction, and hybrid retrieval that combines graph traversal with semantic or vector search.
Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 30, 2026
Recorded sightings
14
Last seen by us
Oct 8, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.