Staff Data Engineer
San Jose, California, United States
We’re checking whether it has closed. This saved copy remains available for reference.
- Pay
Pay amount needs review — pay source
If you want to hear more about what we're building before you apply, you can't. Apply, and we'll tell you plenty. At Archer we aim to attract, retain, and motivate talent that possess the skills and leadership necessary to grow our business. We drive a pay-for-performance culture and reward performance that supports the Company’s business strategy. For this position we are targeting a base pay between $175,00.00 - $215,000.00 . Actual compensation offered will be determined by factors such as job-related knowledge, skills, and experience. Archer is committed to working with and providing reasonable accommodations to job applicants with physical or mental disabilities, and those with sincerely held religious beliefs. Applicants who may require reasonable accommodation for any part of the application or hiring process should provide their name and contact information to Archer’s People Team at people@archer.com. Reasonable accommodations will be determined on a case-by-case basis.
Read the full posting- Work setup
- Unconfirmed
- Employment
- Unconfirmed
Before you apply
- Sponsorship
Visa sponsorship not confirmed — sponsorship source
Certain positions may be eligible for visa sponsorship.
Read the full posting
What you’ll work on
Full postingYou will own the pipelines, storage systems, and data quality mechanisms that sit upstream of our ML platform, ensuring models train on clean, high-throughput, well-governed data.
Build and operate the data lakehouse — defining table formats (Iceberg, Paimon, Parquet), partitioning strategies, and compaction policies optimized for ML consumption patterns.
Work closely with AI researchers, platform engineers, and software engineers to understand data access patterns, optimize query performance, and unblock training runs.
From the employer’s posting
Our sights are set high and our problems are hard, and we believe that diversity in the workplace is what makes us smarter, drives better insights, and will ultimately lift us all to success. We are dedicated to cultivating an equitable and inclusive environment that embraces our differences, and supports and celebrates all of our team members. Archer is best known for Midnight, our electric aircraft. This role isn't on that program. The AI Products Org is building software for the general aviation industry. As a Staff Data Engineer on the AI Platform team, you will design, build, and operate the data infrastructure that powers large-scale model training and inference. You will own the pipelines, storage systems, and data quality mechanisms that sit upstream of our ML platform, ensuring models train on clean, high-throughput, well-governed data. What You’ll Do:
Design and maintain high-throughput, fault-tolerant ingestion and transformation pipelines that feed training workloads at scale, with a focus on latency, throughput, and correctness. Build and operate the data lakehouse — defining table formats (Iceberg, Paimon, Parquet), partitioning strategies, and compaction policies optimized for ML consumption patterns. Instrument pipelines with data quality checks, lineage tracking, and anomaly detection so that model failures trace back to data problems quickly.
Partner with ML engineers to define feature stores, dataset versioning, and experiment-to-production data contracts; integrate with tools like MLflow for dataset and artifact tracking. Work closely with AI researchers, platform engineers, and software engineers to understand data access patterns, optimize query performance, and unblock training runs. What You Need:
What you’ll bring
All qualificationsCore experience
- 5+ years of professional data engineering experience excluding internships.
- Experience with CDC (Change Data Capture) replication from transactional systems (i.e.
- Hands-on experience building production pipelines with tools like Apache Spark, Flink, Airflow, dbt, or similar batch/streaming frameworks.
- Prior experience in aerospace, aviation, or other safety-critical domains where data lineage and auditability are non-negotiable.
- Deep proficiency with columnar formats (i.e.
- Experience with high-throughput messaging systems such as Apache Pulsar or Kafka for real-time data ingestion.
Qualification wording
5+ years of professional data engineering experience excluding internships.
Experience with CDC (Change Data Capture) replication from transactional systems (i.e. PostgreSQL → Lakehouse via Debezium or Airbyte).
Hands-on experience building production pipelines with tools like Apache Spark, Flink, Airflow, dbt, or similar batch/streaming frameworks.
Prior experience in aerospace, aviation, or other safety-critical domains where data lineage and auditability are non-negotiable.
Deep proficiency with columnar formats (i.e. Parquet), open table formats (Apache Iceberg or Apache Paimon), and object storage systems (S3 or equivalent).
Experience with high-throughput messaging systems such as Apache Pulsar or Kafka for real-time data ingestion.
Tools in this posting
- SQL
- AWS
- dbt
- Iceberg
- Kafka
- MLflow
- PostgreSQL
- S3
- Spark
- Airflow
- Airbyte
- Docker
- Kubernetes
- Prefect
- Dagster
Source — Tool mentions in context
- Experience with high-throughput messaging systems such as Apache Pulsar or Kafka for real-time data ingestion. - Strong SQL skills; experience with StarRocks for large-scale analytical queries and real-time analytics over the lakehouse. - Familiarity with AWS data services (S3, Glue, EMR) and containerized workloads (Docker/Kubernetes) in production environments. Airflow/Prefect/Dagster
- Strong SQL skills; experience with StarRocks for large-scale analytical queries and real-time analytics over the lakehouse. - Familiarity with AWS data services (S3, Glue, EMR) and containerized workloads (Docker/Kubernetes) in production environments. Airflow/Prefect/Dagster Bonus Qualifications:
- BS/MS/PhD in Computer Science, Data Engineering, Software Engineering, or a related field. - Hands-on experience building production pipelines with tools like Apache Spark, Flink, Airflow, dbt, or similar batch/streaming frameworks. - Deep proficiency with columnar formats (i.e. Parquet), open table formats (Apache Iceberg or Apache Paimon), and object storage systems (S3 or equivalent).
- Hands-on experience building production pipelines with tools like Apache Spark, Flink, Airflow, dbt, or similar batch/streaming frameworks. - Deep proficiency with columnar formats (i.e. Parquet), open table formats (Apache Iceberg or Apache Paimon), and object storage systems (S3 or equivalent). - Experience with high-throughput messaging systems such as Apache Pulsar or Kafka for real-time data ingestion.
- Deep proficiency with columnar formats (i.e. Parquet), open table formats (Apache Iceberg or Apache Paimon), and object storage systems (S3 or equivalent). - Experience with high-throughput messaging systems such as Apache Pulsar or Kafka for real-time data ingestion. - Strong SQL skills; experience with StarRocks for large-scale analytical queries and real-time analytics over the lakehouse.
- Instrument pipelines with data quality checks, lineage tracking, and anomaly detection so that model failures trace back to data problems quickly. - Partner with ML engineers to define feature stores, dataset versioning, and experiment-to-production data contracts; integrate with tools like MLflow for dataset and artifact tracking. - Work closely with AI researchers, platform engineers, and software engineers to understand data access patterns, optimize query performance, and unblock training runs.
Bonus Qualifications: - Experience with CDC (Change Data Capture) replication from transactional systems (i.e. PostgreSQL → Lakehouse via Debezium or Airbyte). - Exposure to audio or time-series data pipelines, including preprocessing for ASR or speech model training.
Job description
Headquartered in Silicon Valley, California, Archer is a leader in the next-gen aerospace sector building an end-to-end advanced air mobility platform that delivers air taxis, unmanned aircraft systems (“UAS”), aviation-related physical artificial intelligence (“AI”) solutions, and other technologies to customers worldwide across the commercial aerospace and defense sectors.
Our sights are set high and our problems are hard, and we believe that diversity in the workplace is what makes us smarter, drives better insights, and will ultimately lift us all to success. We are dedicated to cultivating an equitable and inclusive environment that embraces our differences, and supports and celebrates all of our team members.
Archer is best known for Midnight, our electric aircraft. This role isn't on that program. The AI Products Org is building software for the general aviation industry. As a Staff Data Engineer on the AI Platform team, you will design, build, and operate the data infrastructure that powers large-scale model training and inference. You will own the pipelines, storage systems, and data quality mechanisms that sit upstream of our ML platform, ensuring models train on clean, high-throughput, well-governed data.
What You’ll Do:
- Design and maintain high-throughput, fault-tolerant ingestion and transformation pipelines that feed training workloads at scale, with a focus on latency, throughput, and correctness.
- Build and operate the data lakehouse — defining table formats (Iceberg, Paimon, Parquet), partitioning strategies, and compaction policies optimized for ML consumption patterns.
- Instrument pipelines with data quality checks, lineage tracking, and anomaly detection so that model failures trace back to data problems quickly.
- Partner with ML engineers to define feature stores, dataset versioning, and experiment-to-production data contracts; integrate with tools like MLflow for dataset and artifact tracking.
- Work closely with AI researchers, platform engineers, and software engineers to understand data access patterns, optimize query performance, and unblock training runs.
- 5+ years of professional data engineering experience excluding internships.
- BS/MS/PhD in Computer Science, Data Engineering, Software Engineering, or a related field.
- Hands-on experience building production pipelines with tools like Apache Spark, Flink, Airflow, dbt, or similar batch/streaming frameworks.
- Deep proficiency with columnar formats (i.e. Parquet), open table formats (Apache Iceberg or Apache Paimon), and object storage systems (S3 or equivalent).
- Experience with high-throughput messaging systems such as Apache Pulsar or Kafka for real-time data ingestion.
- Strong SQL skills; experience with StarRocks for large-scale analytical queries and real-time analytics over the lakehouse.
- Familiarity with AWS data services (S3, Glue, EMR) and containerized workloads (Docker/Kubernetes) in production environments. Airflow/Prefect/Dagster
Bonus Qualifications:
- Experience with CDC (Change Data Capture) replication from transactional systems (i.e. PostgreSQL → Lakehouse via Debezium or Airbyte).
- Exposure to audio or time-series data pipelines, including preprocessing for ASR or speech model training.
- Prior experience in aerospace, aviation, or other safety-critical domains where data lineage and auditability are non-negotiable.
If you want to hear more about what we're building before you apply, you can't. Apply, and we'll tell you plenty.
At Archer we aim to attract, retain, and motivate talent that possess the skills and leadership necessary to grow our business. We drive a pay-for-performance culture and reward performance that supports the Company’s business strategy. For this position we are targeting a base pay between $175,00.00 - $215,000.00 . Actual compensation offered will be determined by factors such as job-related knowledge, skills, and experience.
Your next step
Check the employer’s posting for the current role and application details.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
If you want to hear more about what we're building before you apply, you can't. Apply, and we'll tell you plenty. At Archer we aim to attract, retain, and motivate talent that possess the skills and leadership necessary to grow our business. We drive a pay-for-performance culture and reward performance that supports the Company’s business strategy. For this position we are targeting a base pay between $175,00.00 - $215,000.00 . Actual compensation offered will be determined by factors such as job-related knowledge, skills, and experience. Archer is committed to working with and providing reasonable accommodations to job applicants with physical or mental disabilities, and those with sincerely held religious beliefs. Applicants who may require reasonable accommodation for any part of the application or hiring process should provide their name and contact information to Archer’s People Team at people@archer.com. Reasonable accommodations will be determined on a case-by-case basis.
- Location & working pattern
San Jose, California, United States
Working pattern and location restrictions need checking in the full posting.
- Work authorization
Information collected and processed as part of any job applications you choose to submit is subject to Archer's Candidate Privacy Policy. Certain positions may be eligible for visa sponsorship. Archer is proud to be an Equal Opportunity employer committed to diversity and inclusivity in the workplace. All aspects of employment are decided on the basis of merit, qualifications, and business needs. We do not discriminate based upon race, color, religion, sex, sexual orientation, age, national origin, disability status, protected veteran status, gender identity or any other characteristic protected by federal, state or local laws.
- Status in our records
- Unknown — awaiting fresh evidence
- First seen by us
- Aug 15, 2026
- Recorded sightings
- 34
- Last seen by us
- Sep 27, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.