> job detail
A
👽Other
Staff Data Engineer
Archer · San Jose, California, United States
// classified as
Other (Adjacent or hard to classify.)
posted
2d ago
location
San Jose, California, United States
languages
sql
tools
aws, dbt, iceberg
> stack
sqlawsdbticebergkafkamlflowpostgresqls3sparkairflowdbt
> description
<div class="content-intro"><p><span style="font-weight: 400;">Archer is an aerospace company based in San Jose, California building an all-electric vertical takeoff and landing aircraft with a mission to advance the benefits of sustainable air mobility. We are designing, manufacturing, and operating an all-electric aircraft that can carry four passengers while producing minimal noise.</span></p>
<p><span style="font-weight: 400;">Our sights are set high and our problems are hard, and we believe that diversity in the workplace is what makes us smarter, drives better insights, and will ultimately lift us all to success. We are dedicated to cultivating an equitable and inclusive environment that embraces our differences, and supports and celebrates all of our team members.</span></p></div><p><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;">Archer is best known for Midnight, our electric aircraft. This role isn't on that program. The AI Products Org is building software for the general aviation industry. As a Staff Data Engineer on the AI Platform team, you will design, build, and operate the data infrastructure that powers large-scale model training and inference. You will own the pipelines, storage systems, and data quality mechanisms that sit upstream of our ML platform, ensuring models train on clean, high-throughput, well-governed data.</span></p>
<p style="line-height: 1.5;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><strong><br>What You’ll Do:</strong></span></p>
<ul class="p-rich_text_list p-rich_text_list__bullet p-rich_text_list--nested" data-stringify-type="unordered-list" data-list-tree="true" data-indent="0" data-border="0">
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Design and maintain high-throughput, fault-tolerant ingestion and transformation pipelines that feed training workloads at scale, with a focus on latency, throughput, and correctness.</span></li>
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Build and operate the data lakehouse — defining table formats (Iceberg, Paimon, Parquet), partitioning strategies, and compaction policies optimized for ML consumption patterns.</span></li>
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Instrument pipelines with data quality checks, lineage tracking, and anomaly detection so that model failures trace back to data problems quickly.</span></li>
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"> Partner with ML engineers to define feature stores, dataset versioning, and experiment-to-production data contracts; integrate with tools like MLflow for dataset and artifact tracking.</span></li>
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Work closely with AI researchers, platform engineers, and software engineers to understand data access patterns, optimize query performance, and unblock training runs.</span></li>
</ul>
<div class="p-rich_text_section" style="line-height: 1.5;"> </div>
<div class="p-rich_text_section" style="line-height: 1.5;"><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><strong>What You Need:</strong></span></div>
<div class="p-rich_text_section" style="line-height: 1.5;">
<ul class="p-rich_text_list p-rich_text_list__bullet p-rich_text_list--nested" data-stringify-type="unordered-list" data-list-tree="true" data-indent="0" data-border="0">
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">5+ years of professional data engineering experience excluding internships. </span></li>
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">BS/MS/PhD in Computer Science, Data Engineering, Software Engineering, or a related field.</span></li>
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Hands-on experience building production pipelines with tools like Apache Spark, Flink, Airflow, dbt, or similar batch/streaming frameworks.</span></li>
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Deep proficiency with columnar formats (i.e. Parquet), open table formats (Apache Iceberg or Apache Paimon), and object storage systems (S3 or equivalent).</span></li>
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Experience with high-throughput messaging systems such as Apache Pulsar or Kafka for real-time data ingestion.</span></li>
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Strong SQL skills; experience with StarRocks for large-scale analytical queries and real-time analytics over the lakehouse.</span></li>
<li style="font-size: 12pt; line-height: 1.5; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Familiarity with AWS data services (S3, Glue, EMR) and containerized workloads (Docker/Kubernetes) in production environments. Airflow/Prefect/Dagster</span><br><br></li>
</ul>
</div>
<div class="p-rich_text_section" style="line-height: 1.5;"><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><strong><br>Bonus Qualifications:</strong></span></div>
<div class="p-rich_text_section" style="line-height: 1.5;">
<ul class="p-rich_text_list p-rich_text_list__bullet p-rich_text_list--nested" data-stringify-type="unordered-list" data-list-tree="true" data-indent="0" data-border="0">
<li style="font-size: 12pt; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Experience with CDC (Change Data Capture) replication from transactional systems (i.e. PostgreSQL → Lakehouse via Debezium or Airbyte).</span></li>
<li style="font-size: 12pt; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Exposure to audio or time-series data pipelines, including preprocessing for ASR or speech model training.</span></li>
<li style="font-size: 12pt; font-family: helvetica, arial, sans-serif;" data-stringify-indent="0" data-stringify-border="0"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Prior experience in aerospace, aviation, or other safety-critical domains where data lineage and auditability are non-negotiable.</span></li>
</ul>
</div>
<p style="line-height: 1.5;"> </p>
<div class="p-rich_text_section" style="line-height: 1.5;"><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;"><strong>About The Team:</strong></span></div>
<div class="p-rich_text_section" style="line-height: 1.5;"><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;">AI Products is a roughly 100-person software org inside a company of aerospace engineers, and we expect to grow substantially over the next year. You get the resources and momentum of a real org, but the products themselves are early — which means Data Engineers here get the kind of ownership that usually only exists at startups.</span></div>
<div class="p-rich_text_section"> </div>
<p><span style="font-size: 12pt; font-family: helvetica, arial, sans-serif;">If you want to hear more about what we're building before you apply, you can't. Apply, and we'll tell you plenty.</span></p>
<p><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><br>At Archer we aim to attract, retain, and motivate talent that possess the skills and leadership necessary to grow our business. We drive a pay-for-performance culture and reward performance that supports the Company’s business strategy. For this position we are targeting a base pay between $175,00.00 - $215,000.00 . Actual compensation offered will be determined by factors such as job-related knowledge, skills, and experience.</span></p>
<h5><span style="font-weight: 400; font-family: helvetica, arial, sans-serif; font-size: 12pt;">Archer is committed to working with and providing reasonable accommodations to job applicants with physical or mental disabilities, and those with sincerely held religious beliefs. Applicants who may require reasonable accommodation for any part of the application or hiring process should provide their name and contact information to Archer’s People Team at <a href="mailto:people@archer.com" target="_blank">people@archer.com</a>. Reasonable accommodations will be determined on a case-by-case basis.</span></h5>
<p> </p><div class="content-conclusion"><hr>
<h5><span style="font-weight: 400;">Information collected and processed as part of any job applications you choose to submit is subject to Archer's <a href="https://archer.com/legal/personnel-candidate-privacy-policy">Candidate Privacy Policy</a>.</span></h5>
<h5><span style="font-weight: 400;">Archer is unable to provide work visa sponsorship for this position at the present time.</span></h5>
<h5><span style="font-weight: 400;">Archer is proud to be an Equal Opportunity employer committed to diversity and inclusivity in the workplace. All aspects of employment are decided on the basis of merit, qualifications, and business needs. We do not discriminate based upon race, color, religion, sex, sexual orientation, age, national origin, disability status, protected veteran status, gender identity or any other characteristic protected by federal, state or local laws.</span></h5>
<h5><span style="font-weight: 400;">Archer Aviation does not engage with external recruiting agencies/individual recruiters with whom it does not have a prior written agreement. Archer reserves the right to make use of any unsolicited resumes that it receives and bears no responsibility for payment of any fees asserted from the use of unsolicited resumes. If you are a recruiting agency or individual recruiter wishing to do business with Archer, please reach out to <a href="mailto:People@archer.com" target="_blank">People@archer.com</a>. All employment processes are managed by the Archer People Team.</span></h5></div>