Sr. Data Platform Engineer
Seattle
- Pay
- Salary not listed in the saved posting
- Work setup
- Unconfirmed
- Employment
- Unconfirmed
What you’ll work on
Full postingOwn day-to-day operations of our data infrastructure
Maintain, expand and optimize our postgres database and Iceberg datalake
Partner with product on new feature-driven datasets.
From the employer’s posting
Continue our migration of pipeline orchestration to Prefect Own day-to-day operations of our data infrastructure Extend the pipeline to include new data sources and transformations
Extend the pipeline to include new data sources and transformations Maintain, expand and optimize our postgres database and Iceberg datalake Create the 'connective tissue' for data at Aarden
Cross-team integration Partner with product on new feature-driven datasets. Collaborate with the ML/analytics team to close the loop: anomaly detection → ticket → fix → validation → promotion to production
Tools in this posting
- Python
- SQL
- AWS
- Iceberg
- PostgreSQL
- Prefect
- Airflow
- Dagster
- PySpark
- TypeScript
- DuckDB
Source — Tool mentions in context
Must-have - Have strong Python skills & are comfortable with PySpark or similar distributed data processing - Have a strong sense of how the data you’re working with impacts the end-user
Our Stack Languages: Python and SQL. TypeScript/Node is a plus for our Application layer and AWS ingest paths. Orchestration & compute
- Prefect 3 for pipeline orchestration (YAML/config-driven flows, retries, logging) - Coiled for elastic EC2 workers on GDAL-heavy and batch Python jobs - Wherobots (managed Apache Sedona / PySpark) for large-scale spatial joins, parcel ingest, and lakehouse work
- Extend the pipeline to include new data sources and transformations - Maintain, expand and optimize our postgres database and Iceberg datalake - Create the 'connective tissue' for data at Aarden
- Have experience with geospatial data (GeoParquet, PostGIS, Apache Sedona, or similar) - Have worked with table formats like Apache Iceberg and lakehouse architectures - Have worked on workflow orchestration (Prefect, Airflow, Dagster, or similar)
- Wherobots (managed Apache Sedona / PySpark) for large-scale spatial joins, parcel ingest, and lakehouse work Data lake & formats: Apache Iceberg, Cloud-Optimized GeoTIFF (COG), and PMTiles. Queried with PySpark and DuckDB. Geospatial: GDAL, rasterio, GeoPandas, and tippecanoe. Large-scale spatial work runs on Sedona/Spark via Wherobots.
Geospatial: GDAL, rasterio, GeoPandas, and tippecanoe. Large-scale spatial work runs on Sedona/Spark via Wherobots. Databases & serving: PostgreSQL + PostGIS (and pgvector on the app side) as the production store. Working at Aarden
- Have worked with table formats like Apache Iceberg and lakehouse architectures - Have worked on workflow orchestration (Prefect, Airflow, Dagster, or similar) - Are comfortable working in a git-based, CI-friendly workflow
- Coiled for elastic EC2 workers on GDAL-heavy and batch Python jobs - Wherobots (managed Apache Sedona / PySpark) for large-scale spatial joins, parcel ingest, and lakehouse work Data lake & formats: Apache Iceberg, Cloud-Optimized GeoTIFF (COG), and PMTiles. Queried with PySpark and DuckDB.
About aarden-ai
Aarden is a land intelligence platform that helps landowners, investors, and developers figure out what a piece of land can actually be used for, and how to market it.
In the employer’s words · Read in context
Job description
About Us
The role
What you’ll do
- Continue our migration of pipeline orchestration to Prefect
- Own day-to-day operations of our data infrastructure
- Extend the pipeline to include new data sources and transformations
- Maintain, expand and optimize our postgres database and Iceberg datalake
- Create the 'connective tissue' for data at Aarden
- Partner with product on new feature-driven datasets.
- Collaborate with the ML/analytics team to close the loop: anomaly detection → ticket → fix → validation → promotion to production
- Develop cross-team tooling/infra to keep GitHub, Notion, Linear, and Slack connected so pipeline issues, docs, and fixes stay linked
- Implement run-over-run data observability (row counts, key column distributions) to catch anomalies and bugs
- Expose accuracy/quality metrics as first-class artifacts so changes can be evaluated automatically, by a human or an agent
- Write and maintain AI-context documentation (schema docs, pipeline architecture, known patterns/quirks, "what not to do")
You might be a good fit if you…
- Have strong Python skills & are comfortable with PySpark or similar distributed data processing
- Have a strong sense of how the data you’re working with impacts the end-user
- Are curious and excited about AI and the impact it can have on our ways of working as developers
- Have experience with geospatial data (GeoParquet, PostGIS, Apache Sedona, or similar)
- Have worked with table formats like Apache Iceberg and lakehouse architectures
- Have worked on workflow orchestration (Prefect, Airflow, Dagster, or similar)
- Are comfortable working in a git-based, CI-friendly workflow
- Have worked in full-stack environments, where your work can directly impact the application layer
- Have experience with Apache Sedona or other cloud spatial-compute platforms
- Have built observability/logging layers for data pipelines (not just app services)
- Have experience with property, parcel, real estate, or land data specifically
- Have experience using AI agents to improve data architecture in a real production codebase
- Experience in real estate/land, energy, forestry, or agriculture tech
Our Stack
- Prefect 3 for pipeline orchestration (YAML/config-driven flows, retries, logging)
- Coiled for elastic EC2 workers on GDAL-heavy and batch Python jobs
- Wherobots (managed Apache Sedona / PySpark) for large-scale spatial joins, parcel ingest, and lakehouse work
Working at Aarden
- At least 2 in-person days per week at our office in Capitol Hill | We’ve found that while heads-down time at home is fantastic for task-related productivity, in-person time is magic for longer-form productivity. Our in-person days are used to plan, troubleshoot, and check-in with each other on progress and questions. Expect team lunches and whiteboarding.
- Focused ownership in your role | The rest of the team is here to help you and cares deeply about the long-term functionality of our applications. With that said, we’ll be looking to you to own your lane, go deep, and develop a strong stance on what it takes to make our applications best-in-class.
- Dedicated monthly AI tooling budget | We’re in a golden era of AI-powered developer tooling. We strongly encourage augmenting your output with AI tools, and have a dedicated & flexible budget for every team member to support that setup. We care about what you ship, not how.
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
- Ask the employer about the salary range before committing time to the process.
Complete your application on jobs.gem.com. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
No pay amount identified in the saved description.
- Location & working pattern
Seattle
Working pattern and location restrictions need checking in the full posting.
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Jul 28, 2026
- Recorded sightings
- 20
- Last seen by us
- Oct 1, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.