Back to jobs

Member of Technical Staff (Software Engineer, Data Platform)

San Francisco, California, United States

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Perplexity

What you’ll work on

Full posting

The Data Platform team owns the end-to-end data lifecycle at Perplexity, from ingestion through processing, storage, and serving, powering product features, analytics, experimentation, AI workloads, and the company’s data lake.

The team defines the architecture for batch and streaming systems, the orchestration and observability stack, and a self-serve data platform, while thoughtfully combining platforms such as Databricks and Snowflake with open-source technologies including Spark, Kafka, Flink, Airflow, Dagster, dbt, Iceberg, Delta Lake, and ClickHouse.

  • Mentor engineers, review designs, and raise the technical bar for data infrastructure through thoughtful feedback, documentation, and hands-on collaboration.

From the employer’s posting
The Data Platform team owns the end-to-end data lifecycle at Perplexity, from ingestion through processing, storage, and serving, powering product features, analytics, experimentation, AI workloads, and the company’s data lake.
The team defines the architecture for batch and streaming systems, the orchestration and observability stack, and a self-serve data platform, while thoughtfully combining platforms such as Databricks and Snowflake with open-source technologies including Spark, Kafka, Flink, Airflow, Dagster, dbt, Iceberg, Delta Lake, and ClickHouse.
Drive architectural decisions across storage, compute, orchestration, and data APIs, partnering closely with product engineering and data science to align the data ecosystem with Perplexity’s roadmap. Mentor engineers, review designs, and raise the technical bar for data infrastructure through thoughtful feedback, documentation, and hands-on collaboration. Qualifications

What you’ll bring

All qualifications

Core experience

  • 5+ years (Senior) or 8+ years (Staff) of software engineering experience.
  • Strong experience building production data infrastructure systems.
  • Hands-on experience with batch and/or streaming data processing at scale.
  • Deep familiarity with data orchestration systems (Airflow, Dagster, or similar).
  • Proficiency in Python and at least one additional backend language (Go, TypeScript, etc.).
  • Experience supporting ML/AI workflows, training pipelines, or evaluation systems.
Qualification wording
5+ years (Senior) or 8+ years (Staff) of software engineering experience.
Strong experience building production data infrastructure systems.
Hands-on experience with batch and/or streaming data processing at scale.
Deep familiarity with data orchestration systems (Airflow, Dagster, or similar).
Proficiency in Python and at least one additional backend language (Go, TypeScript, etc.).
Experience supporting ML/AI workflows, training pipelines, or evaluation systems.

Tools in this posting

  • Go
  • Python
  • TypeScript
  • Databricks
  • dbt
  • Delta
  • Iceberg
  • Kafka
  • Snowflake
  • Spark
  • Airflow
  • Dagster
  • ClickHouse
Source — Tool mentions in context
- Deep familiarity with data orchestration systems (Airflow, Dagster, or similar). - Proficiency in Python and at least one additional backend language (Go, TypeScript, etc.). - Strong systems thinking around reliability, latency, cost, and complexity tradeoffs.
The Data Platform team owns the end-to-end data lifecycle at Perplexity, from ingestion through processing, storage, and serving, powering product features, analytics, experimentation, AI workloads, and the company’s data lake. The team defines the architecture for batch and streaming systems, the orchestration and observability stack, and a self-serve data platform, while thoughtfully combining platforms such as Databricks and Snowflake with open-source technologies including Spark, Kafka, Flink, Airflow, Dagster, dbt, Iceberg, Delta Lake, and ClickHouse. In this senior/staff role, you will shape architecture, set standards, and drive the long-term technical direction of Perplexity’s data ecosystem.
- Design and operate large-scale batch and streaming data pipelines that directly power Perplexity product features, AI training and evaluation workflows, analytics, and experimentation. - Build event-driven and streaming systems (Kafka, Kinesis, PubSub, or similar) for real-time ingestion, transformation, and delivery, alongside batch frameworks for backfills, aggregations, and offline computation. - Lead the architecture of data orchestration using tools like Airflow or Dagster, owning scheduling, dependency management, retries, SLAs, and end-to-end observability for critical data flows.
- Build event-driven and streaming systems (Kafka, Kinesis, PubSub, or similar) for real-time ingestion, transformation, and delivery, alongside batch frameworks for backfills, aggregations, and offline computation. - Lead the architecture of data orchestration using tools like Airflow or Dagster, owning scheduling, dependency management, retries, SLAs, and end-to-end observability for critical data flows. - Set and enforce guarantees for data correctness, freshness, lineage, and recoverability, designing systems that handle rapid scale growth, partial failures, and evolving schemas without disrupting AI workloads or product experiences.
- Hands-on experience with batch and/or streaming data processing at scale. - Deep familiarity with data orchestration systems (Airflow, Dagster, or similar). - Proficiency in Python and at least one additional backend language (Go, TypeScript, etc.).

Job description

View original posting ↗

About the Role

The Data Platform team owns the end-to-end data lifecycle at Perplexity, from ingestion through processing, storage, and serving, powering product features, analytics, experimentation, AI workloads, and the company’s data lake.

The team defines the architecture for batch and streaming systems, the orchestration and observability stack, and a self-serve data platform, while thoughtfully combining platforms such as Databricks and Snowflake with open-source technologies including Spark, Kafka, Flink, Airflow, Dagster, dbt, Iceberg, Delta Lake, and ClickHouse.

In this senior/staff role, you will shape architecture, set standards, and drive the long-term technical direction of Perplexity’s data ecosystem.

Key Responsibilities

  • Design and operate large-scale batch and streaming data pipelines that directly power Perplexity product features, AI training and evaluation workflows, analytics, and experimentation.

  • Build event-driven and streaming systems (Kafka, Kinesis, PubSub, or similar) for real-time ingestion, transformation, and delivery, alongside batch frameworks for backfills, aggregations, and offline computation.

  • Lead the architecture of data orchestration using tools like Airflow or Dagster, owning scheduling, dependency management, retries, SLAs, and end-to-end observability for critical data flows.

  • Set and enforce guarantees for data correctness, freshness, lineage, and recoverability, designing systems that handle rapid scale growth, partial failures, and evolving schemas without disrupting AI workloads or product experiences.

  • Build self-serve data platforms that let engineers, data scientists, and analysts safely discover data, define contracts, and create and operate their own pipelines with minimal friction.

  • Improve developer experience through better abstractions, opinionated paved paths, and standards for data modeling, testing, validation, and deployment, treating the data platform as a product used by many teams.

  • Drive architectural decisions across storage, compute, orchestration, and data APIs, partnering closely with product engineering and data science to align the data ecosystem with Perplexity’s roadmap.

  • Mentor engineers, review designs, and raise the technical bar for data infrastructure through thoughtful feedback, documentation, and hands-on collaboration.

Qualifications

  • 5+ years (Senior) or 8+ years (Staff) of software engineering experience.

  • Strong experience building production data infrastructure systems.

  • Hands-on experience with batch and/or streaming data processing at scale.

  • Deep familiarity with data orchestration systems (Airflow, Dagster, or similar).

  • Proficiency in Python and at least one additional backend language (Go, TypeScript, etc.).

  • Strong systems thinking around reliability, latency, cost, and complexity tradeoffs.

  • Experience supporting ML/AI workflows, training pipelines, or evaluation systems.

  • Familiarity with data quality, lineage, observability, and governance tooling.

  • Prior ownership of internal platforms used by many teams.

If you’re excited about this role, we encourage you to apply even if your experience doesn’t match every qualification listed above.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

San Francisco, California, United States

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Jun 4, 2026
Recorded sightings
220
Last seen by us
Oct 9, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.