Back to jobs

Lead Machine Learning Engineer, Evaluations

Mountain View

Pay
USD 170,000–190,000/year — pay source
Salary 170,000 – 190,000 USD per year Benefits include:
Read the full posting
Work setup
Unconfirmed
Employment
Full-time — employment source
Employment type Full-time
Read the full posting
Apply at Asapp-2

What you’ll work on

Full posting
  • Partner closely with Research, Product, and Platform teams to productize experiments into robust AI solutions

  • Mentor and support other engineers through design reviews, feedback, and knowledge sharing.

From the employer’s posting
Build the data infrastructure evaluation depends on: annotation and labeling pipelines, dataset versioning, data quality checks, and tooling that lets researchers and product teams run and interpret experiments without needing platform team help. Partner closely with Research, Product, and Platform teams to productize experiments into robust AI solutions Represent the eval platform to stakeholders outside the immediate team- set expectations on what "good" looks like for a model/agent release, and report on platform health and coverage.
Stay current with advancements in ML, NLP, voice, and LLM systems, and contribute actively to technical discussions across teams. Mentor and support other engineers through design reviews, feedback, and knowledge sharing. What you'll need

What you’ll bring

All qualifications

Core experience

  • Demonstrated experience leading the technical direction of a project or small team: setting architecture, driving design reviews, and being accountable for a system's long-term health (not just shipping features).
  • Experience designing data pipelines for ML evaluation- labeling/annotation workflows, dataset versioning and quality control, and reproducible benchmarking.
  • Experience building and evaluating agentic systems at scale.
  • Experience with voice/audio quality evaluations.
  • Familiarity with large-scale ML experimentation, benchmarking, or simulation frameworks.
  • Experience with conversational/customer-support AI domains (e.g., containment rate, conversation quality, goal completion).
Qualification wording
Demonstrated experience leading the technical direction of a project or small team: setting architecture, driving design reviews, and being accountable for a system's long-term health (not just shipping features).
Experience designing data pipelines for ML evaluation- labeling/annotation workflows, dataset versioning and quality control, and reproducible benchmarking.
Experience building and evaluating agentic systems at scale.
Experience with voice/audio quality evaluations.
Familiarity with large-scale ML experimentation, benchmarking, or simulation frameworks.
Experience with conversational/customer-support AI domains (e.g., containment rate, conversation quality, goal completion).
Education & alternatives
- Experience designing data pipelines for ML evaluation- labeling/annotation workflows, dataset versioning and quality control, and reproducible benchmarking. - A Bachelor’s Degree in CS or other related fields - Demonstrated technical mentorship of junior and mid-level engineers, driving adoption of best practices and architectural alignment for scalability and extensibility.

Tools in this posting

  • Python
  • AWS
  • Kafka
  • Kubernetes
  • Docker
Source — Tool mentions in context
- Demonstrated experience leading the technical direction of a project or small team: setting architecture, driving design reviews, and being accountable for a system's long-term health (not just shipping features). - Strong architectural skills, with proven experience designing complex, data-intensive software systems and production experience with Python, AWS, Kubernetes, and/or Docker. - Experience designing data pipelines for ML evaluation- labeling/annotation workflows, dataset versioning and quality control, and reproducible benchmarking.
- Knowledge of techniques for optimizing model architectures for faster inference. - Experience with AWS, CI/CD, Kafka, Athena Compensation package also includes a performance bonus on top of the listed salary range

Job description

View original posting ↗

At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we’re guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed, ownership, and a relentless focus on outcomes. ASAPP’s AI Engineering team is seeking an enterprising, talented and curious machine learning engineer.
 
The AI Engineering team is responsible for working closely with the research and modeling teams to create state-of-the-art NLP models for specific tasks, and deploy them in a production setting designed to serve our customers at scale. We are looking for a Machine Learning Engineer to help build and evaluate the core intelligence behind our agentic AI systems. This role will play a key part in designing and owning evaluation frameworks that ensure quality, safety, and performance across complex agentic systems.

We're looking for a Lead Machine Learning Engineer to own and grow the evaluation platform that measures quality, safety, and performance across ASAPP's agentic AI systems- the infrastructure that tells us, with confidence, whether a model or agent change is actually an improvement before it reaches customers.

This a hybrid role with 10-12 days of in-office presence per month to balance flexibility with collaboration.

What you'll do

  • Help develop the technical roadmap and architecture for the evaluation platform, from offline benchmarking to online/production monitoring of agentic and LLM-based systems.

  • Design eval methodologies appropriate to different stages of the pipeline: golden/regression test sets, human-in-the-loop review workflows, LLM-as-judge approaches, and automated metrics for task success, safety, and hallucinations.

  • Build the data infrastructure evaluation depends on: annotation and labeling pipelines, dataset versioning, data quality checks, and tooling that lets researchers and product teams run and interpret experiments without needing platform team help.

  • Partner closely with Research, Product, and Platform teams to productize experiments into robust AI solutions

  • Represent the eval platform to stakeholders outside the immediate team- set expectations on what "good" looks like for a model/agent release, and report on platform health and coverage.

  • Stay current with advancements in ML, NLP, voice, and LLM systems, and contribute actively to technical discussions across teams.

  • Mentor and support other engineers through design reviews, feedback, and knowledge sharing.

What you'll need

  • Deep, hands-on experience building and operating evaluation systems for modern ML/LLM/agentic systems- not just consuming existing eval tools.

  • Demonstrated experience leading the technical direction of a project or small team: setting architecture, driving design reviews, and being accountable for a system's long-term health (not just shipping features).

  • Strong architectural skills, with proven experience designing complex, data-intensive software systems and production experience with Python, AWS, Kubernetes, and/or Docker.

  • Experience designing data pipelines for ML evaluation- labeling/annotation workflows, dataset versioning and quality control, and reproducible benchmarking.

  • A Bachelor’s Degree in CS or other related fields

  • Demonstrated technical mentorship of junior and mid-level engineers, driving adoption of best practices and architectural alignment for scalability and extensibility.

  • Desire to learn, teach, and collaborate closely with cross-functional peers.

What we'd like to see

  • Experience building and evaluating agentic systems at scale. 

  • Experience with voice/audio quality evaluations.

  • Production experience with LLM-centric services (e.g., inference, orchestration, evaluation, monitoring)

  • Familiarity with large-scale ML experimentation, benchmarking, or simulation frameworks.

  • Experience with conversational/customer-support AI domains (e.g., containment rate, conversation quality, goal completion).

  • Knowledge of techniques for optimizing model architectures for faster inference.

  • Experience with AWS, CI/CD, Kafka, Athena

Compensation package also includes a performance bonus on top of the listed salary range
 
Separately, we also offer a compelling equity grant comprised of stock options

Salary

170,000 – 190,000 USD per year

Benefits include:
 
Competitive compensation with stock options
Comprehensive medical, vision, and dental insurance
401k matching
Fitness and wellness stipend
Mental well-being benefits
Professional learning and development stipend
Parental leave, including adoptive and foster parents
3 weeks paid time off (increases with tenure) along with sick leave, bereavement and jury duty
 
ASAPP is committed to creating a diverse environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, disability, age, or veteran status. If you have a disability and need assistance with our employment application process, please email us at careers@asapp.com to obtain assistance. #LI-SL1 #LI-Hybrid

Employment type

Full-time

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.

Complete your application on jobs.lever.co. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay
Salary 170,000 – 190,000 USD per year Benefits include:
Location & working pattern

Mountain View

We're looking for a Lead Machine Learning Engineer to own and grow the evaluation platform that measures quality, safety, and performance across ASAPP's agentic AI systems- the infrastructure that tells us, with confidence, whether a model or agent change is actually an improvement before it reaches customers. This a hybrid role with 10-12 days of in-office presence per month to balance flexibility with collaboration. What you'll do
More source context
3 weeks paid time off (increases with tenure) along with sick leave, bereavement and jury duty ASAPP is committed to creating a diverse environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, disability, age, or veteran status. If you have a disability and need assistance with our employment application process, please email us at careers@asapp.com to obtain assistance. #LI-SL1 #LI-Hybrid Employment type
Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 12, 2026
Recorded sightings
14
Last seen by us
Oct 8, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.