Back to jobs

Staff ML Engineer, Frontier AI

San Francisco, California, United States

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Ambiencehealthcare

What you’ll work on

Full posting
  • You’ll work closely with clinicians, product managers, researchers, and engineers to translate cutting-edge research into reliable, production-grade AI systems.

  • Drive Production Model Improvement: Identify high-impact failure modes across products and lead improvements across prompting, retrieval, context, routing, data, and post-training.

  • Design experiments that determine the right intervention and prove that improvements hold in production.

From the employer’s posting
The Role: As a Staff Machine Learning Engineer at Ambience, you will help set the technical direction for the AI systems that power our clinical products. You’ll identify the highest-impact opportunities to improve model behavior, evaluation, post-training, and agentic systems, and lead the design and execution of cross-cutting initiatives. This is a highly hands-on role with broad technical influence. You’ll work closely with clinicians, product managers, researchers, and engineers to translate cutting-edge research into reliable, production-grade AI systems. Our engineering roles are hybrid — working onsite at our San Francisco office three days per week. What You’ll Own:
Define the AI Quality and Evaluation Strategy: Establish how we measure production AI quality across LLM and agentic systems, including automated graders, regression testing, human evaluation, failure taxonomies, and offline-to-online validation. Drive Production Model Improvement: Identify high-impact failure modes across products and lead improvements across prompting, retrieval, context, routing, data, and post-training. Design experiments that determine the right intervention and prove that improvements hold in production. Advance Post-Training Capabilities: Shape how Ambience uses techniques such as supervised fine-tuning, preference optimization, reinforcement learning, distillation, and synthetic data to improve model behavior.

What you’ll bring

All qualifications

Core experience

  • Deep expertise with modern LLMs, transformers, and complex production AI systems.
  • Experience with realtime voice, conversational AI, or multimodal systems.
  • 401(k) with a company match of up to 3% of base salary
  • Experience interviewing or hiring ML engineers.
Qualification wording
Expert in Production AI Systems 7+ years in production ML, research engineering, or applied AI. Deep expertise with modern LLMs, transformers, and complex production AI systems. Have led or materially shaped consequential AI systems with real production impact. Able to reason across models, data, orchestration, evaluation, serving, and product behavior.
Experience with realtime voice, conversational AI, or multimodal systems.
401(k) with a company match of up to 3% of base salary
Experience interviewing or hiring ML engineers.

Tools in this posting

  • Python
  • PyTorch
Source — Tool mentions in context
- Post-Training and Model Improvement Expertise Have materially improved model behavior through post-training or model adaptation techniques such as supervised fine-tuning, preference optimization, reinforcement learning, distillation, synthetic data, or related approaches. Understand how data quality, objective design, evaluation, and post-training choices affect downstream model behavior. Know when post-training is the right intervention versus prompting, retrieval, context optimization, routing, or changes to the surrounding system. Have taken model improvements from experimentation through rigorous evaluation and production deployment. - Production-Grade Software Engineer Proficient in Python and modern ML frameworks; PyTorch preferred. Comfortable with modern MLOps, CI/CD, observability, and containerized deployments. Remains deeply hands-on: writes code, inspects traces, analyzes failures, and prototypes solutions. - Data-Centric AI Developer Skilled at building large, high-quality datasets and feedback loops. Experienced using production failures, user feedback, and active learning to improve model and system quality. Able to design data and evaluation infrastructure that creates leverage across multiple teams.

About Ambiencehealthcare

Here at Ambience, we never set out to be just another scribe.

In the employer’s words · Read in context

Job description

View original posting ↗

About Us:

Here at Ambience, we never set out to be just another scribe. We’re building the AI intelligence platform that restores humanity to healthcare and drives meaningful ROI for health systems across the country.

Our technology helps providers focus on delivering great care by removing the administrative burden that pulls them away from patients and away from their most impactful work. Ambience delivers real-time coding-aware documentation and clinical workflow support across ambulatory, emergency and inpatient settings at the top health systems in North America.

Our teams operate relentlessly with extreme ownership to build the best solutions for our health system partners. We value candor, positivity and deep thought — and we expect a lot from each other because we know the problems we’re solving truly matter.

Ambience was ranked #1 for Improving the Clinician Experience in the KLAS Research Emerging Solutions Top 20 Report, recognized by Fast Company as one of the Next Big Things in Tech, named one of the best AI companies in healthcare by Inc., and selected as a LinkedIn Top Startup in 2024 and 2025. We’re backed by Oak HC/FT, Andreessen Horowitz (a16z), OpenAI Startup Fund, and Kleiner Perkins — and we’re just getting started.

The Role:

As a Staff Machine Learning Engineer at Ambience, you will help set the technical direction for the AI systems that power our clinical products. You’ll identify the highest-impact opportunities to improve model behavior, evaluation, post-training, and agentic systems, and lead the design and execution of cross-cutting initiatives.
This is a highly hands-on role with broad technical influence. You’ll work closely with clinicians, product managers, researchers, and engineers to translate cutting-edge research into reliable, production-grade AI systems.
Our engineering roles are hybrid — working onsite at our San Francisco office three days per week.

What You’ll Own:

  • Define the AI Quality and Evaluation Strategy: Establish how we measure production AI quality across LLM and agentic systems, including automated graders, regression testing, human evaluation, failure taxonomies, and offline-to-online validation.

  • Drive Production Model Improvement: Identify high-impact failure modes across products and lead improvements across prompting, retrieval, context, routing, data, and post-training. Design experiments that determine the right intervention and prove that improvements hold in production.

  • Advance Post-Training Capabilities: Shape how Ambience uses techniques such as supervised fine-tuning, preference optimization, reinforcement learning, distillation, and synthetic data to improve model behavior.

  • Architect Agentic AI Systems: Set technical direction for production systems involving tool use, retrieval, context and state management, routing, orchestration, tracing, and failure recovery.

  • Build Self-Improving AI Loops: Design the systems that turn production failures and user feedback into better datasets, evaluations, and model behavior, creating scalable improvement flywheels across products.

  • Shape the AI Roadmap: Co-develop the 12-month technical roadmap, balancing near-term product wins with longer-term research and platform investments.

  • Raise the Technical Bar: Lead technical reviews, mentor engineers, establish best practices, and create reusable systems and frameworks that improve the effectiveness of the broader AI team.

  • Stay at the Cutting Edge: Distill insights from recent research in LLMs, agents, post-training, NLP, speech, and multimodal AI, and drive experiments that keep Ambience at the forefront of clinical AI.

Who You Are:

  • Expert in Production AI Systems
    7+ years in production ML, research engineering, or applied AI.
    Deep expertise with modern LLMs, transformers, and complex production AI systems.
    Have led or materially shaped consequential AI systems with real production impact.
    Able to reason across models, data, orchestration, evaluation, serving, and product behavior.

  • Deep Evaluation Rigor
    Significant experience designing evaluations for LLMs, agents, or other complex AI systems.
    Can turn ambiguous product-quality problems into robust metrics, datasets, and experiments.
    Experienced with grader bias, leakage, contamination, misleading aggregate metrics, regression detection, and offline-online mismatch.
    Knows how to establish whether a model or system improvement actually translates into better user outcomes.

  • Agentic Systems Expertise
    Have built or architected production systems involving multiple models, tools, retrieval, context and state management, routing, orchestration, tracing, and failure recovery.
    Understand how to evaluate and debug complex agent behavior across model, tool, and system boundaries.
    Comfortable making architectural decisions that span multiple products or teams.

  • Post-Training and Model Improvement Expertise
    Have materially improved model behavior through post-training or model adaptation techniques such as supervised fine-tuning, preference optimization, reinforcement learning, distillation, synthetic data, or related approaches.
    Understand how data quality, objective design, evaluation, and post-training choices affect downstream model behavior.
    Know when post-training is the right intervention versus prompting, retrieval, context optimization, routing, or changes to the surrounding system.
    Have taken model improvements from experimentation through rigorous evaluation and production deployment.

  • Production-Grade Software Engineer
    Proficient in Python and modern ML frameworks; PyTorch preferred.
    Comfortable with modern MLOps, CI/CD, observability, and containerized deployments.
    Remains deeply hands-on: writes code, inspects traces, analyzes failures, and prototypes solutions.

  • Data-Centric AI Developer
    Skilled at building large, high-quality datasets and feedback loops.
    Experienced using production failures, user feedback, and active learning to improve model and system quality.
    Able to design data and evaluation infrastructure that creates leverage across multiple teams.

  • Effective Technical Leader
    Able to work closely with clinicians, product managers, researchers, and engineers.
    Strong communicator who can simplify complex AI concepts and build alignment around technical decisions.
    Comfortable operating in ambiguity, identifying the right problems to solve, and driving execution across multiple workstreams.

  • Mission-Aligned
    Passion for healthcare or other mission-driven industries.
    Thrives in a fast-paced, early-stage environment.
    Takes broad ownership of technical outcomes and raises the effectiveness of the people around them.

    Nice-to-Haves

  • Experience with realtime voice, conversational AI, or multimodal systems.

  • Prior work in healthcare, clinical AI, or other regulated, high-stakes industries.

  • Experience interviewing or hiring ML engineers.

  • Open-source contributions to ML, agent, evaluation, or post-training tooling.

Why Here:

Our products power specialty-specific note generation, chart-aware diagnosis prediction, and real-time clinical decision support in real clinical settings. The work is deeply technical, but the goal is simple: help clinicians do great work with less friction.
To keep improving, we can't wait around for bigger models. The fastest path is building intelligence that gets better with every encounter — learning from how clinicians actually use the products, what they change, and what outcomes follow.
You will own the hardest model quality problems across our clinical AI suite: coding models that navigate a proprietary million-term ontology with multi-objective precision, a scribe that learns from edit signals without introducing regressions, long-context chart understanding that stays faithful under real clinical complexity, and population-level reasoning that surfaces patterns across patients in a way that's auditable and actionable.
This is not applied ML on clean benchmarks — it's research-grade model work with production stakes, where your improvements directly shape what clinicians experience every day.

Life at Ambience

Working at Ambience means opting into a high-ownership, high-trust environment built for people who want to grow fast, operate decisively and focus on work that matters. This could be the right place for you if you want to

  • Work on mission-critical AI technology that directly improves clinicians’ day-to-day lives and health system financial health across some of the most complex, high-stakes workflows in the world.

  • Join a “dream team” culture where we hire exceptional people, expect exceptional outcomes and invest deeply in feedback and continuous growth. We operate as a championship team, and that means being ok with hard, uncomfortable, ambiguous problems that lead to real greatness.

  • Operate with real ownership and accountability in an environment where there are no bystanders: If something is broken, we fix it! You will have meaningful autonomy and be expected to drive work to completion.

To help you do your best work, we pair these expectations with benefits intentionally designed to help you feel supported and safe at Ambience and beyond. Some of our key benefits include

  • Comprehensive medical, dental, and vision coverage for you and your dependents

  • 401(k) with a company match of up to 3% of base salary

  • A remote-friendly culture (with a San Francisco HQ) and full equipment provisioning to ensure you can work effectively from wherever you’re based.

  • Parental leave to support your family needs

  • Annual company-wide off-sites, team off-sites and regular team lunches and all-hands gatherings, with travel, lodging and meals covered

  • Flexible time off with no annual cap, company-wide holidays and an annual holiday shutdown from December 24–January 1 designed to support real rest and long-term sustainability.

Ambience Healthcare is an equal opportunity employer and is committed to building a diverse and inclusive workplace. We do not discriminate on the basis of race, color, religion, sex, gender identity or expression, sexual orientation, national origin, age, disability, veteran status, genetic information, or any other legally protected status. We encourage applicants from all backgrounds to apply.

Ambience is committed to supporting every candidate’s ability to fully participate in our hiring process. If you need any accommodations during your application or interviews, please reach out to our Recruiting team at accommodations@ambiencehealthcare.com. We’ll handle your request confidentially and work with you to ensure an accessible and equitable experience for all candidates.


Ambience Healthcare has become aware of scams targeting jobseekers with fake jobs and even interviewing people. Our emails will always come from @ambiencehealthcare.com. We would never our ask candidates to download apps or make any form of payment(s). If you are contacted through WhatsApp, Telegram, similar but fake email domains, or asked to make a payment, these contacts are not legitimate. Report the issue immediately to LinkedIn and the FBI.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

San Francisco, California, United States

The Role: As a Staff Machine Learning Engineer at Ambience, you will help set the technical direction for the AI systems that power our clinical products. You’ll identify the highest-impact opportunities to improve model behavior, evaluation, post-training, and agentic systems, and lead the design and execution of cross-cutting initiatives. This is a highly hands-on role with broad technical influence. You’ll work closely with clinicians, product managers, researchers, and engineers to translate cutting-edge research into reliable, production-grade AI systems. Our engineering roles are hybrid — working onsite at our San Francisco office three days per week. What You’ll Own:
More source context
- 401(k) with a company match of up to 3% of base salary - A remote-friendly culture (with a San Francisco HQ) and full equipment provisioning to ensure you can work effectively from wherever you’re based. - Parental leave to support your family needs
Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Jun 2, 2026
Recorded sightings
61
Last seen by us
Oct 10, 2026
Employer says posted
Mar 17, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.