Back to jobs

Machine Learning Engineer: Evaluation

San Francisco, California, United States

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Bedrock-Robotics

What you’ll work on

Full posting
  • You'll work alongside construction veterans and world-class engineers to solve physical-world problems that simulations can't touch.

  • Design and maintain eval systems:

  • Implement infrastructure and classifiers - to self-annotate data and allow creation of datasets for a variety of training and evaluation use cases.

From the employer’s posting
We're not here debating the future of AI. We're deploying it in the real world. In just two years, we've raised $350M and achieved the first fully autonomous excavator deployments in construction. This is where algorithms meet steel-toed boots. You'll work alongside construction veterans and world-class engineers to solve physical-world problems that simulations can't touch. If you're ready to do meaningful work on hard problems, we'd love to have you join us. Machine Learning Engineer: Evaluation
What you’ll do: Design and maintain eval systems: Build pipelines for measuring system performance – across open loop and closed loop simulation, hardware in the loop systems, and field data from Bedrock Operator equipped machinery. Excite other teams to gain insights earlier in the development cycle through streamlined workflows.
Classify data sources for training and testing: Implement infrastructure and classifiers - to self-annotate data and allow creation of datasets for a variety of training and evaluation use cases. Leverage models to source rich annotations for massive datasets to accelerate model iteration. Predict system performance:

Tools in this posting

  • Python
Source — Tool mentions in context
- 2+ years of professional experience analyzing modern ML or robotics system performance on real-world problems - Proficiency in Python and a data warehouse query language and comfort with development on infrastructure within parallelized cloud-based frameworks - Strong statistical analysis skills (e.g. classification, model fit bias determination, hypothesis testing, and uncertainty quantification)

Job description

View original posting ↗

Join the team bringing advanced autonomy to the built world

At Bedrock, we're moving AI out of the lab and into the real world. Our team includes veterans who helped launch Waymo, scaled Segment to a $3.2B acquisition, and grew Uber Freight to $5B in revenue. Today, we're deploying autonomous systems on heavy construction equipment across the country, improving safety on job sites and accelerating schedules on critical infrastructure projects.

We're not here debating the future of AI. We're deploying it in the real world. In just two years, we've raised $350M and achieved the first fully autonomous excavator deployments in construction.

This is where algorithms meet steel-toed boots. You'll work alongside construction veterans and world-class engineers to solve physical-world problems that simulations can't touch. If you're ready to do meaningful work on hard problems, we'd love to have you join us.

Machine Learning Engineer:  Evaluation

Bedrock is bringing autonomy to the construction industry! We’re a group of veterans from the autonomous vehicle industry who are passionate about bringing the benefits of automation to areas in the construction industry currently underserved by the market.

We’re looking for a highly motivated engineer with experience evaluating complex ML systems deployed in the real world. Your Mission: Translate the infinite nuance of the built world into actionable, AI-native evaluations that accelerate Bedrock Operator adoption. 

The ideal candidate has hands-on experience in building evaluation systems and designing and executing statistical tests to gauge performance deltas between system iterations. More importantly, you’ve iterated on complex ML systems run in production environments, and you understand the complexities that come with it.

What you’ll do:

Design and maintain eval systems:

  • Build pipelines for measuring system performance – across open loop and closed loop simulation, hardware in the loop systems, and field data from Bedrock Operator equipped machinery. Excite other teams to gain insights earlier in the development cycle through streamlined workflows.

Develop metrics:

  • Connect product goals and system behavior - by bridging real-world specification to measurable indicators from logged data. Empower confident decision making from parameter tuning to program planning by slicing through the noise and delivering objective insights.

Classify data sources for training and testing:

  • Implement infrastructure and classifiers - to self-annotate data and allow creation of datasets for a variety of training and evaluation use cases. Leverage models to source rich annotations for massive datasets to accelerate model iteration.

Predict system performance:

  • Model metrics and interpret results - from various sources ranging from raw sensor data to key leading indicators. Determine whether new construction sites pose hidden challenges and drive business decisions about deployment readiness.

What we’re looking for:

  • Engineers who are currently Senior or Staff level with 5+ years of professional software engineering, data science, or research experience 

  • 2+ years of professional experience analyzing modern ML or robotics system performance on real-world problems 

  • Proficiency in Python and a data warehouse query language and comfort with development on infrastructure within parallelized cloud-based frameworks

  • Strong statistical analysis skills (e.g. classification, model fit bias determination, hypothesis testing, and uncertainty quantification)

  • Experience working with large datasets

***Bonus points: We’re especially interested in engineers who have applied statistical backgrounds to ML research or real-world robotics applications.

Bedrock Robotics is an Equal Opportunity Employer

We’re committed to building a diverse and inclusive workplace. We consider all qualified applicants for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, age, disability, veteran status, genetic information, or any other protected characteristic.

Reasonable Accommodations

We want our hiring process to be accessible to everyone. If you need an accommodation to participate in the application or interview process, please let your recruiter know so we can support you.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

San Francisco, California, United States

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Jun 2, 2026
Recorded sightings
81
Last seen by us
Oct 8, 2026
Employer says posted
Jan 31, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.