Machine Learning Engineer: Evaluation
San Francisco, California, United States
- Pay
- Salary not listed in the saved posting
- Work setup
- Unconfirmed
- Employment
- Unconfirmed
What you’ll work on
Full postingYou'll work alongside construction veterans and world-class engineers to solve physical-world problems that simulations can't touch.
Design and maintain eval systems:
Implement infrastructure and classifiers - to self-annotate data and allow creation of datasets for a variety of training and evaluation use cases.
From the employer’s posting
We're not here debating the future of AI. We're deploying it in the real world. In just two years, we've raised $350M and achieved the first fully autonomous excavator deployments in construction. This is where algorithms meet steel-toed boots. You'll work alongside construction veterans and world-class engineers to solve physical-world problems that simulations can't touch. If you're ready to do meaningful work on hard problems, we'd love to have you join us. Machine Learning Engineer: Evaluation
What you’ll do: Design and maintain eval systems: Build pipelines for measuring system performance – across open loop and closed loop simulation, hardware in the loop systems, and field data from Bedrock Operator equipped machinery. Excite other teams to gain insights earlier in the development cycle through streamlined workflows.
Classify data sources for training and testing: Implement infrastructure and classifiers - to self-annotate data and allow creation of datasets for a variety of training and evaluation use cases. Leverage models to source rich annotations for massive datasets to accelerate model iteration. Predict system performance:
Tools in this posting
- Python
Source — Tool mentions in context
- 2+ years of professional experience analyzing modern ML or robotics system performance on real-world problems - Proficiency in Python and a data warehouse query language and comfort with development on infrastructure within parallelized cloud-based frameworks - Strong statistical analysis skills (e.g. classification, model fit bias determination, hypothesis testing, and uncertainty quantification)
Job description
Join the team bringing advanced autonomy to the built world
At Bedrock, we're moving AI out of the lab and into the real world. Our team includes veterans who helped launch Waymo, scaled Segment to a $3.2B acquisition, and grew Uber Freight to $5B in revenue. Today, we're deploying autonomous systems on heavy construction equipment across the country, improving safety on job sites and accelerating schedules on critical infrastructure projects.
We're not here debating the future of AI. We're deploying it in the real world. In just two years, we've raised $350M and achieved the first fully autonomous excavator deployments in construction.
This is where algorithms meet steel-toed boots. You'll work alongside construction veterans and world-class engineers to solve physical-world problems that simulations can't touch. If you're ready to do meaningful work on hard problems, we'd love to have you join us.
Machine Learning Engineer: Evaluation
Bedrock is bringing autonomy to the construction industry! We’re a group of veterans from the autonomous vehicle industry who are passionate about bringing the benefits of automation to areas in the construction industry currently underserved by the market.
We’re looking for a highly motivated engineer with experience evaluating complex ML systems deployed in the real world. Your Mission: Translate the infinite nuance of the built world into actionable, AI-native evaluations that accelerate Bedrock Operator adoption.
The ideal candidate has hands-on experience in building evaluation systems and designing and executing statistical tests to gauge performance deltas between system iterations. More importantly, you’ve iterated on complex ML systems run in production environments, and you understand the complexities that come with it.
What you’ll do:
Design and maintain eval systems:
Build pipelines for measuring system performance – across open loop and closed loop simulation, hardware in the loop systems, and field data from Bedrock Operator equipped machinery. Excite other teams to gain insights earlier in the development cycle through streamlined workflows.
Develop metrics:
Connect product goals and system behavior - by bridging real-world specification to measurable indicators from logged data. Empower confident decision making from parameter tuning to program planning by slicing through the noise and delivering objective insights.
Classify data sources for training and testing:
Implement infrastructure and classifiers - to self-annotate data and allow creation of datasets for a variety of training and evaluation use cases. Leverage models to source rich annotations for massive datasets to accelerate model iteration.
Predict system performance:
Model metrics and interpret results - from various sources ranging from raw sensor data to key leading indicators. Determine whether new construction sites pose hidden challenges and drive business decisions about deployment readiness.
What we’re looking for:
Engineers who are currently Senior or Staff level with 5+ years of professional software engineering, data science, or research experience
2+ years of professional experience analyzing modern ML or robotics system performance on real-world problems
Proficiency in Python and a data warehouse query language and comfort with development on infrastructure within parallelized cloud-based frameworks
Strong statistical analysis skills (e.g. classification, model fit bias determination, hypothesis testing, and uncertainty quantification)
Experience working with large datasets
***Bonus points: We’re especially interested in engineers who have applied statistical backgrounds to ML research or real-world robotics applications.
Bedrock Robotics is an Equal Opportunity Employer
We’re committed to building a diverse and inclusive workplace. We consider all qualified applicants for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, age, disability, veteran status, genetic information, or any other protected characteristic.
Reasonable Accommodations
We want our hiring process to be accessible to everyone. If you need an accommodation to participate in the application or interview process, please let your recruiter know so we can support you.
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
- Ask the employer about the salary range before committing time to the process.
Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
No pay amount identified in the saved description.
- Location & working pattern
San Francisco, California, United States
Working pattern and location restrictions need checking in the full posting.
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Jun 2, 2026
- Recorded sightings
- 81
- Last seen by us
- Oct 8, 2026
- Employer says posted
- Jan 31, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.