Research Scientist, Agentic Data & Benchmarking
Sunnyvale, CA
- Pay
USD 150,000–450,000/year — pay source
Salary 150,000 – 450,000 USD per year We encourage you to apply even if you don't meet every qualification listed. Strong candidates rarely match every line, and we'd rather hear from you than have you rule yourself out.
Read the full posting- Work setup
- Unconfirmed
- Employment
Full-time — employment source
Employment type Full-time
Read the full posting
What you’ll work on
Full postingThe Agents team trains advanced agentic language models that use reasoning and tool use to complete real tasks on a computer.
You'll own the agentic data pipeline end-to-end — sourcing and generating high-quality trajectories, tool-use data, and RL environments — and the evaluation suite that tells us, rigorously and reproducibly, what our agents can actually do.
Build and harden evaluation harnesses so benchmarks run reliably at scale against training checkpoints, with clear signal on regressions and model health.
Design and scale RL environments and reward signals, and measure their impact on model performance.
Partner with research and product teams to translate capability goals into measurable data and evaluation artifacts.
From the employer’s posting
The Agents team trains advanced agentic language models that use reasoning and tool use to complete real tasks on a computer. This is a specialist role at the center of the loop that drives those models: the data we train on and the benchmarks we measure against.
You'll own the agentic data pipeline end-to-end — sourcing and generating high-quality trajectories, tool-use data, and RL environments — and the evaluation suite that tells us, rigorously and reproducibly, what our agents can actually do. These two halves are inseparable: benchmarks expose where models fail, and targeted data closes the gap. The agents are only as good as the data they learn from and the evals that keep us honest, and this role owns both.
Design and run evaluations of agentic capabilities — multi-step reasoning, tool use, long-horizon planning, computer use, and safety properties — turning ambiguous notions of "intelligence" into defensible, reproducible metrics. Build and harden evaluation harnesses so benchmarks run reliably at scale against training checkpoints, with clear signal on regressions and model health. Run experiments characterizing how prompting, sampling, scaffolding, and environment design affect agentic performance on internal and public benchmarks.
Source, generate, and curate high-quality agentic training data: trajectories, tool-use traces, and task datasets for new capabilities. Design and scale RL environments and reward signals, and measure their impact on model performance. Manage technical relationships with external data vendors and domain experts, evaluating data quality and iterating quickly on feedback.
Contribute to technical reports, research publications, and open-source benchmarks and tooling. Partner with research and product teams to translate capability goals into measurable data and evaluation artifacts. Qualifications
What you’ll bring
All qualificationsCore experience
- 2+ years of experience with a clear emphasis on evaluations and/or training-data curation for ML systems (related areas: LLM training/fine-tuning, RL, or distributed ML systems).
- Strong Python and PyTorch development experience.
- Demonstrated experience designing and deep-diving into evaluations, or curating and generating training datasets — ideally both.
- Hands-on experience using LLM agents in your personal or professional work.
Preferred experience
- Experience with reinforcement learning, reward design, or RL environment construction for LLMs.
- Experience with large-scale dataset sourcing, curation, and processing, including working with external vendors or domain experts.
- Strong knowledge of the literature on agent evaluation, RL, LLM reasoning, and tool use.
- Experience building or operating data pipelines and evaluation infrastructure reliable at scale (e.g., PyTorch, Ray).
Qualification wording
2+ years of experience with a clear emphasis on evaluations and/or training-data curation for ML systems (related areas: LLM training/fine-tuning, RL, or distributed ML systems).
Strong Python and PyTorch development experience.
Demonstrated experience designing and deep-diving into evaluations, or curating and generating training datasets — ideally both.
Hands-on experience using LLM agents in your personal or professional work.
Experience with reinforcement learning, reward design, or RL environment construction for LLMs.
Experience with large-scale dataset sourcing, curation, and processing, including working with external vendors or domain experts.
Strong knowledge of the literature on agent evaluation, RL, LLM reasoning, and tool use.
Experience building or operating data pipelines and evaluation infrastructure reliable at scale (e.g., PyTorch, Ray).
Education & alternatives
Academic qualifications - BS, MS, or PhD (or equivalent experience) in Computer Science, Machine Learning, or a related field. Minimum qualifications
Tools in this posting
- Python
- PyTorch
Source — Tool mentions in context
- 2+ years of experience with a clear emphasis on evaluations and/or training-data curation for ML systems (related areas: LLM training/fine-tuning, RL, or distributed ML systems). - Strong Python and PyTorch development experience. - Demonstrated experience designing and deep-diving into evaluations, or curating and generating training datasets — ideally both.
- Strong knowledge of the literature on agent evaluation, RL, LLM reasoning, and tool use. - Experience building or operating data pipelines and evaluation infrastructure reliable at scale (e.g., PyTorch, Ray). - Experience evaluating or generating data for software-engineering or computer-use agents.
About Ifm-Us
The Institute of Foundation Models (IFM) is a dedicated research lab for building, understanding, using, and risk-managing foundation models.
In the employer’s words · Read in context
Job description
About the Institute of Foundation Models
The Institute of Foundation Models (IFM) is a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
As part of our team, you'll work at the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You'll help build groundbreaking AI systems with the potential to reshape entire industries, and contribute to establishing MBZUAI as a global hub for high-performance computing and deep learning.
About the role
The Agents team trains advanced agentic language models that use reasoning and tool use to complete real tasks on a computer. This is a specialist role at the center of the loop that drives those models: the data we train on and the benchmarks we measure against.
You'll own the agentic data pipeline end-to-end — sourcing and generating high-quality trajectories, tool-use data, and RL environments — and the evaluation suite that tells us, rigorously and reproducibly, what our agents can actually do. These two halves are inseparable: benchmarks expose where models fail, and targeted data closes the gap. The agents are only as good as the data they learn from and the evals that keep us honest, and this role owns both.
This is a research scientist position for someone who wants depth in data and measurement rather than breadth across the whole stack. You should be the kind of person who reads through datasets line by line, distrusts a metric until it's been validated, and gets satisfaction from making an eval suite that nobody questions.
Key responsibilities
-
Design and run evaluations of agentic capabilities — multi-step reasoning, tool use, long-horizon planning, computer use, and safety properties — turning ambiguous notions of "intelligence" into defensible, reproducible metrics.
-
Build and harden evaluation harnesses so benchmarks run reliably at scale against training checkpoints, with clear signal on regressions and model health.
-
Run experiments characterizing how prompting, sampling, scaffolding, and environment design affect agentic performance on internal and public benchmarks.
-
Diagnose anomalous eval results mid-training-run — determine whether the cause is the model, the data, the harness, or the infrastructure — and communicate the answer clearly.
-
Source, generate, and curate high-quality agentic training data: trajectories, tool-use traces, and task datasets for new capabilities.
-
Design and scale RL environments and reward signals, and measure their impact on model performance.
-
Manage technical relationships with external data vendors and domain experts, evaluating data quality and iterating quickly on feedback.
-
Develop QA frameworks that catch reward hacking, label noise, and contamination, keeping data and benchmark quality high.
-
Contribute to technical reports, research publications, and open-source benchmarks and tooling.
-
Partner with research and product teams to translate capability goals into measurable data and evaluation artifacts.
Benchmarking & evaluation
Agentic data
Across both
Qualifications
-
BS, MS, or PhD (or equivalent experience) in Computer Science, Machine Learning, or a related field.
-
2+ years of experience with a clear emphasis on evaluations and/or training-data curation for ML systems (related areas: LLM training/fine-tuning, RL, or distributed ML systems).
-
Strong Python and PyTorch development experience.
-
Demonstrated experience designing and deep-diving into evaluations, or curating and generating training datasets — ideally both.
-
Hands-on experience using LLM agents in your personal or professional work.
-
A habit of reading through raw data and trajectories to understand them and spot issues, and an instinct to distrust a metric until it's validated.
-
Experience with reinforcement learning, reward design, or RL environment construction for LLMs.
-
Background in statistics and experimental design — a feel for signal-to-noise, statistical power, and contamination in evaluations.
-
Experience with large-scale dataset sourcing, curation, and processing, including working with external vendors or domain experts.
-
Strong knowledge of the literature on agent evaluation, RL, LLM reasoning, and tool use.
-
Experience building or operating data pipelines and evaluation infrastructure reliable at scale (e.g., PyTorch, Ray).
-
Experience evaluating or generating data for software-engineering or computer-use agents.
-
Contributions to published research, public benchmarks, and/or open-source ML software.
Academic qualifications
Minimum qualifications
Preferred qualifications
Representative projects
-
Stand up a new agentic benchmark from scratch — define the task, build the dataset and scoring, validate against known signals, and ship a view that makes the result legible to researchers and leadership.
-
Build an RL environment for a new high-value capability: design the reward, generate and QA the trajectory data, and measure the lift on model performance.
-
Diagnose a mid-training regression: an eval suite returns anomalous numbers and you determine whether it's the model, the harness, the data, or the infrastructure.
-
Partner with an external data vendor or domain expert to source high-quality trajectories, then build the QA framework that keeps reward hacking and contamination out.
-
Take a flaky distributed eval pipeline and make it reliable — better retries, better observability, faster feedback to researchers.
Salary
150,000 – 450,000 USD per year
Employment type
Full-time
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
Complete your application on jobs.lever.co. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
Salary 150,000 – 450,000 USD per year We encourage you to apply even if you don't meet every qualification listed. Strong candidates rarely match every line, and we'd rather hear from you than have you rule yourself out.
- Location & working pattern
Sunnyvale, CA
Working pattern and location restrictions need checking in the full posting.
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Jun 9, 2026
- Recorded sightings
- 24
- Last seen by us
- Oct 6, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.