Back to jobs

Machine Learning Engineer - Model Evaluation & Experimentation

United States

Pay
$60–90/hour — pay source
This role is for one of our clients Compensation: $60-$90 per hour Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Machine Learning Engineers and Researchers to bring hands-on expertise in model development, experimentation, and evaluation to create rigorous benchmark tasks for advanced AI systems.
Read the full posting
Work setup
Remote stated — work setup source
In this role, you will design sophisticated, multi-step machine learning challenges inspired by real-world research workflows. From implementing experimental ideas and running training pipelines to analyzing model behavior and validating results, you will help establish high-quality evaluation benchmarks that reveal the strengths and limitations of frontier AI models. This is a fully remote, full-time engagement requiring approximately 35 hours per week. Requirements
Read the full posting
Employment
Unconfirmed
Apply at Weekday AI

What you’ll work on

Full posting
  • Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis.

  • Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria.

  • Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes.

From the employer’s posting
Key Responsibilities Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis. Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria.
Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis. Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria. Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes.
Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria. Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes. Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable.

What you’ll bring

All qualifications

Core experience

  • Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline.
  • Practical experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis.
  • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies.
  • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments.
  • Ability to commit approximately 35 hours per week on a consistent basis.

Preferred experience

  • Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior—is preferred.
  • Experience developing or evaluating large language models, foundation models, or generative AI systems.
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
  • Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation.
Qualification wording
Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline.
Practical experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis.
Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies.
Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments.
Ability to commit approximately 35 hours per week on a consistent basis.
Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior—is preferred.
Experience developing or evaluating large language models, foundation models, or generative AI systems.
Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation.
Education & alternatives
Required Qualifications - Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline. - Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role.

Tools in this posting

  • Python
Source — Tool mentions in context
- Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria. - Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes. - Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable.
- Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies. - Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments. - Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior—is preferred.

Job description

View original posting ↗

This role is for one of our clients

Compensation: $60-$90 per hour

Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Machine Learning Engineers and Researchers to bring hands-on expertise in model development, experimentation, and evaluation to create rigorous benchmark tasks for advanced AI systems.

In this role, you will design sophisticated, multi-step machine learning challenges inspired by real-world research workflows. From implementing experimental ideas and running training pipelines to analyzing model behavior and validating results, you will help establish high-quality evaluation benchmarks that reveal the strengths and limitations of frontier AI models.

This is a fully remote, full-time engagement requiring approximately 35 hours per week.

Requirements

Key Responsibilities

  • Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis.
  • Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria.
  • Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes.
  • Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable.
  • Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions.
  • Collaborate with AI researchers and fellow subject matter experts to continuously improve benchmark quality, technical rigor, and evaluation consistency.

Required Qualifications

  • Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline.
  • Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role.
  • Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows.
  • Practical experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis.
  • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies.
  • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments.
  • Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior—is preferred.
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
  • Excellent analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended technical problems independently.
  • Strong written communication skills for documenting experimental methodologies and technical findings.
  • Ability to commit approximately 35 hours per week on a consistent basis.

Preferred Qualifications

  • Experience developing or evaluating large language models, foundation models, or generative AI systems.
  • Background in reinforcement learning, deep learning, distributed training, or model optimization.
  • Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation.
  • Experience contributing to research publications, open-source machine learning projects, or advanced AI systems.

Why Join

  • Help shape how next-generation AI systems are evaluated through rigorous machine learning experimentation.
  • Collaborate with leading AI researchers developing frontier evaluation benchmarks.
  • Apply your expertise to improve AI reasoning, model quality, and experimental reliability.
  • Contribute directly to benchmark development that advances the capabilities of state-of-the-art AI systems.
  • Enjoy the flexibility of a fully remote engagement while working on impactful AI research initiatives.

Equal Opportunity

We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.

Contract & Engagement Details

  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of approximately 35 hours per week.
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
  • Work does not require access to confidential or proprietary information from any current or former employer.
  • Payments are issued weekly based on approved work completed.
  • At this time, we are unable to support H1-B or STEM OPT candidates.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.

Complete your application on apply.workable.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay
This role is for one of our clients Compensation: $60-$90 per hour Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Machine Learning Engineers and Researchers to bring hands-on expertise in model development, experimentation, and evaluation to create rigorous benchmark tasks for advanced AI systems.
Location & working pattern

United States

In this role, you will design sophisticated, multi-step machine learning challenges inspired by real-world research workflows. From implementing experimental ideas and running training pipelines to analyzing model behavior and validating results, you will help establish high-quality evaluation benchmarks that reveal the strengths and limitations of frontier AI models. This is a fully remote, full-time engagement requiring approximately 35 hours per week. Requirements
More source context
- Contribute directly to benchmark development that advances the capabilities of state-of-the-art AI systems. - Enjoy the flexibility of a fully remote engagement while working on impactful AI research initiatives. Equal Opportunity

More relevant text appears in the full description.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Aug 15, 2026
Recorded sightings
173
Last seen by us
Oct 9, 2026
Employer says posted
Jul 30, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.