Data Science & Quantitative Analysis Expert
United States
- Pay
$60–90/hour — pay source
This role is for one of our clients Compensation: $60-$90 per hour Join a pioneering AI initiative focused on developing the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Data Scientists and Quantitative Analysts to bring real-world analytical rigor to AI evaluation by designing sophisticated benchmark tasks based on practical data science workflows.
Read the full posting- Work setup
Remote stated — work setup source
In this role, you will create complex, multi-step analytical challenges that mirror real research and business scenarios—from cleaning datasets and comparing statistical methods to interpreting results and presenting actionable insights. Working closely with AI researchers, you'll help identify where advanced AI models succeed, where they fail, and how evaluation benchmarks can better measure analytical reasoning. This is a fully remote, full-time engagement requiring approximately 35 hours per week. Requirements
Read the full posting- Employment
- Unconfirmed
What you’ll work on
Full postingDevelop reproducible reference analyses using Jupyter Notebooks or Google Colab, documenting methodologies and findings with clarity.
Create benchmark tasks that require objective comparisons between analytical techniques, supported by statistical validation and evidence-based recommendations.
Evaluate AI-generated analyses for correctness, statistical validity, reasoning quality, and interpretation accuracy.
From the employer’s posting
Design realistic data analysis challenges inspired by real-world research and analytical workflows, including data preparation, statistical modeling, hypothesis testing, and comparative analysis. Develop reproducible reference analyses using Jupyter Notebooks or Google Colab, documenting methodologies and findings with clarity. Create benchmark tasks that require objective comparisons between analytical techniques, supported by statistical validation and evidence-based recommendations.
Develop reproducible reference analyses using Jupyter Notebooks or Google Colab, documenting methodologies and findings with clarity. Create benchmark tasks that require objective comparisons between analytical techniques, supported by statistical validation and evidence-based recommendations. Evaluate AI-generated analyses for correctness, statistical validity, reasoning quality, and interpretation accuracy.
Create benchmark tasks that require objective comparisons between analytical techniques, supported by statistical validation and evidence-based recommendations. Evaluate AI-generated analyses for correctness, statistical validity, reasoning quality, and interpretation accuracy. Identify analytical errors, flawed assumptions, and reasoning gaps that experienced data professionals would immediately recognize.
What you’ll bring
All qualificationsCore experience
- Master's degree, PhD, or equivalent practical experience in Data Science, Statistics, Mathematics, Economics, Operations Research, or another quantitative STEM discipline.
- Proficiency with Jupyter Notebooks or Google Colab for building reproducible analytical workflows.
- Ability to commit approximately 35 hours per week on a consistent basis.
Preferred experience
- Experience with AI evaluation, benchmark development, AI model assessment, or task authoring is preferred.
- Experience designing reproducible research workflows or analytical evaluation frameworks.
- Familiarity with machine learning, large language models, or AI-assisted data analysis.
- Experience reviewing analytical work, mentoring analysts, or contributing to research publications.
Qualification wording
Master's degree, PhD, or equivalent practical experience in Data Science, Statistics, Mathematics, Economics, Operations Research, or another quantitative STEM discipline.
Proficiency with Jupyter Notebooks or Google Colab for building reproducible analytical workflows.
Ability to commit approximately 35 hours per week on a consistent basis.
Experience with AI evaluation, benchmark development, AI model assessment, or task authoring is preferred.
Experience designing reproducible research workflows or analytical evaluation frameworks.
Familiarity with machine learning, large language models, or AI-assisted data analysis.
Experience reviewing analytical work, mentoring analysts, or contributing to research publications.
Education & alternatives
Required Qualifications - Master's degree, PhD, or equivalent practical experience in Data Science, Statistics, Mathematics, Economics, Operations Research, or another quantitative STEM discipline. - Minimum 1 year of professional experience in research, research engineering, quantitative analysis, data science, or another data-intensive analytical role.
Tools in this posting
- Python
- NumPy
- pandas
- Scipy
- scikit-learn
Source — Tool mentions in context
- Proficiency with Jupyter Notebooks or Google Colab for building reproducible analytical workflows. - Strong programming skills in Python, including experience with libraries such as pandas, NumPy, SciPy, scikit-learn, or similar analytical frameworks. - Working knowledge of Git and collaborative software development practices.
Job description
This role is for one of our clients
Compensation: $60-$90 per hour
Join a pioneering AI initiative focused on developing the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Data Scientists and Quantitative Analysts to bring real-world analytical rigor to AI evaluation by designing sophisticated benchmark tasks based on practical data science workflows.
In this role, you will create complex, multi-step analytical challenges that mirror real research and business scenarios—from cleaning datasets and comparing statistical methods to interpreting results and presenting actionable insights. Working closely with AI researchers, you'll help identify where advanced AI models succeed, where they fail, and how evaluation benchmarks can better measure analytical reasoning.
This is a fully remote, full-time engagement requiring approximately 35 hours per week.
Requirements
Key Responsibilities
- Design realistic data analysis challenges inspired by real-world research and analytical workflows, including data preparation, statistical modeling, hypothesis testing, and comparative analysis.
- Develop reproducible reference analyses using Jupyter Notebooks or Google Colab, documenting methodologies and findings with clarity.
- Create benchmark tasks that require objective comparisons between analytical techniques, supported by statistical validation and evidence-based recommendations.
- Evaluate AI-generated analyses for correctness, statistical validity, reasoning quality, and interpretation accuracy.
- Identify analytical errors, flawed assumptions, and reasoning gaps that experienced data professionals would immediately recognize.
- Collaborate with AI researchers and fellow subject matter experts to improve benchmark quality, consistency, and analytical rigor.
Required Qualifications
- Master's degree, PhD, or equivalent practical experience in Data Science, Statistics, Mathematics, Economics, Operations Research, or another quantitative STEM discipline.
- Minimum 1 year of professional experience in research, research engineering, quantitative analysis, data science, or another data-intensive analytical role.
- Strong hands-on experience with data cleaning, exploratory data analysis, statistical testing, correlation analysis, predictive modeling, and interpretation of analytical results.
- Proficiency with Jupyter Notebooks or Google Colab for building reproducible analytical workflows.
- Strong programming skills in Python, including experience with libraries such as pandas, NumPy, SciPy, scikit-learn, or similar analytical frameworks.
- Working knowledge of Git and collaborative software development practices.
- Excellent written communication skills with the ability to present analytical findings clearly to both technical and non-technical audiences.
- Experience with AI evaluation, benchmark development, AI model assessment, or task authoring is preferred.
- Exceptional analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended problems independently.
- Ability to commit approximately 35 hours per week on a consistent basis.
Preferred Qualifications
- Experience designing reproducible research workflows or analytical evaluation frameworks.
- Familiarity with machine learning, large language models, or AI-assisted data analysis.
- Background in benchmark design, statistical modeling, or quantitative research methodology.
- Experience reviewing analytical work, mentoring analysts, or contributing to research publications.
Why Join
- Help shape the future of AI by improving how advanced models are evaluated on real-world analytical reasoning.
- Collaborate with leading AI researchers developing frontier evaluation benchmarks.
- Apply your expertise in statistics and data science to advance AI reliability and decision-making capabilities.
- Contribute directly to benchmark development that influences the evolution of next-generation AI systems.
- Enjoy the flexibility of a fully remote engagement while working on impactful AI research initiatives.
Equal Opportunity
We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.
Contract & Engagement Details
- Independent contractor engagement.
- Fully remote with flexible working hours.
- Expected commitment of approximately 35 hours per week.
- Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
- Work does not require access to confidential or proprietary information from any current or former employer.
- Payments are issued weekly based on approved work completed.
- At this time, we are unable to support H1-B or STEM OPT candidates.
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
Complete your application on apply.workable.com. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
This role is for one of our clients Compensation: $60-$90 per hour Join a pioneering AI initiative focused on developing the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Data Scientists and Quantitative Analysts to bring real-world analytical rigor to AI evaluation by designing sophisticated benchmark tasks based on practical data science workflows.
- Location & working pattern
United States
In this role, you will create complex, multi-step analytical challenges that mirror real research and business scenarios—from cleaning datasets and comparing statistical methods to interpreting results and presenting actionable insights. Working closely with AI researchers, you'll help identify where advanced AI models succeed, where they fail, and how evaluation benchmarks can better measure analytical reasoning. This is a fully remote, full-time engagement requiring approximately 35 hours per week. Requirements
More source context
- Contribute directly to benchmark development that influences the evolution of next-generation AI systems. - Enjoy the flexibility of a fully remote engagement while working on impactful AI research initiatives. Equal Opportunity
More relevant text appears in the full description.
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Aug 15, 2026
- Recorded sightings
- 172
- Last seen by us
- Oct 9, 2026
- Employer says posted
- Jul 29, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.