Back to jobs
Apple

Machine Learning Engineer - Agentic AI Evaluation Frameworks

Cupertino

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Pay, work setup, and employment type unconfirmed

Not confirmed in this saved copy: pay, work setup, employment type. Check the full posting

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.

Already applied? Track this application

About applying

Apply opens the employer’s site in a new tab. Add your outcome here after you submit.

Source details & eligibility

Before you apply

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Cupertino

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Posting history
Status in our records
Active
First seen by us
Sep 9, 2026
Recorded sightings
2

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

Job description

Imagine what you could do here. At Apple, great ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your work, and there’s no telling what you could accomplish. The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered experiences. In this role, you will develop evaluation frameworks, datasets, tooling, and quality signals that enable teams to understand and continuously improve Generative AI and LLM-powered products. You will work closely with Machine Learning, Software Engineering, Quality Engineering, Product, Human Interface, Data Science, and domain experts to establish rigorous evaluation practices throughout the AI product lifecycle. You will help define how we measure the quality of AI experiences across the Commerce domain, including Store AI, Shopping AI, Learning AI, Content GenAI, Conversational AI, and Platform Self-Service. This is an opportunity to work at the intersection of machine learning, software engineering, data, and product quality, helping ensure our AI experiences are accurate, relevant, grounded, reliable, and useful for users around the world.