Back to jobs

On-device ML Performance Engineer, Graphics, Games and Machine Learning

Cupertino

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Check the employer’s page ↗

Tools in this posting

  • PyTorch
Source — Tool mentions in context
The On-device ML Performance team has the responsibility to analyze latency, memory, power and numerical correctness of the latest machine learning models running on Apple SoC’s, and to make Apple’s ML software stack take full advantage of the capabilities in Apple’s ML accelerators. The work from this cross functional team enables model developers’ decisions to optimize performance via advanced techniques such as different model authoring techniques, quantization, sparsity, performance and accuracy tradeoffs. The work of this team impacts all new Apple HW and ML Inference on them. Our group is looking for an On-device ML Performance Engineer, with technical expertise in computer architecture, performance, memory, power, ML model architectures, ML frameworks such as PyTorch, and on-device ML inference. The role entails deep analysis of ML models and their architecture, the implementation of the models in the ML SW stack for optimum performance, power and memory usage, and debug involving the performance and power consumption of CPU, GPU, and Apple Neural Engine.

Job description

View original posting ↗

The On-Device Machine Learning team at Apple is responsible for enabling the Research to Production lifecycle of cutting edge machine learning models that power magical user experiences on Apple’s hardware and software platforms. Apple is the best place to do on-device machine learning, and this team sits at the heart of that discipline, interfacing with research, SW engineering, HW engineering, and products. The On-device ML Performance team has the responsibility to analyze latency, memory, power and numerical correctness of the latest machine learning models running on Apple SoC’s, and to make Apple’s ML software stack take full advantage of the capabilities in Apple’s ML accelerators. The work from this cross functional team enables model developers’ decisions to optimize performance via advanced techniques such as different model authoring techniques, quantization, sparsity, performance and accuracy tradeoffs. The work of this team impacts all new Apple HW and ML Inference on them. Our group is looking for an On-device ML Performance Engineer, with technical expertise in computer architecture, performance, memory, power, ML model architectures, ML frameworks such as PyTorch, and on-device ML inference. The role entails deep analysis of ML models and their architecture, the implementation of the models in the ML SW stack for optimum performance, power and memory usage, and debug involving the performance and power consumption of CPU, GPU, and Apple Neural Engine.

Your next step

Check the employer’s posting for the current role and application details.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Cupertino

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Aug 29, 2026
Recorded sightings
216
Last seen by us
Oct 8, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error
AI answers unavailable

We couldn’t identify enough role detail in this saved description to support an AI answer. Read the full posting