Back to jobs

Senior Inference Optimization ML Engineer

Mountain View, California, United States

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Rhoda-Ai

What you’ll work on

Full posting
  • Own inference performance end-to-end — diagnose and improve latency, throughput, and efficiency of large foundation models in production

  • Own the optimization layer that determines how quickly and efficiently our foundation models run in the real world — high ownership, high impact, small elite team

  • Build systematic performance attribution: latency decomposition (compute vs.

From the employer’s posting
What You'll Do Own inference performance end-to-end — diagnose and improve latency, throughput, and efficiency of large foundation models in production Build systematic performance attribution: latency decomposition (compute vs. memory bandwidth vs. I/O), bottleneck identification, and prioritization across model families
Direct leverage on research velocity and real-world robot performance — every efficiency gain you make accelerates model iteration and tightens the loop between model and robot behavior Own the optimization layer that determines how quickly and efficiently our foundation models run in the real world — high ownership, high impact, small elite team
Own inference performance end-to-end — diagnose and improve latency, throughput, and efficiency of large foundation models in production Build systematic performance attribution: latency decomposition (compute vs. memory bandwidth vs. I/O), bottleneck identification, and prioritization across model families Apply and develop optimization techniques including quantization, pruning, distillation, operator fusion, and model compilation (e.g., TensorRT, torch.compile, XLA)

Tools in this posting

  • PyTorch
Source — Tool mentions in context
- 3+ years of experience in inference optimization, ML systems, or a closely related field - Deep hands-on experience with modern ML stacks (PyTorch required; JAX a plus) - Strong understanding of compute, memory bandwidth, and I/O bottlenecks in large model inference

Job description

View original posting ↗

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

We're looking for an Inference Optimization MLE to help build and operate the systems that make our foundation models run fast and efficiently in production. You'll be responsible for squeezing maximum performance out of large multimodal models, across cloud and on-robot deployment targets. You will working closely with research and robotics teams to close the gap between training and real-world deployment.

What You'll Do

  • Own inference performance end-to-end — diagnose and improve latency, throughput, and efficiency of large foundation models in production

  • Build systematic performance attribution: latency decomposition (compute vs. memory bandwidth vs. I/O), bottleneck identification, and prioritization across model families

  • Apply and develop optimization techniques including quantization, pruning, distillation, operator fusion, and model compilation (e.g., TensorRT, torch.compile, XLA)

  • Optimize attention mechanisms, KV caching, and memory layouts for large multimodal models (vision, video, language, proprioception)

  • Work with kernel-level tooling (e.g., CUDA, Triton) to identify hotspots and implement or tune custom kernels where needed

  • Build benchmarking and regression detection infrastructure: latency baselines, throughput curves, and automated detection of performance regressions across model versions

  • Collaborate closely with research engineers to translate model innovations into optimized, deployment-ready implementations

What We're Looking For

  • 3+ years of experience in inference optimization, ML systems, or a closely related field

  • Deep hands-on experience with modern ML stacks (PyTorch required; JAX a plus)

  • Strong understanding of compute, memory bandwidth, and I/O bottlenecks in large model inference

  • Experience with model optimization techniques: quantization (INT8/FP8/AWQ), distillation, pruning, and compilation

  • Familiarity with inference serving frameworks (e.g., Triton, TensorRT, vLLM, TorchServe)

  • Exceptional debugging and measurement ability: turn "inference is slow" into clear bottlenecks, experiments, and validated improvements

  • High ownership mindset and comfort in a fast-moving environment

Nice to Have (But Not Required)

  • GPU kernel or compiler-level experience (CUDA, Triton, graph capture, operator fusion)

  • Experience with multimodal or video model inference (variable-length sequences, packing/bucketing)

  • Familiarity with edge/cloud hybrid deployment patterns and on-robot inference constraints

  • Experience with speculative decoding, continuous batching, or other LLM serving optimizations

  • Background in streaming or low-latency systems relevant to real-time robot control

Why This Role

  • Direct leverage on research velocity and real-world robot performance — every efficiency gain you make accelerates model iteration and tightens the loop between model and robot behavior

  • Own the optimization layer that determines how quickly and efficiently our foundation models run in the real world — high ownership, high impact, small elite team

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Mountain View, California, United States

- Experience with multimodal or video model inference (variable-length sequences, packing/bucketing) - Familiarity with edge/cloud hybrid deployment patterns and on-robot inference constraints - Experience with speculative decoding, continuous batching, or other LLM serving optimizations
Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 6, 2026
Recorded sightings
187
Last seen by us
Oct 8, 2026
Employer says posted
May 12, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.