Machine Learning Engineer — Training Optimization
Remote (world)
- Pay
- Salary not listed in the saved posting
- Work setup
Remote stated — work setup source
Listed location: Remote (world)
Read the full posting- Employment
- Unconfirmed
What you’ll work on
Full postingWe’re looking for an ML Engineer focused on training optimization to help us scale and improve large-scale model training.
This is a high-impact role with real ownership: your work directly affects how fast we can iterate, how large we can scale, and how efficiently we deploy new models.
Optimize large-scale model training pipelines (throughput, convergence, stability, and cost)
Work on cutting-edge models and training systems at scale
Improve distributed training strategies (data, model, and pipeline parallelism)
From the employer’s posting
We’re looking for an ML Engineer focused on training optimization to help us scale and improve large-scale model training. You’ll work at the intersection of research and production, optimizing training pipelines for speed, stability, and cost—while collaborating closely with researchers pushing model architecture and capability forward.
This is a high-impact role with real ownership: your work directly affects how fast we can iterate, how large we can scale, and how efficiently we deploy new models.
What You’ll Do Optimize large-scale model training pipelines (throughput, convergence, stability, and cost) Improve distributed training strategies (data, model, and pipeline parallelism)
Real ownership at Series-A stage — your work shapes the company’s trajectory Work on cutting-edge models and training systems at scale Small, highly technical team with fast feedback loops
Optimize large-scale model training pipelines (throughput, convergence, stability, and cost) Improve distributed training strategies (data, model, and pipeline parallelism) Tune optimizers, schedulers, batch sizing, and precision (bf16 / fp16 / fp8)
Tools in this posting
- PyTorch
Source — Tool mentions in context
- Distributed systems for ML training - Experience with PyTorch (required) - Comfort working close to hardware (GPUs, memory, networking constraints)
Job description
About the Role
We’re looking for an ML Engineer focused on training optimization to help us scale and improve large-scale model training. You’ll work at the intersection of research and production, optimizing training pipelines for speed, stability, and cost—while collaborating closely with researchers pushing model architecture and capability forward.
This is a high-impact role with real ownership: your work directly affects how fast we can iterate, how large we can scale, and how efficiently we deploy new models.
What You’ll Do
Optimize large-scale model training pipelines (throughput, convergence, stability, and cost)
Improve distributed training strategies (data, model, and pipeline parallelism)
Tune optimizers, schedulers, batch sizing, and precision (bf16 / fp16 / fp8)
Reduce training time and compute cost via profiling, bottleneck analysis, and systems-level improvements
Collaborate with researchers on architecture-aware training strategies
Build and maintain robust training infrastructure (checkpointing, fault tolerance, reproducibility)
Evaluate and integrate new training techniques (e.g. gradient checkpointing, ZeRO, FSDP, custom kernels)
Own training performance metrics and continuously push them forward
What We’re Looking For
Strong experience training large neural networks (LLMs or similarly large models)
Hands-on experience with training optimization (not just model usage)
Solid understanding of:
Backpropagation, optimization algorithms, and training dynamics
Distributed systems for ML training
Experience with PyTorch (required)
Comfort working close to hardware (GPUs, memory, networking constraints)
Ability to move fluidly between research ideas and production-ready code
Nice to Have
Experience with large-scale distributed training (multi-node, multi-GPU)
Familiarity with DeepSpeed, FSDP, Megatron, or custom training stacks
Experience optimizing training on AMD or NVIDIA GPUs
Contributions to open-source ML infrastructure or research codebases
Exposure to non-Transformer architectures (RNNs, hybrid models, etc.)
Why Join Us
Real ownership at Series-A stage — your work shapes the company’s trajectory
Work on cutting-edge models and training systems at scale
Small, highly technical team with fast feedback loops
Strong emphasis on engineering quality and research rigor
Competitive compensation + meaningful equity
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
- Ask the employer about the salary range before committing time to the process.
Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
No pay amount identified in the saved description.
- Location & working pattern
Remote (world)
- Contributions to open-source ML infrastructure or research codebases - Exposure to non-Transformer architectures (RNNs, hybrid models, etc.) Why Join Us
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Jun 2, 2026
- Recorded sightings
- 23
- Last seen by us
- Sep 28, 2026
- Employer says posted
- Jan 22, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.