Back to jobs

Machine Learning Engineer — Training Optimization

Remote (world)

Pay
Salary not listed in the saved posting
Work setup
Remote stated — work setup source
Listed location: Remote (world)
Read the full posting
Employment
Unconfirmed
Apply at Featherlessai

What you’ll work on

Full posting

We’re looking for an ML Engineer focused on training optimization to help us scale and improve large-scale model training.

This is a high-impact role with real ownership: your work directly affects how fast we can iterate, how large we can scale, and how efficiently we deploy new models.

  • Optimize large-scale model training pipelines (throughput, convergence, stability, and cost)

  • Work on cutting-edge models and training systems at scale

  • Improve distributed training strategies (data, model, and pipeline parallelism)

From the employer’s posting
We’re looking for an ML Engineer focused on training optimization to help us scale and improve large-scale model training. You’ll work at the intersection of research and production, optimizing training pipelines for speed, stability, and cost—while collaborating closely with researchers pushing model architecture and capability forward.
This is a high-impact role with real ownership: your work directly affects how fast we can iterate, how large we can scale, and how efficiently we deploy new models.
What You’ll Do Optimize large-scale model training pipelines (throughput, convergence, stability, and cost) Improve distributed training strategies (data, model, and pipeline parallelism)
Real ownership at Series-A stage — your work shapes the company’s trajectory Work on cutting-edge models and training systems at scale Small, highly technical team with fast feedback loops
Optimize large-scale model training pipelines (throughput, convergence, stability, and cost) Improve distributed training strategies (data, model, and pipeline parallelism) Tune optimizers, schedulers, batch sizing, and precision (bf16 / fp16 / fp8)

Tools in this posting

  • PyTorch
Source — Tool mentions in context
- Distributed systems for ML training - Experience with PyTorch (required) - Comfort working close to hardware (GPUs, memory, networking constraints)

Job description

View original posting ↗

About the Role

We’re looking for an ML Engineer focused on training optimization to help us scale and improve large-scale model training. You’ll work at the intersection of research and production, optimizing training pipelines for speed, stability, and cost—while collaborating closely with researchers pushing model architecture and capability forward.

This is a high-impact role with real ownership: your work directly affects how fast we can iterate, how large we can scale, and how efficiently we deploy new models.

What You’ll Do

  • Optimize large-scale model training pipelines (throughput, convergence, stability, and cost)

  • Improve distributed training strategies (data, model, and pipeline parallelism)

  • Tune optimizers, schedulers, batch sizing, and precision (bf16 / fp16 / fp8)

  • Reduce training time and compute cost via profiling, bottleneck analysis, and systems-level improvements

  • Collaborate with researchers on architecture-aware training strategies

  • Build and maintain robust training infrastructure (checkpointing, fault tolerance, reproducibility)

  • Evaluate and integrate new training techniques (e.g. gradient checkpointing, ZeRO, FSDP, custom kernels)

  • Own training performance metrics and continuously push them forward

What We’re Looking For

  • Strong experience training large neural networks (LLMs or similarly large models)

  • Hands-on experience with training optimization (not just model usage)

  • Solid understanding of:

    • Backpropagation, optimization algorithms, and training dynamics

    • Distributed systems for ML training

  • Experience with PyTorch (required)

  • Comfort working close to hardware (GPUs, memory, networking constraints)

  • Ability to move fluidly between research ideas and production-ready code

Nice to Have

  • Experience with large-scale distributed training (multi-node, multi-GPU)

  • Familiarity with DeepSpeed, FSDP, Megatron, or custom training stacks

  • Experience optimizing training on AMD or NVIDIA GPUs

  • Contributions to open-source ML infrastructure or research codebases

  • Exposure to non-Transformer architectures (RNNs, hybrid models, etc.)

Why Join Us

  • Real ownership at Series-A stage — your work shapes the company’s trajectory

  • Work on cutting-edge models and training systems at scale

  • Small, highly technical team with fast feedback loops

  • Strong emphasis on engineering quality and research rigor

  • Competitive compensation + meaningful equity

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Remote (world)

- Contributions to open-source ML infrastructure or research codebases - Exposure to non-Transformer architectures (RNNs, hybrid models, etc.) Why Join Us
Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Jun 2, 2026
Recorded sightings
23
Last seen by us
Sep 28, 2026
Employer says posted
Jan 22, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.