Back to jobs

Senior Software Engineer, ML Infrastructure

San Francisco, California, United States

Pay
$500Pay period needs review — pay source
Who We Are Voxel is building the future of Computer Vision and Machine Learning for operations, risk, and safety. We use computer vision and AI to enable existing security cameras to automatically detect hazards and high-risk activities, keep people safe and drive operational efficiencies. Our technology addresses the key cost drivers for workers’ compensation, general liability, and property damage, which cost US employers over $500 billion annually. Our customers include Fortune 500 companies across grocery, retail, manufacturing, food and beverage, logistics, and pharmaceutical distribution. We’ve passed $10M ARR with strong expansion revenue. Based in SF, backed by industry-leading VCs. About the Role
Read the full posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Voxel

What you’ll work on

Full posting

Voxel’s perception system is the technical core of everything we ship.

We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models.

  • You’ll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production.

  • Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new architectures.

  • Write performant code that scales well in production environments.

From the employer’s posting
Voxel’s perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team.
We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models. You’ll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production. You'll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers.
Voxel’s perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team. We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models. You’ll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production. You'll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers. What You'll Do
What You'll Do Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new architectures. Own the train-to-deploy handoff - export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment.
Experience with AWS (S3, EC2, EKS, or similar) for ML workloads. Strong Python. Write performant code that scales well in production environments. Track record of owning infrastructure end-to-end: scoping, building, shipping, and improving systems that internal teams depend on.

Tools in this posting

  • Python
  • AWS
  • MLflow
  • Prefect
  • S3
  • PyTorch
Source — Tool mentions in context
- Experience with AWS (S3, EC2, EKS, or similar) for ML workloads. - Strong Python. Write performant code that scales well in production environments. - Track record of owning infrastructure end-to-end: scoping, building, shipping, and improving systems that internal teams depend on.
- Establish ML experiment tracking and lifecycle management - pick the right tools (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently. - Establish DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely. - Understand the infra needs of applied ML/CV engineers and design scalable solutions that support model development.
- Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar). - Experience with AWS (S3, EC2, EKS, or similar) for ML workloads. - Strong Python. Write performant code that scales well in production environments.
- Own the train-to-deploy handoff - export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment. - Establish ML experiment tracking and lifecycle management - pick the right tools (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently. - Establish DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely.
- Hands-on experience building ML training pipelines in PyTorch. - Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar). - Experience with AWS (S3, EC2, EKS, or similar) for ML workloads.
Nice to Have - Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar) - Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar)
- 4+ years of experience building and shipping large scale software solutions. - Hands-on experience building ML training pipelines in PyTorch. - Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar).
- Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar) - Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar) - Background in computer vision model training

Benefits in the posting

Full benefits wording
  • Equity through Voxel’s Equity Incentive Plan
  • Total compensation includes base salary, annual bonus, and equity
  • Comprehensive health, dental, and vision insurance
  • Competitive paid parental leave
  • Unlimited PTO and flexible work arrangements

From the employer’s posting.

About Voxel

Voxel is building the future of Computer Vision and Machine Learning for operations, risk, and safety.

In the employer’s words · Read in context

Job description

View original posting ↗

Who We Are

Voxel is building the future of Computer Vision and Machine Learning for operations, risk, and safety. We use computer vision and AI to enable existing security cameras to automatically detect hazards and high-risk activities, keep people safe and drive operational efficiencies. Our technology addresses the key cost drivers for workers’ compensation, general liability, and property damage, which cost US employers over $500 billion annually. Our customers include Fortune 500 companies across grocery, retail, manufacturing, food and beverage, logistics, and pharmaceutical distribution. We’ve passed $10M ARR with strong expansion revenue. Based in SF, backed by industry-leading VCs.

 

About the Role

Voxel’s perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team.

We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models. You’ll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production. You'll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers.

What You'll Do

  • Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new architectures.

  • Own the train-to-deploy handoff - export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment.

  • Establish ML experiment tracking and lifecycle management - pick the right tools (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently.

  • Establish DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely.

  • Understand the infra needs of applied ML/CV engineers and design scalable solutions that support model development.

What We're Looking For

  • 4+ years of experience building and shipping large scale software solutions.

  • Hands-on experience building ML training pipelines in PyTorch.

  • Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar).

  • Experience with AWS (S3, EC2, EKS, or similar) for ML workloads.

  • Strong Python. Write performant code that scales well in production environments.

  • Track record of owning infrastructure end-to-end: scoping, building, shipping, and improving systems that internal teams depend on.

  • Bias toward shipping. You'd rather ship something good this week than something perfect next quarter.

  • Strong communication skills.

Nice to Have

  • Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar)

  • Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar)

  • Background in computer vision model training

Compensation & Benefits

  • Equity through Voxel’s Equity Incentive Plan

  • Total compensation includes base salary, annual bonus, and equity

  • Comprehensive health, dental, and vision insurance

  • Competitive paid parental leave

  • Unlimited PTO and flexible work arrangements

  • Daily meals in-office, team events, annual company onsite

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.

Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay
Who We Are Voxel is building the future of Computer Vision and Machine Learning for operations, risk, and safety. We use computer vision and AI to enable existing security cameras to automatically detect hazards and high-risk activities, keep people safe and drive operational efficiencies. Our technology addresses the key cost drivers for workers’ compensation, general liability, and property damage, which cost US employers over $500 billion annually. Our customers include Fortune 500 companies across grocery, retail, manufacturing, food and beverage, logistics, and pharmaceutical distribution. We’ve passed $10M ARR with strong expansion revenue. Based in SF, backed by industry-leading VCs. About the Role
Location & working pattern

San Francisco, California, United States

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Jun 2, 2026
Recorded sightings
49
Last seen by us
Oct 6, 2026
Employer says posted
Apr 14, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.