ML Infrastructure Engineer
San Francisco, California, USA
- Pay
- Salary not listed in the saved posting
- Work setup
- Unconfirmed
- Employment
- Unconfirmed
What you’ll work on
Full postingYou'll build training pipelines that handle deep transformer models on hundreds of terabytes of 3D point cloud and image data.
Develop and maintain reliable, reproducible ML training and data generation pipelines.
Create CI/CD workflows for validating data pipelines and model training runs, including automated correctness checks and regression detection.
From the employer’s posting
This role is ideal for mid-career ML infrastructure engineers with experience building for both training and inference. You'll build training pipelines that handle deep transformer models on hundreds of terabytes of 3D point cloud and image data. You'll also architect our inference infrastructure, delivering both heavy offline detection algorithms and real-time responsive inference that integrates directly with our CAD software. Responsibilities
Design and build a centralized system for versioning training data, generated datasets, and model artifacts, with full lineage tracking from raw source data through to trained model outputs. Develop and maintain reliable, reproducible ML training and data generation pipelines. Refactor and harden existing training and data generation scripts into composable, testable, and maintainable components.
Refactor and harden existing training and data generation scripts into composable, testable, and maintainable components. Create CI/CD workflows for validating data pipelines and model training runs, including automated correctness checks and regression detection. Build tooling that enables ML engineers to launch, monitor, and debug training jobs with minimal friction.
What you’ll bring
All qualificationsCore experience
- 3+ years of work experience in relevant fields.
- Familiarity with AWS infrastructure services.
- Bachelor's or Master's degree in Computer Science, Engineering, or equivalent experience.
- Experience with containerized ML workflows and GPU-accelerated training environments.
- Strong communication skills and the ability to work closely with ML researchers and engineers to understand their workflows and translate them into robust systems.
- Experience with model optimization techniques (e.g., quantization, TensorRT, ONNX Runtime, distillation).
Qualification wording
3+ years of work experience in relevant fields.
Familiarity with AWS infrastructure services.
Bachelor's or Master's degree in Computer Science, Engineering, or equivalent experience.
Experience with containerized ML workflows and GPU-accelerated training environments.
Strong communication skills and the ability to work closely with ML researchers and engineers to understand their workflows and translate them into robust systems.
Experience with model optimization techniques (e.g., quantization, TensorRT, ONNX Runtime, distillation).
Tools in this posting
- Python
- AWS
- Prefect
- Airflow
- Terraform
- PyTorch
Source — Tool mentions in context
- Ability to read and refactor ML training code — you don't need to design model architectures, but you need to understand what training pipelines are doing well enough to make them reliable. - Proficient with Python, PyTorch. Bonus qualifications
Bonus qualifications - Familiarity with AWS infrastructure services. - Experience with containerized ML workflows and GPU-accelerated training environments.
- Experience with model optimization techniques (e.g., quantization, TensorRT, ONNX Runtime, distillation). - Knowledge of infrastructure-as-code tools (e.g., AWS CDK, Terraform). - Experience building or operating ML systems that handle large unstructured datasets (imagery, 3D data, sensor data).
- Experience designing and building data versioning, artifact management, or dataset lineage systems (e.g., DVC, LakeFS, Weights & Biases, or custom solutions). - Hands-on experience with ML pipeline orchestration tools (e.g., Airflow, Prefect, Metaflow, or similar). - Experience with model serving and inference optimization — profiling latency, reducing memory footprint, or scaling serving infrastructure to meet real-time constraints.
Job description
At Mach9, ML infrastructure engineers build and maintain the systems that power production AI models for civil engineering and surveying. Our ML pipeline spans 10,000+ miles of labeled survey data, image segmentation networks, and 3D prediction models serving real-time inference to surveyors and engineers in the field.
This role is ideal for mid-career ML infrastructure engineers with experience building for both training and inference.
You'll build training pipelines that handle deep transformer models on hundreds of terabytes of 3D point cloud and image data. You'll also architect our inference infrastructure, delivering both heavy offline detection algorithms and real-time responsive inference that integrates directly with our CAD software.
ResponsibilitiesDesign and build a centralized system for versioning training data, generated datasets, and model artifacts, with full lineage tracking from raw source data through to trained model outputs.
Develop and maintain reliable, reproducible ML training and data generation pipelines.
Refactor and harden existing training and data generation scripts into composable, testable, and maintainable components.
Create CI/CD workflows for validating data pipelines and model training runs, including automated correctness checks and regression detection.
Build tooling that enables ML engineers to launch, monitor, and debug training jobs with minimal friction.
Optimize and scale real-time model inference services to meet latency and throughput requirements in production, including profiling, batching strategies, and resource-efficient serving.
Own the deployment path from trained model artifact to production endpoint, ensuring reliable rollouts, rollback, and monitoring.
3+ years of work experience in relevant fields.
Bachelor's or Master's degree in Computer Science, Engineering, or equivalent experience.
Strong communication skills and the ability to work closely with ML researchers and engineers to understand their workflows and translate them into robust systems.
Experience designing and building data versioning, artifact management, or dataset lineage systems (e.g., DVC, LakeFS, Weights & Biases, or custom solutions).
Hands-on experience with ML pipeline orchestration tools (e.g., Airflow, Prefect, Metaflow, or similar).
Experience with model serving and inference optimization — profiling latency, reducing memory footprint, or scaling serving infrastructure to meet real-time constraints.
Ability to read and refactor ML training code — you don't need to design model architectures, but you need to understand what training pipelines are doing well enough to make them reliable.
Proficient with Python, PyTorch.
Familiarity with AWS infrastructure services.
Experience with containerized ML workflows and GPU-accelerated training environments.
Experience with model optimization techniques (e.g., quantization, TensorRT, ONNX Runtime, distillation).
Knowledge of infrastructure-as-code tools (e.g., AWS CDK, Terraform).
Experience building or operating ML systems that handle large unstructured datasets (imagery, 3D data, sensor data).
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
- Ask the employer about the salary range before committing time to the process.
Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
No pay amount identified in the saved description.
- Location & working pattern
San Francisco, California, USA
Working pattern and location restrictions need checking in the full posting.
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Jun 2, 2026
- Recorded sightings
- 31
- Last seen by us
- Oct 8, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.