Back to jobs

CV/ML Platform Engineer

Austin, TX

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Allen Control Systems

What you’ll work on

Full posting
  • Manage NVIDIA GPU clusters for ML training.

  • Own the ACS CV/ML CI/CD pipeline.

  • Improve and maintain core ML infrastructure, such as model registration and versioning, experiment tracking, and model and data provenance tracking.

From the employer’s posting
Deploy and operate Kubernetes clusters on bare-metal infrastructure hosting 130+ NVIDIA GPUs, with hybrid burst capability to AWS for scalable compute and storage workloads. Manage NVIDIA GPU clusters for ML training. Own the ACS CV/ML CI/CD pipeline.
Manage NVIDIA GPU clusters for ML training. Own the ACS CV/ML CI/CD pipeline. Improve and maintain core ML infrastructure, such as model registration and versioning, experiment tracking, and model and data provenance tracking.
Own the ACS CV/ML CI/CD pipeline. Improve and maintain core ML infrastructure, such as model registration and versioning, experiment tracking, and model and data provenance tracking. Improve and maintain ML model testing, performance analysis, and reporting tools.

Tools in this posting

  • AWS
  • Kubernetes
  • Grafana
  • MLflow
Source — Tool mentions in context
What You'll Do: - Deploy and operate Kubernetes clusters on bare-metal infrastructure hosting 130+ NVIDIA GPUs, with hybrid burst capability to AWS for scalable compute and storage workloads. - Manage NVIDIA GPU clusters for ML training.
Position Overview We are seeking an experienced CV/ML Platform Engineer with specialization in Computer Vision and Machine Learning (CV/ML) to design, build, and own the data, model, and compute infrastructure powering ACS CV/ML team. You will help manage a 130+ GPU bare-metal Kubernetes cluster, own CV/ML CI/CD pipelines, and ensure ML model training proceeds at high volume with low friction. What You'll Do:
- Experience with model optimization toolchains, including TensorRT, ONNX, and quantization techniques, specifically for cross-compilation to ARM targets like NVIDIA Jetson. - Proficiency with observability stacks (ELK, Prometheus/Grafana) adapted for ML, including monitoring GPU health, training throughput, and model inference metrics. - Strong Linux systems knowledge (Debian/Ubuntu), including networking for high-throughput data, storage, and security hardening for defense-grade production environments.
- Hands-on experience with NVIDIA GPU infrastructure, including managing CUDA libraries and development environments, GPU Operator, device plugins, and scheduling (MIG, Volcano, or fractional GPU sharing). - Experience implementing and maintaining MLOps platforms such as Kubeflow, MLflow, Weights & Biases (W&B), or DVC for experiment tracking and model versioning. - Familiarity with high-performance storage solutions (e.g., MinIO, WEKA, or Ceph) and data orchestration tools capable of handling terabytes of video/image data.

Benefits in the posting

Full benefits wording
  • Competitive salary
  • Health, Dental, Vision Insurance
  • Paid Time Off

From the employer’s posting.

Job description

View original posting ↗

Company Overview 

Allen Control Systems (ACS) is a cutting-edge defense startup founded by two former Navy electrical engineers with a proven track record in robotics and software. We are developing a small, autonomous gun turret that employs advanced computer vision and control systems to precisely target and neutralize small drones and loitering munitions. Our innovative approach requires overcoming significant technical challenges, making this an exciting and dynamic environment for experienced engineers. 

With an engineering-first culture, ACS values technical excellence and innovation. Backed by our founders’ successful exits from two previous ventures acquired for a combined $180M in 2022, we are committed to ensuring that the groundbreaking technologies we develop have a real-world impact. 

Position Overview

We are seeking an experienced CV/ML Platform Engineer with specialization in Computer Vision and Machine Learning (CV/ML) to design, build, and own the data, model, and compute infrastructure powering ACS CV/ML team. You will help manage a 130+ GPU bare-metal Kubernetes cluster, own CV/ML CI/CD pipelines, and ensure ML model training proceeds at high volume with low friction. 

What You'll Do: 

  • Deploy and operate Kubernetes clusters on bare-metal infrastructure hosting 130+ NVIDIA GPUs, with hybrid burst capability to AWS for scalable compute and storage workloads. 
  • Manage NVIDIA GPU clusters for ML training.
  • Own the ACS CV/ML CI/CD pipeline.
  • Improve and maintain core ML infrastructure, such as model registration and versioning, experiment tracking, and model and data provenance tracking.
  • Improve and maintain ML model testing, performance analysis, and reporting tools.
  • Automate repetitive model training and testing tasks to increase developer velocity.
  • Work with Software Team Platform Engineers to ensure efficient coordination and minimal duplication between CV/ML infrastructure and wider Software infrastructure.
  • Collaborate with the Software Team to automate the optimization of models (TensorRT/quantization) for deployment on NVIDIA Jetson and other edge hardware. 

Required Technical Skills: 

  • 2+ years of experience in Platform Engineering or DevOps/MLOps. 
  • Strong programming skills are required for automating ML lifecycles and building custom CLI tools for CV engineers. 
  • Hands-on experience with NVIDIA GPU infrastructure, including managing CUDA libraries and development environments, GPU Operator, device plugins, and scheduling (MIG, Volcano, or fractional GPU sharing). 
  • Experience implementing and maintaining MLOps platforms such as Kubeflow, MLflow, Weights & Biases (W&B), or DVC for experiment tracking and model versioning. 
  • Familiarity with high-performance storage solutions (e.g., MinIO, WEKA, or Ceph) and data orchestration tools capable of handling terabytes of video/image data. 
  • Proven track record building CI/CD pipelines that include automated model validation, performance benchmarking, and artifact management for both cloud and edge targets. 
  • Experience with model optimization toolchains, including TensorRT, ONNX, and quantization techniques, specifically for cross-compilation to ARM targets like NVIDIA Jetson. 
  • Proficiency with observability stacks (ELK, Prometheus/Grafana) adapted for ML, including monitoring GPU health, training throughput, and model inference metrics. 
  • Strong Linux systems knowledge (Debian/Ubuntu), including networking for high-throughput data, storage, and security hardening for defense-grade production environments.  

What We Offer 

  • Competitive salary
  • Health, Dental, Vision Insurance
  • Paid Time Off 

Allen Control Systems is an Equal Opportunity Employer, providing equal employment opportunities to all employees and applicants for employment. Allen Control Systems prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.  

#LI-AS1 

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on job-boards.greenhouse.io. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

Austin, TX

What You'll Do: - Deploy and operate Kubernetes clusters on bare-metal infrastructure hosting 130+ NVIDIA GPUs, with hybrid burst capability to AWS for scalable compute and storage workloads. - Manage NVIDIA GPU clusters for ML training.
Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Mar 21, 2026
Recorded sightings
193
Last seen by us
Oct 6, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.