Back to jobs

MLOps / DevOps Engineer – AI Infrastructure

New York City, New York, United States

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at Creatoriq

Tools in this posting

  • Bash
  • Python
  • AWS
  • Databricks
  • Grafana
  • Kubernetes
  • S3
  • SageMaker
  • Terraform
  • Google Cloud (GCP)
Source — Tool mentions in context
- Solid understanding of networking fundamentals, including routing, load balancing, network security, and related concepts. - Scripting experience with Python, Bash, or similar languages to automate infrastructure and operational tasks. Nice to have
- 2+ years of hands-on experience supporting production ML/AI infrastructure, including model serving, training pipelines, and model monitoring. - 2+ Hands-on experience with Databricks (clusters, jobs, ML pipelines, Model Serving) and/or AWS SageMaker (endpoints, inference, model deployment, provisioned throughput, autoscaling) in a production environment. - 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar services.
- 2+ Hands-on experience with Databricks (clusters, jobs, ML pipelines, Model Serving) and/or AWS SageMaker (endpoints, inference, model deployment, provisioned throughput, autoscaling) in a production environment. - 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar services. - 2+ years of experience working with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS.
- Familiarity with Helm and service mesh technologies such as Istio, Linkerd, Traefik, or similar tools. - Experience designing serverless and event-driven architectures (beyond basic usage) using AWS Lambda, API Gateway, and SQS. - Exposure to cloud and infrastructure security practices, including vulnerability management and tools such as Nessus, Prowler, Trivy, or firewalls.
MLOps & ML Platform Infrastructure - Operate and scale ML platform infrastructure, including Databricks and Sagemaker clusters, jobs compute, ML pipelines, and Model Serving endpoints. - Manage production model-serving infrastructure, including compute capacity, provisioned throughput, and autoscaling for high-throughput inference workloads.
Observability, Incident Response & Engineering Collaboration - Maintain monitoring, logging, metrics, and alerting solutions using tools such as Prometheus, Grafana, Coralogix, and CloudWatch. - Support incident response and perform Root Cause Analysis (RCA) for infrastructure and deployment-related issues.
- Knowledge of security standards, compliance requirements, and cloud security best practices. - Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, Grafana, or similar solutions. - FinOps experience, including cloud cost monitoring, optimization, and accountability practices.
- 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar services. - 2+ years of experience working with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS. - Strong experience building and maintaining CI/CD pipelines using tools such as GitLab CI/CD or Jenkins.
- Support and maintain scalable, highly available, and secure cloud infrastructure in accordance with company policies and standards. - Provision and manage cloud resources using Infrastructure as Code (Terraform, Terragrunt, CloudFormation). - Implement cloud security best practices, including IAM/role-based access controls, encryption, vulnerability management, and secure infrastructure configurations.
- Strong experience building and maintaining CI/CD pipelines using tools such as GitLab CI/CD or Jenkins. - Hands-on experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar technologies. - Strong Linux system administration and troubleshooting skills.
Nice to have - Experience with Google Cloud, particularly for candidates who have worked across multi-cloud environments. - Familiarity with Helm and service mesh technologies such as Istio, Linkerd, Traefik, or similar tools.

About Creatoriq

CreatorIQ is the operating system for creator-led growth, helping global brands and agencies transform creator marketing into an intelligence-driven growth engine.

In the employer’s words · Read in context

Job description

View original posting ↗

CreatorIQ is the operating system for creator-led growth trusted by more than 1,300 global brands and agencies.

We’re on a mission to make businesses more human, and humans more impactful. We operate by our values — be intentional, pursue excellence every day, embrace the journey together, and be a good human — every day. CreatorIQ has earned the title of best companies to work for in multiple programs, including BuiltIn LA and NY. It’s been named a Fastest-Growing Company in North America on the Deloitte Technology Fast 500™ for four years, was named a leader in IDC MarketScape: Worldwide Influencer Marketing Platforms for Large Enterprises in 2025, was named a Leader by The Forrester New Wave™: Influencer Marketing Solutions, and has been consistently recognized by G2 as a Leader, and is rated 5 stars on Influencer MarketingHub. We operate in a flexible work model that combines both in-person and remote work to boost collaboration, enhance innovation, and adapt to individual work styles.

We're seeking passionate, innovative minds to join our journey. Be a part of our dynamic team and let's transform the industry together!

MLOps / DevOps Engineer – AI Infrastructure

The DevOps Engineer is responsible for supporting and improving cloud and ML/AI infrastructure, automating deployments, and maintaining CI/CD pipelines to ensure efficient, secure, and scalable development workflows. This role plays a crucial part in infrastructure automation, monitoring, and cloud security while collaborating with Software and ML Engineers, Product support, QA, and Security teams.
As a key member of the DevOps team, the DevOps Engineer helps manage cloud environments, CI/CD pipelines, and Infrastructure as Code (IaC), ensuring high availability and compliance with security best practices.


In this role, you’ll get to:

Cloud Infrastructure, Security & Reliability

  • Support and maintain scalable, highly available, and secure cloud infrastructure in accordance with company policies and standards.

  • Provision and manage cloud resources using Infrastructure as Code (Terraform, Terragrunt, CloudFormation).

  • Implement cloud security best practices, including IAM/role-based access controls, encryption, vulnerability management, and secure infrastructure configurations.

  • Support containerized environments and orchestration platforms.

  • Apply DevSecOps principles across infrastructure and deployment workflows.

  • Participate in disaster recovery planning, testing, and recovery activities.

CI/CD, Automation & Deployment

  • Maintain and optimize CI/CD pipelines using tools such as GitLab CI/CD and Jenkins, supporting application and ML model deployments.

  • Improve deployment reliability and support zero-downtime deployment strategies.

  • Automate configuration management, infrastructure provisioning, and routine operational processes.

  • Troubleshoot deployment and pipeline issues and implement improvements to prevent recurrence.

  • Develop scripts and automation to reduce manual work and improve engineering efficiency.

AI & Agentic Infrastructure

  • Help design, deploy, operate, and secure infrastructure supporting AI and agentic products, including MCP, agents, integrations, internal tooling, and customer-facing use cases.

  • Use AI-assisted engineering tools, coding copilots, and AI-driven troubleshooting to improve DevOps productivity and reduce repetitive operational work.

  • Evaluate and adopt practical AI-enabled workflows that improve infrastructure management, troubleshooting, and operational efficiency.

MLOps & ML Platform Infrastructure

  • Operate and scale ML platform infrastructure, including Databricks and Sagemaker clusters, jobs compute, ML pipelines, and Model Serving endpoints.

  • Manage production model-serving infrastructure, including compute capacity, provisioned throughput, and autoscaling for high-throughput inference workloads.

  • Maintain infrastructure-level monitoring for model drift, data quality, inference performance, and serving health, while partnering with ML Engineering on model evaluation, quality thresholds, and model correctness.

  • Partner with ML Engineering to support reliable CI/CD and production deployment of ML models.

Observability, Incident Response & Engineering Collaboration

  • Maintain monitoring, logging, metrics, and alerting solutions using tools such as Prometheus, Grafana, Coralogix, and CloudWatch.

  • Support incident response and perform Root Cause Analysis (RCA) for infrastructure and deployment-related issues.

  • Improve system observability through effective log aggregation, metrics collection, monitoring, and alerting.

  • Partner with Software Engineers, ML Engineers, QA, and Software Engineers in Test to improve deployment workflows and integrate automated testing into CI/CD pipelines.

  • Collaborate with IT Security to maintain secure cloud operations and infrastructure policies.

  • Respond to engineering and Product Support requests in a timely manner and provide technical infrastructure support when needed.

  • Maintain accurate internal technical and operational documentation.

  • Collaborate effectively with international teams across multiple time zones

Who you are and what you’ll need for this position:

  • 3+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-focused role.

  • 2+ years of hands-on experience supporting production ML/AI infrastructure, including model serving, training pipelines, and model monitoring.

  • 2+ Hands-on experience with Databricks (clusters, jobs, ML pipelines, Model Serving) and/or AWS SageMaker (endpoints, inference, model deployment, provisioned throughput, autoscaling) in a production environment.

  • 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar services.

  • 2+ years of experience working with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS.

  • Strong experience building and maintaining CI/CD pipelines using tools such as GitLab CI/CD or Jenkins.

  • Hands-on experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar technologies.

  • Strong Linux system administration and troubleshooting skills.

  • Solid understanding of networking fundamentals, including routing, load balancing, network security, and related concepts.

  • Scripting experience with Python, Bash, or similar languages to automate infrastructure and operational tasks.

Nice to have

  • Experience with Google Cloud, particularly for candidates who have worked across multi-cloud environments.

  • Familiarity with Helm and service mesh technologies such as Istio, Linkerd, Traefik, or similar tools.

  • Experience designing serverless and event-driven architectures (beyond basic usage) using AWS Lambda, API Gateway, and SQS.

  • Exposure to cloud and infrastructure security practices, including vulnerability management and tools such as Nessus, Prowler, Trivy, or firewalls.

  • Knowledge of security standards, compliance requirements, and cloud security best practices.

  • Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, Grafana, or similar solutions.

  • FinOps experience, including cloud cost monitoring, optimization, and accountability practices.

  • Experience with API gateway or API management platforms such as Kong or Apigee.

  • Hands-on experience using AI tools to improve engineering workflows, automation, troubleshooting, or agentic use cases.

Confidence can sometimes hold us back from applying for a job. But we'll let you in on a secret: there's no such thing as a 'perfect' candidate. Have 50% of the criteria? Excited about this opportunity? Passionate about what we do at CreatorIQ? Please apply! CreatorIQ is a place where everyone can grow.

What you will get from us:

  • People: Work with talented, collaborative, and friendly people who love what they do.

  • Guidance: Utilize our learning platform to fully get the training and tools you'll need to become successful here from your first day with us.

  • Work/life harmony: 15 days of vacation, floating and company holidays, wellness benefits, and paid parental leave.

  • Whole Health Package: Comprehensive medical, dental, vision, life, and disability insurance, plus additional wellness benefits.

  • Planning for the future: A 401(k) plan to help you plan ahead.

  • Work from home stipend: To assist you in setting up a home office that works for you.

Who we are:

CreatorIQ is the operating system for creator-led growth, helping global brands and agencies transform creator marketing into an intelligence-driven growth engine. Powered by the Creator Graph™, which processes more than 250 million social posts daily across more than 15 million creators worldwide, CreatorIQ unifies fragmented platform data into a centralized intelligence layer and system of record for creator relationships, performance, governance, and commerce. More than 1,300 organizations—including Dentsu, Delta Air Lines, Google, Beiersdorf, Nestlé, and Wella—rely on CreatorIQ as the infrastructure to run and scale their creator programs globally. CreatorIQ is a global company headquartered in Los Angeles with offices in Austin, New York, San Francisco, London, Manila, and Warsaw. Learn more at www.creatoriq.com and follow us on LinkedIn and Instagram.

At CreatorIQ, we believe that diversity is the key to unlocking our full potential. We are committed to fostering an inclusive, equitable, and empowering work environment where everyone can thrive, regardless of race, ethnicity, gender, sexual orientation, age, religion, disability, or any other characteristic that makes us unique. By embracing our core values of being intentional, pursuing excellence every day, embracing the journey together, being a good human, and staying focused on what’s important, we create an atmosphere that promotes collaboration and growth. Join us to celebrate differences, innovate together, and be a part of a business that is disrupting the marketing industry.

Compensation, benefits, and beyond:

We understand that a comprehensive benefits package plays a significant role in your overall compensation. To gain more insight into the various components of our total compensation, we invite you to review our benefits and perks.

AI Transparency Notice

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications and note taking during interviews. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please refer to our Global Candidate Privacy Notice.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

New York City, New York, United States

CreatorIQ is the operating system for creator-led growth trusted by more than 1,300 global brands and agencies. We’re on a mission to make businesses more human, and humans more impactful. We operate by our values — be intentional, pursue excellence every day, embrace the journey together, and be a good human — every day. CreatorIQ has earned the title of best companies to work for in multiple programs, including BuiltIn LA and NY. It’s been named a Fastest-Growing Company in North America on the Deloitte Technology Fast 500™ for four years, was named a leader in IDC MarketScape: Worldwide Influencer Marketing Platforms for Large Enterprises in 2025, was named a Leader by The Forrester New Wave™: Influencer Marketing Solutions, and has been consistently recognized by G2 as a Leader, and is rated 5 stars on Influencer MarketingHub. We operate in a flexible work model that combines both in-person and remote work to boost collaboration, enhance innovation, and adapt to individual work styles. We're seeking passionate, innovative minds to join our journey. Be a part of our dynamic team and let's transform the industry together!
Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 27, 2026
Recorded sightings
9
Last seen by us
Oct 8, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.