← back to jobs
> job detail
F
🤖ML Engineer

AI/ML Engineer

Ford Motor · Chennai, Tamil Nadu, India
// classified as
ML Engineer (Productionizing models, serving, MLOps.)
posted
1d ago
location
Chennai, Tamil Nadu, India
languages
python
tools
bigquery, docker, grafana
> stack
pythonbigquerydockergrafanakubernetesterraformpytorchtensorflow
> description
  • Seeking an experienced Senior AI/ML Engineer specialized in Generative AI and Agentic Systems. 

  • Role focuses on the end-to-end design, development, optimization, and deployment of autonomous and semi-autonomous AI agents. 

  • Architect multi-agent orchestration systems, design enterprise-grade Retrieval-Augmented Generation (RAG) pipelines, and implement production-ready MLOps infrastructure utilizing Google Cloud Platform (GCP) and containerized environments. 

  • Design, develop, and deploy autonomous and semi-autonomous AI agents utilizing leading frameworks such as LangChain, LangGraph, LlamaIndex, CrewAI, or Google's Agent Development Kit/Vertex AI Agent Builder to automate and optimize enterprise business processes. 

  • Architect complex multi-agent systems, establishing robust orchestration patterns, task decomposition methods, cognitive planning loops, and inter-agent communication protocols like Model Context Protocol, function-calling, and structured tool execution. 

  • Build, secure, and maintain integrations between AI agents and external enterprise systems, APIs (REST/GraphQL), databases, and Google Cloud services such as BigQuery, Cloud Functions, Cloud Run, and Vertex AI APIs. 

  • Implement advanced memory management paradigms including short-term, long-term, episodic, and semantic memory to maintain context, state, and historical execution metadata across user sessions. 

  • Author, test, and optimize advanced prompt templates and system instructions while establishing and maintaining reusable prompt libraries to ensure deterministic, safe, and repeatable agent behaviors. 

  • Evaluate, select, and fine-tune foundation models such as Gemini or other open-source/proprietary models via Vertex AI Model Garden to balance model capability, execution latency, and API inference costs. 

  • Design and optimize high-throughput, low-latency RAG pipelines, ensuring clean document ingestion, smart text chunking (semantic/character-based), and high-quality embedding generation. 

  • Integrate and manage scalable vector databases such as Vertex AI Vector Search, AlloyDB, or similar vector stores to ground AI agent responses in verified enterprise knowledge. 

  • Analyze, clean, and pre-process complex structured and unstructured data sources to optimize model ingestion and ensure efficient data access. 

  • Write, test, and maintain declarative Infrastructure as Code scripts using Terraform to provision secure GCP environments, including GKE clusters, Cloud Run services, Vertex AI endpoints, storage buckets, and networking components. 

  • Build and maintain CI/CD pipelines using Cloud Build, GitLab CI, or Jenkins for automated testing, container building, and seamless multi-environment deployment of agent configurations, prompt files, and backend tools. 

  • Package application code, agents, and dependencies into secure Docker containers, orchestrating deployments on Google Kubernetes Engine (GKE) or deploying serverless workflows via Cloud Run. 

  • Configure, schedule, and maintain workflow orchestrators such as Vertex AI Pipelines or Cloud Composer/Airflow to automate scheduled agent evaluation cycles, model fine-tuning, and data ingestion processes. 

  • Establish comprehensive evaluation metrics and testing pipelines to measure task completion rates, reasoning depth, tool-calling precision, latency, token consumption, and hallucination rates utilizing LLM-as-a-judge and Vertex AI evaluation tools. 

  • Implement robust input/output content filtering, moderation tools, grounding validators, prompt-injection defenses, data privacy checks, and human-in-the-loop approval gates. 

  • Implement production monitoring, logging, distributed tracing, and real-time alerting using Cloud Monitoring, Cloud Logging, Cloud Trace, Prometheus, Grafana, or dedicated LLM observability tools like LangSmith. 

  • Continually audit and optimize agent workflows, model parameters, caching strategies, and underlying infrastructure to maximize cost-efficiency and performance under high loads. 

  • Adhere to strict version control standards using Git, managing branching, pull requests, and code review workflows for code, prompts, Dockerfiles, and Terraform scripts. 

  • Partner closely with Data Scientists, Software Engineers, and business units to translate complex operational requirements into scalable production-grade AI solutions. 

  • Maintain comprehensive technical documentation, including system architecture diagrams, agent flowcharts, tool definitions, prompt engineering strategies, containerization guidelines, and standard operating procedures. 

  • Minimum of 3 years of professional experience in Machine Learning or AI Engineering with a strong foundation in MLOps practices. 

  • At least 1 to 2 years of hands-on experience specifically designing and implementing LLM-based, generative AI, or agentic systems in production. 

  • Expert-level Python programming skills and experience with standard machine learning libraries such as PyTorch, TensorFlow, or Scikit-learn. 

  • Strong theoretical and practical understanding of deep learning, Natural Language Processing (NLP), and Transformer architectures. 

  • Hands-on experience using agentic frameworks such as LangChain, LangGraph, LlamaIndex, CrewAI, or Vertex AI Agent Builder. 

  • Proven experience with Google Cloud Platform (GCP) and container services including Docker, Kubernetes/GKE, Cloud Run, and Cloud Functions. 

  • Proven experience working with vector indexing, semantic search, and databases like Vertex AI Vector Search, AlloyDB, or equivalents. 

  • Solid understanding of database systems (SQL/NoSQL) and building or consuming REST and GraphQL APIs. 

  • Experience with Terraform, Git, and automated CI/CD tools like Cloud Build, GitLab CI, or Jenkins. 

  • Experience setting up monitoring solutions such as Prometheus, Cloud Logging, or LangSmith, and evaluating LLM outputs for quality and safety. 

  • Excellent troubleshooting, debugging, and analytical skills for diagnosing complex agent behavior, tool failures, and infrastructure bottlenecks. 

  • Exceptional collaborative and communication skills to effectively translate complex technical constraints to cross-functional stakeholders. 

  • GCP Professional Machine Learning Engineer or Google Professional Cloud Architect certifications are highly desired.