AI/ML Engineer
-
Seeking an experienced Senior AI/ML Engineer specialized in Generative AI and Agentic Systems.
-
Role focuses on the end-to-end design, development, optimization, and deployment of autonomous and semi-autonomous AI agents.
-
Architect multi-agent orchestration systems, design enterprise-grade Retrieval-Augmented Generation (RAG) pipelines, and implement production-ready MLOps infrastructure utilizing Google Cloud Platform (GCP) and containerized environments.
Design, develop, and deploy autonomous and semi-autonomous AI agents utilizing leading frameworks such as LangChain, LangGraph, LlamaIndex, CrewAI, or Google's Agent Development Kit/Vertex AI Agent Builder to automate and optimize enterprise business processes.
Architect complex multi-agent systems, establishing robust orchestration patterns, task decomposition methods, cognitive planning loops, and inter-agent communication protocols like Model Context Protocol, function-calling, and structured tool execution.
Build, secure, and maintain integrations between AI agents and external enterprise systems, APIs (REST/GraphQL), databases, and Google Cloud services such as BigQuery, Cloud Functions, Cloud Run, and Vertex AI APIs.
Implement advanced memory management paradigms including short-term, long-term, episodic, and semantic memory to maintain context, state, and historical execution metadata across user sessions.
Author, test, and optimize advanced prompt templates and system instructions while establishing and maintaining reusable prompt libraries to ensure deterministic, safe, and repeatable agent behaviors.
Evaluate, select, and fine-tune foundation models such as Gemini or other open-source/proprietary models via Vertex AI Model Garden to balance model capability, execution latency, and API inference costs.
Design and optimize high-throughput, low-latency RAG pipelines, ensuring clean document ingestion, smart text chunking (semantic/character-based), and high-quality embedding generation.
Integrate and manage scalable vector databases such as Vertex AI Vector Search, AlloyDB, or similar vector stores to ground AI agent responses in verified enterprise knowledge.
Analyze, clean, and pre-process complex structured and unstructured data sources to optimize model ingestion and ensure efficient data access.
Write, test, and maintain declarative Infrastructure as Code scripts using Terraform to provision secure GCP environments, including GKE clusters, Cloud Run services, Vertex AI endpoints, storage buckets, and networking components.
Build and maintain CI/CD pipelines using Cloud Build, GitLab CI, or Jenkins for automated testing, container building, and seamless multi-environment deployment of agent configurations, prompt files, and backend tools.
Package application code, agents, and dependencies into secure Docker containers, orchestrating deployments on Google Kubernetes Engine (GKE) or deploying serverless workflows via Cloud Run.
Configure, schedule, and maintain workflow orchestrators such as Vertex AI Pipelines or Cloud Composer/Airflow to automate scheduled agent evaluation cycles, model fine-tuning, and data ingestion processes.
Establish comprehensive evaluation metrics and testing pipelines to measure task completion rates, reasoning depth, tool-calling precision, latency, token consumption, and hallucination rates utilizing LLM-as-a-judge and Vertex AI evaluation tools.
Implement robust input/output content filtering, moderation tools, grounding validators, prompt-injection defenses, data privacy checks, and human-in-the-loop approval gates.
Implement production monitoring, logging, distributed tracing, and real-time alerting using Cloud Monitoring, Cloud Logging, Cloud Trace, Prometheus, Grafana, or dedicated LLM observability tools like LangSmith.
Continually audit and optimize agent workflows, model parameters, caching strategies, and underlying infrastructure to maximize cost-efficiency and performance under high loads.
Adhere to strict version control standards using Git, managing branching, pull requests, and code review workflows for code, prompts, Dockerfiles, and Terraform scripts.
Partner closely with Data Scientists, Software Engineers, and business units to translate complex operational requirements into scalable production-grade AI solutions.
Maintain comprehensive technical documentation, including system architecture diagrams, agent flowcharts, tool definitions, prompt engineering strategies, containerization guidelines, and standard operating procedures.
Minimum of 3 years of professional experience in Machine Learning or AI Engineering with a strong foundation in MLOps practices.
At least 1 to 2 years of hands-on experience specifically designing and implementing LLM-based, generative AI, or agentic systems in production.
Expert-level Python programming skills and experience with standard machine learning libraries such as PyTorch, TensorFlow, or Scikit-learn.
Strong theoretical and practical understanding of deep learning, Natural Language Processing (NLP), and Transformer architectures.
Hands-on experience using agentic frameworks such as LangChain, LangGraph, LlamaIndex, CrewAI, or Vertex AI Agent Builder.
Proven experience with Google Cloud Platform (GCP) and container services including Docker, Kubernetes/GKE, Cloud Run, and Cloud Functions.
Proven experience working with vector indexing, semantic search, and databases like Vertex AI Vector Search, AlloyDB, or equivalents.
Solid understanding of database systems (SQL/NoSQL) and building or consuming REST and GraphQL APIs.
Experience with Terraform, Git, and automated CI/CD tools like Cloud Build, GitLab CI, or Jenkins.
Experience setting up monitoring solutions such as Prometheus, Cloud Logging, or LangSmith, and evaluating LLM outputs for quality and safety.
Excellent troubleshooting, debugging, and analytical skills for diagnosing complex agent behavior, tool failures, and infrastructure bottlenecks.
Exceptional collaborative and communication skills to effectively translate complex technical constraints to cross-functional stakeholders.
GCP Professional Machine Learning Engineer or Google Professional Cloud Architect certifications are highly desired.