Back to jobs

ML/AI Research Engineer — Agentic AI Lab (Founding Team)

San Francisco Bay Area · San Francisco Bay Area, California, United States

Pay
Salary not listed in the saved posting
Work setup
On-site — work setup source
Location type: On-site
From the employer’s posting
Employment
Full-time — employment source
Employment type: Full-time
From the employer’s posting
Team
Engineering — team source
Department: Engineering
From the employer’s posting

Before you apply

Sponsorship
Visa sponsorship not confirmed — sponsorship source
Do you now or will you in the future require visa sponsorship to work?
From the employer’s application form
Apply at Fabrion

What you’ll work on

Full posting

We’re designing the future of enterprise AI infrastructure — grounded in agents, retrieval-augmented generation (RAG), knowledge graphs, and multi-tenant governance.

We’re looking for an ML/AI Research Engineer to join our AI Lab and lead the design, training, evaluation, and optimization of agent-native AI models.

What you’ll bring

All qualifications

Preferred experience

  • Deep experience fine-tuning open-source LLMs using HuggingFace Transformers, DeepSpeed, vLLM, FSDP, LoRA/QLoRA
  • Experience training or customizing agent frameworks with multi-step reasoning and memory
  • Experience running models under quantized (int4/int8) or multi-GPU settings with inference tuning (vLLM, TGI)
  • Experience building enterprise-grade RAG pipelines integrated with real-time or contextual data
Qualification wording
Deep experience fine-tuning open-source LLMs using HuggingFace Transformers, DeepSpeed, vLLM, FSDP, LoRA/QLoRA
Experience training or customizing agent frameworks with multi-step reasoning and memory
Experience running models under quantized (int4/int8) or multi-GPU settings with inference tuning (vLLM, TGI)
Experience building enterprise-grade RAG pipelines integrated with real-time or contextual data

Tools in this posting

  • Python
  • Rust
  • SQL
  • Delta
  • DuckDB
  • Kubernetes
  • Neo4j
  • SageMaker
  • Huggingface
  • Transformers
  • JavaScript
  • PostgreSQL
Source — Tool mentions in context
- Compute: Ray, Kubernetes, TGI, Sagemaker, LambdaLabs, Modal - Languages: Python (core), optionally Rust (for inference layers) or JS (for UX experimentation) Soft Skills & Mindset
- Familiar with LangChain, LangGraph, LlamaIndex, and open-source vector DBs (Weaviate, Qdrant, FAISS) - Experience grounding models with structured data (SQL, graph, metadata) + unstructured sources - Bonus: Worked with Neo4j, Puppygraph, RDF, OWL, or other semantic modeling systems
- Graph Knowledge Systems: Neo4j, Puppygraph, RDF, Gremlin, JSON-LD - Storage & Access: Iceberg, DuckDB, Postgres, Parquet, Delta Lake - Evaluation: OpenLLM Evals, Trulens, Ragas, LangSmith, Weight & Biases
- Evaluation: OpenLLM Evals, Trulens, Ragas, LangSmith, Weight & Biases - Compute: Ray, Kubernetes, TGI, Sagemaker, LambdaLabs, Modal - Languages: Python (core), optionally Rust (for inference layers) or JS (for UX experimentation)
- Experience grounding models with structured data (SQL, graph, metadata) + unstructured sources - Bonus: Worked with Neo4j, Puppygraph, RDF, OWL, or other semantic modeling systems Agent Intelligence:
- Vector DBs: Weaviate, Qdrant, FAISS, Pinecone, Chroma - Graph Knowledge Systems: Neo4j, Puppygraph, RDF, Gremlin, JSON-LD - Storage & Access: Iceberg, DuckDB, Postgres, Parquet, Delta Lake
Model Training: - Deep experience fine-tuning open-source LLMs using HuggingFace Transformers, DeepSpeed, vLLM, FSDP, LoRA/QLoRA - Worked with both base and instruction-tuned models; familiar with SFT, RLHF, DPO pipelines
Preferred Tech Stack - LLM Training & Inference: HuggingFace Transformers, DeepSpeed, vLLM, FlashAttention, FSDP, LoRA - Agent Orchestration: LangChain, LangGraph, ReAct, OpenAgents, LlamaIndex

Job description

View original posting ↗

ML/AI Research Engineer — Agentic AI Lab (Founding Team)

Location: San Francisco Bay Area
Type: Full-Time
Compensation: Competitive salary + meaningful equity (founding tier)

Backed by 8VC, we're building a world-class team to tackle one of the industry’s most critical infrastructure problems.

About the Role

We’re designing the future of enterprise AI infrastructure — grounded in agents, retrieval-augmented generation (RAG), knowledge graphs, and multi-tenant governance.

We’re looking for an ML/AI Research Engineer to join our AI Lab and lead the design, training, evaluation, and optimization of agent-native AI models. You'll work at the intersection of LLMs, vector search, graph reasoning, and reinforcement learning — building the intelligence layer that sits on top of our enterprise data fabric.

This isn’t a prompt engineer role. It’s full-cycle ML: from data curation and fine-tuning to evaluation, interpretability, and deployment — with cost-awareness, alignment, and agent coordination all in scope.

Core Responsibilities

  • Fine-tune and evaluate open-source LLMs (e.g. LLaMA 3, Mistral, Falcon, Mixtral) for enterprise use cases with both structured and unstructured data

  • Build and optimize RAG pipelines using LangChain, LangGraph, LlamaIndex, or Dust — integrated with our vector DBs and internal knowledge graph

  • Train agent architectures (ReAct, AutoGPT, BabyAGI, OpenAgents) using enterprise task data

  • Develop embedding-based memory and retrieval chains with token-efficient chunking strategies

  • Create reinforcement learning pipelines to optimize agent behaviors (e.g. RLHF, DPO, PPO)

  • Establish scalable evaluation harnesses for LLM and agent performance, including synthetic evals, trace capture, and explainability tools

  • Contribute to model observability, drift detection, error classification, and alignment

  • Optimize inference latency and GPU resource utilization across cloud and on-prem environments

Desired Experience

Model Training:

  • Deep experience fine-tuning open-source LLMs using HuggingFace Transformers, DeepSpeed, vLLM, FSDP, LoRA/QLoRA

  • Worked with both base and instruction-tuned models; familiar with SFT, RLHF, DPO pipelines

  • Comfortable building and maintaining custom training datasets, filters, and eval splits

  • Understand tradeoffs in batch size, token window, optimizer, precision (FP16, bfloat16), and quantization

RAG + Knowledge Graphs:

  • Experience building enterprise-grade RAG pipelines integrated with real-time or contextual data

  • Familiar with LangChain, LangGraph, LlamaIndex, and open-source vector DBs (Weaviate, Qdrant, FAISS)

  • Experience grounding models with structured data (SQL, graph, metadata) + unstructured sources

  • Bonus: Worked with Neo4j, Puppygraph, RDF, OWL, or other semantic modeling systems

Agent Intelligence:

  • Experience training or customizing agent frameworks with multi-step reasoning and memory

  • Understand common agent loop patterns (e.g. Plan→Act→Reflect), memory recall, and tools

  • Familiar with self-correction, multi-agent communication, and agent ops logging

Optimization:

  • Strong background in token cost optimization, chunking strategies, reranking (e.g. Cohere, Jina), compression, and retrieval latency tuning

  • Experience running models under quantized (int4/int8) or multi-GPU settings with inference tuning (vLLM, TGI)

Preferred Tech Stack

  • LLM Training & Inference: HuggingFace Transformers, DeepSpeed, vLLM, FlashAttention, FSDP, LoRA

  • Agent Orchestration: LangChain, LangGraph, ReAct, OpenAgents, LlamaIndex

  • Vector DBs: Weaviate, Qdrant, FAISS, Pinecone, Chroma

  • Graph Knowledge Systems: Neo4j, Puppygraph, RDF, Gremlin, JSON-LD

  • Storage & Access: Iceberg, DuckDB, Postgres, Parquet, Delta Lake

  • Evaluation: OpenLLM Evals, Trulens, Ragas, LangSmith, Weight & Biases

  • Compute: Ray, Kubernetes, TGI, Sagemaker, LambdaLabs, Modal

  • Languages: Python (core), optionally Rust (for inference layers) or JS (for UX experimentation)

Soft Skills & Mindset

  • Startup DNA: resourceful, fast-moving, and capable of working in ambiguity

  • Deep curiosity about agent-based architectures and real-world enterprise complexity

  • Comfortable owning model performance end-to-end: from dataset to deployment

  • Strong instincts around explainability, safety, and continuous improvement

  • Enjoy pair-designing with product and UX to shape capabilities, not just APIs

Why This Role Matters

This role is foundational to our thesis: that agents + enterprise data + knowledge modeling can create intelligent infrastructure for real-world, multi-billion-dollar workflows. Your work won’t be buried in research reports — it will be productionized and activated by hundreds of users and hundreds of thousands of decisions. If this is your dream role - we would love to hear from you.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

San Francisco Bay Area, California, United States

Working pattern and location restrictions need checking in the full posting.

Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Jun 2, 2026
Recorded sightings
22
Last seen by us
Sep 29, 2026
Employer says posted
Aug 28, 2025

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.