Software Engineer — Agentic data pipelines
San Diego, California, United States
- Pay
- Salary not listed in the saved posting
- Work setup
- Unconfirmed
- Employment
- Unconfirmed
What you’ll work on
Full postingDevelop LLM-based pipelines for data cleaning, normalization, and formatting across diverse data modalities (e.g., molecular, genomic, clinical, literature)
Implement automated quality-control workflows that detect anomalies, flag inconsistencies, and enforce data standards
Collaborate with ML scientists on the Enchant team to understand data requirements and translate them into scalable acquisition and processing systems
From the employer’s posting
KEY RESPONSIBILITIES Design, build, and maintain agentic systems that turn a pointer to a biomedical data source (web page, S3 path, GitHub repository, a table in a paper) into a reviewed, versioned dataset. The agent's output is committed pipeline code and its test suite, so re-running it later reproduces the same dataset. Develop LLM-based pipelines for data cleaning, normalization, and formatting across diverse data modalities (e.g., molecular, genomic, clinical, literature) Implement automated quality-control workflows that detect anomalies, flag inconsistencies, and enforce data standards
Design, build, and maintain agentic systems that turn a pointer to a biomedical data source (web page, S3 path, GitHub repository, a table in a paper) into a reviewed, versioned dataset. The agent's output is committed pipeline code and its test suite, so re-running it later reproduces the same dataset. Develop LLM-based pipelines for data cleaning, normalization, and formatting across diverse data modalities (e.g., molecular, genomic, clinical, literature) Implement automated quality-control workflows that detect anomalies, flag inconsistencies, and enforce data standards Evaluate and iterate on agent architectures, prompting strategies, tool definitions, validation loops, and evaluation harnesses that make agent-generated code trustworthy, improving reliability and throughput over time
Evaluate and iterate on agent architectures, prompting strategies, tool definitions, validation loops, and evaluation harnesses that make agent-generated code trustworthy, improving reliability and throughput over time Collaborate with ML scientists on the Enchant team to understand data requirements and translate them into scalable acquisition and processing systems Monitor and maintain distributed data pipelines in production, diagnosing failures and improving robustness over time
What you’ll bring
All qualificationsCore experience
- Master’s degree in a computational STEM field, or a Bachelor's with 2+ years of industry experience
- Strong Python engineering skills, including experience building and maintaining production-quality software
- Hands-on experience with LLM APIs (e.g., Claude, GPT) and agentic patterns such as tool use, orchestration, and multi-step reasoning
- Familiarity with biomedical or chemical data sources and formats (e.g., PDB, UniProt, ChEMBL, SDF/MOL, FASTA, or similar)
- Hands-on experience with Python testing frameworks (e.g., pytest fixtures, parametrization)
Preferred experience
- Experience with agent orchestration frameworks, and with evaluation harnesses for LLM-generated code
- Familiarity with cloud infrastructure and workflow orchestration (e.g., AWS, Docker, Kubernetes)
- Knowledge of multimodal biomedical data—spanning small molecules, proteins, assays, images, ‘omics, and/or clinical records
- Experience with large-scale dataset construction or curation for ML model training
Qualification wording
Master’s degree in a computational STEM field, or a Bachelor's with 2+ years of industry experience
Strong Python engineering skills, including experience building and maintaining production-quality software
Hands-on experience with LLM APIs (e.g., Claude, GPT) and agentic patterns such as tool use, orchestration, and multi-step reasoning
Familiarity with biomedical or chemical data sources and formats (e.g., PDB, UniProt, ChEMBL, SDF/MOL, FASTA, or similar)
Hands-on experience with Python testing frameworks (e.g., pytest fixtures, parametrization)
Experience with agent orchestration frameworks, and with evaluation harnesses for LLM-generated code
Familiarity with cloud infrastructure and workflow orchestration (e.g., AWS, Docker, Kubernetes)
Knowledge of multimodal biomedical data—spanning small molecules, proteins, assays, images, ‘omics, and/or clinical records
Experience with large-scale dataset construction or curation for ML model training
Tools in this posting
- Python
- AWS
- Docker
- S3
- Kubernetes
Source — Tool mentions in context
- Master’s degree in a computational STEM field, or a Bachelor's with 2+ years of industry experience - Strong Python engineering skills, including experience building and maintaining production-quality software - Hands-on experience with LLM APIs (e.g., Claude, GPT) and agentic patterns such as tool use, orchestration, and multi-step reasoning
- Comfort with data engineering fundamentals: ETL design, data validation, and working with structured and unstructured data at scale - Hands-on experience with Python testing frameworks (e.g., pytest fixtures, parametrization) PREFERRED QUALIFICATIONS
- Experience with agent orchestration frameworks, and with evaluation harnesses for LLM-generated code - Familiarity with cloud infrastructure and workflow orchestration (e.g., AWS, Docker, Kubernetes) - Knowledge of multimodal biomedical data—spanning small molecules, proteins, assays, images, ‘omics, and/or clinical records
KEY RESPONSIBILITIES - Design, build, and maintain agentic systems that turn a pointer to a biomedical data source (web page, S3 path, GitHub repository, a table in a paper) into a reviewed, versioned dataset. The agent's output is committed pipeline code and its test suite, so re-running it later reproduces the same dataset. Develop LLM-based pipelines for data cleaning, normalization, and formatting across diverse data modalities (e.g., molecular, genomic, clinical, literature) - Implement automated quality-control workflows that detect anomalies, flag inconsistencies, and enforce data standards
About Iambic-Therapeutics
Our mission is to deliver better medicines through innovations in AI-based discovery technologies.
In the employer’s words · Read in context
Job description
JOB SUMMARY
We are seeking a software engineer to join our team at Iambic Therapeutics, working on data acquisition and curation for Enchant, our multimodal transformer model trained at scale on a wide variety of biomedical data. In this role, you will design and build agentic systems that generate code to acquire, clean, format, quality-control, and generate auditable data reports for the large-scale datasets that power Enchant training. The model writes the code, the code runs the pipeline. You will work at the intersection of LLM-based automation and biomedical data engineering—developing AI agents that can navigate heterogeneous data sources, enforce quality standards, and operate reliably at scale.
This role is ideal for candidates who combine strong software engineering instincts with scientific understanding of biomedical data, and who are excited about using LLMs as tools to solve practical data problems.
KEY RESPONSIBILITIES
Design, build, and maintain agentic systems that turn a pointer to a biomedical data source (web page, S3 path, GitHub repository, a table in a paper) into a reviewed, versioned dataset. The agent's output is committed pipeline code and its test suite, so re-running it later reproduces the same dataset. Develop LLM-based pipelines for data cleaning, normalization, and formatting across diverse data modalities (e.g., molecular, genomic, clinical, literature)
Implement automated quality-control workflows that detect anomalies, flag inconsistencies, and enforce data standards
Evaluate and iterate on agent architectures, prompting strategies, tool definitions, validation loops, and evaluation harnesses that make agent-generated code trustworthy, improving reliability and throughput over time
Collaborate with ML scientists on the Enchant team to understand data requirements and translate them into scalable acquisition and processing systems
Monitor and maintain distributed data pipelines in production, diagnosing failures and improving robustness over time
Document data provenance, processing decisions, and quality metrics to support reproducibility and auditing
Operate the agents safely with sandboxed execution, least-privilege credentials, restricted network access, audit logs, and raising potential security risks to the team
QUALIFICATIONS
Master’s degree in a computational STEM field, or a Bachelor's with 2+ years of industry experience
Strong Python engineering skills, including experience building and maintaining production-quality software
Hands-on experience with LLM APIs (e.g., Claude, GPT) and agentic patterns such as tool use, orchestration, and multi-step reasoning
Familiarity with biomedical or chemical data sources and formats (e.g., PDB, UniProt, ChEMBL, SDF/MOL, FASTA, or similar)
Comfort with data engineering fundamentals: ETL design, data validation, and working with structured and unstructured data at scale
Hands-on experience with Python testing frameworks (e.g., pytest fixtures, parametrization)
PREFERRED QUALIFICATIONS
Experience with agent orchestration frameworks, and with evaluation harnesses for LLM-generated code
Familiarity with cloud infrastructure and workflow orchestration (e.g., AWS, Docker, Kubernetes)
Knowledge of multimodal biomedical data—spanning small molecules, proteins, assays, images, ‘omics, and/or clinical records
Experience with large-scale dataset construction or curation for ML model training
Knowledge of agent security practices: sandboxing, scoped credentials, prompt injection
Interest in a longer-term project: a natural language orchestrator that lets drug prosecution team members request inference, fine-tuning, virtual screens, and dataset analysis without writing code using our internal tools
ABOUT IAMBIC THERAPEUTICS
Iambic is a clinical-stage life-science and technology company developing novel medicines using its AI-driven discovery and development platform. Based in San Diego and founded in 2020, Iambic has assembled a world-class team that unites pioneering AI experts and experienced drug hunters. The Iambic platform has demonstrated delivery of new drug candidates to human clinical trials with unprecedented speed and across multiple target classes and mechanisms of action. Iambic is advancing a pipeline of potential best-in-class and first-in-class clinical assets, both internally and in partnership, to address urgent unmet patient need. Learn more about the Iambic team, platform, pipeline, and partnerships at iambic.ai.
MISSION & CORE VALUES
Our mission is to deliver better medicines through innovations in AI-based discovery technologies. The culture and work at Iambic Therapeutics are profoundly strengthened by the diversity of our people and our differences in background, culture, national origin, religion, sexual orientation, and life experiences. We are committed to building an inclusive environment where a diverse group of talented humans work together to discover therapeutics and create technologies.
PAY AND BENEFITS
We offer industry leading competitive pay, company paid healthcare, flexible spending accounts, voluntary life insurance, 401K matching, and uncapped vacation to our team. We are in a brand-new state-of-the art facility in beautiful San Diego with an onsite gym, dining, and easy access to great places to live and play.
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
- Ask the employer about the salary range before committing time to the process.
Complete your application on jobs.ashbyhq.com. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
No pay amount identified in the saved description.
- Location & working pattern
San Diego, California, United States
Working pattern and location restrictions need checking in the full posting.
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Sep 5, 2026
- Recorded sightings
- 21
- Last seen by us
- Oct 8, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.