← back to jobs
> job detail
I
🧪Data Scientist

Research Intern

Infrrd · Bangalore, Karnataka, India
// classified as
Data Scientist (Modeling, experiments, research.)
location
Bangalore, Karnataka, India
languages
python
tools
> stack
python
> description

About Infrrd

Infrrd (pronounced In-fur-d) is an Enterprise AI company that automates document-heavy workflows for customers in mortgage, insurance, and finance. Our Research team works on the next generation of document intelligence: agentic systems that read, reason over, and audit complex documents with outputs that can be trusted and verified. We are looking for a Research Intern to join the team and contribute to experiments that shape what we ship.


About the Role

As a Research Intern, you will work on well-scoped research tasks under the guidance of senior researchers, across areas such as agentic document extraction, LLM-based auditing of mortgage documents, table extraction with calibrated trust scores, and verifiable evaluation of model outputs without ground truth. You will run experiments end to end: preparing and checking data, building prototypes, analysing errors and reasoning traces, and writing up what you found. This is a hands-on role for someone who enjoys rigorous experimentation and wants exposure to real enterprise-scale document AI problems.

Education details: 10+ 2/PUC mandatory (No Diploma), B.E/B.Tech/M.Tech students from all Computer Science related backgrounds with a focus on machine learning, NLP, or computer vision.

Year of Graduation: 2027

Percentage criteria: minimum 60% aggregate and higher throughout academics.

Internship duration: 1 year with an opportunity to convert to a full-time role based on performance.



What You Will Do

  • Experimentation and Prototyping: Design and run experiments to validate document processing and agentic extraction approaches; build prototypes and proof-of-concept implementations using LLM and vision-language model APIs.
  • Evaluation and Verification: Help build evaluation harnesses and verifier checks (cross-field consistency, structural invariants, multi-pass agreement) that measure whether an extraction or audit verdict can be trusted, including when no ground truth is available.
  • Error and Trace Analysis: Conduct in-depth error analysis on model outputs and agent reasoning traces to identify failure modes, categorise them, and propose fixes.
  • Data Quality and EDA: Verify the quality of datasets and synthetic document packages used in experiments; perform exploratory analysis to understand document characteristics and edge cases.
  • Rule and Checklist Work: Assist in converting domain checklist rules into executable, testable checks and in measuring their precision and recall on real documents.
  • Literature Tracking: Read and summarise recent papers on document AI, agent harnesses, RL post-training, and evaluation; present findings in internal paper discussions.
  • Tooling and Workflow: Use AI coding assistants (Claude Code, Copilot, or similar) and internal tools effectively; track progress in Jira; participate actively in stand-ups and code reviews.
  • Documentation and Communication: Document methodology, experiment setup, and results clearly so they are reproducible; contribute to technical reports, Confluence pages, and internal presentations.


Who You Are

  • Strong mathematical, statistical, and probabilistic foundation with a solid grasp of core ML concepts.
  • Strong Python skills, including writing clean, testable code within a larger codebase.
  • Working knowledge of Transformer-based language models and how to use LLM APIs (prompting, structured outputs, tool or function calling).
  • Familiarity with evaluation methodology: designing metrics, building test sets, and analysing results with rigour rather than anecdotes.
  • Academic or project experience in NLP, computer vision, or document understanding (OCR, layout, tables, forms).
  • Ability to run experiments scientifically, keep track of what was tried, and communicate outcomes clearly.
  • Curiosity about agentic systems and initiative in learning new techniques and applying them to real problems.


Good to Have

  • Experience with vision-language models or document-specific models for extraction and layout understanding.
  • Exposure to agent frameworks, multi-agent orchestration, or harness design for LLM-based systems.
  • Familiarity with RL post-training methods (GRPO, RLVR) or model fine-tuning.
  • Experience with table extraction, PDF parsing, or synthetic data generation.
  • Contributions to open source, published work, or a portfolio of research projects.


Pay Stubs Data Extraction

NLP-powered Table Extraction for Insurance Policy Data

Ally | Agentic AI for Mortgage

Automated Data Extraction from Engineering & Construction Drawings