Data Engineer, Red Tape Index
About Labrynth
Labrynth accelerates progress by streamlining regulatory complexity. We build AI-powered platforms that navigate complex regulations, generate audit-level documentation, and provide certainty, not shortcuts. Our technology serves clients across heavily regulated industries including energy, compliance, and government regulations.
We operate as a forward-deployed engineering organization: small, high-velocity teams embedded directly with clients to rapidly discover needs and ship production-quality solutions.
About the Role
We are hiring a Data Engineer to build the data platform behind our regulatory indices: acquiring fragmented public data, transforming it into clean, auditable datasets, and constructing the index methodology that turns it into published rankings.
This is a data platform role more than a pure pipeline or backend role. You will sit close to the raw sources and close to the math. The work spans three modes:
Acquire: source data from fragmented and often hostile places, including government open-data portals, APIs, HTML, PDFs, legacy Excel formats, login-protected portals, and commercial sites behind anti-bot protection.
Transform: normalize inconsistent jurisdictional data through bronze β silver β gold pipelines with idempotent ingestion, content hashing, and run-level lineage.
Construct: turn clean data into transparent, auditable indices through winsorization, percentile ranks, weighting, composites, and sensitivity testing.
What You'll Do
Ship scrapers and ingestion flows against messy, sometimes adversarial sources, using HTTP/2 clients, TLS-fingerprint evasion, and browser automation fallbacks, and keep them resilient as sources change
Own Postgres schema design and migrations end to end across per-country and per-domain schemas
Build and maintain medallion (bronze β silver β gold) transforms that are idempotent, content-hashed, and lineage-tracked
Implement and defend index methodology: normalization, weighting, and composite construction where the math verifiably says what it claims (our scoring core is held to 100% test coverage)
Assess data feasibility early, clarify requirements with partners, and convert ambiguous index ideas into executable plans
Take an index end to end: sourcing, validation, methodology, publication, and refresh planning
Operate pipelines on our orchestration stack (Prefect dispatching per-flow ECS Fargate tasks) with observability everywhere
What We're Looking For
Our stack is deliberately modern (Python 3.14, uv, ruff, ty, polars, Prefect 3, marimo). We don't filter on those exact tools; we hire for Python and data depth and expect a short ramp.
Strong Python and SQL; you have designed Postgres schemas and owned migrations (SQLAlchemy and Alembic, or equivalents) in production
Data pipeline experience with a lakehouse/medallion mindset: idempotent ingestion, content hashing, and lineage are habits, not aspirations
Web scraping beyond requests: anti-bot evasion, browser automation, and resilience against messy or hostile sources
Statistics literacy for index methodology: winsorization, normalization, weighting, and sensitivity testing, and you can reason about whether an index's math supports its claims
Comfort with modern Python tooling and CI discipline: typing, linting, coverage gates, and conventional commits
Product discovery instincts: you talk with partners in plain language, assess data feasibility before committing, and flag what is proven versus assumed
End-to-end ownership: you are a pragmatic generalist who moves across data, backend, infrastructure, and basic product decisions in an uncertain environment
Nice to Have
Prefect experience, or Airflow/Dagster with willingness to switch
AWS (ECS, S3) and Terraform
polars, pyarrow, and marimo or a Jupyter background
LLM-in-pipeline experience (pydantic-ai, AWS Bedrock, evals)
Actuarial, quantitative research, or data science background in ranking or index construction
Experience with government open data (permits, energy, environmental, or economic datasets)
Comfort working alongside AI tooling; our repos are agent-forward (Claude agent teams, spec-driven docs)
What We Offer
High-impact work at the intersection of AI and critical infrastructure regulation
End-to-end ownership of indices, from raw source to published methodology
Small team with outsized influence; your feasibility calls shape what we build
Modern AI-native development environment (Claude Code, Cursor, multi-model orchestration)
Remote-first
Competitive compensation
Values We Hire For
Character: integrity and trustworthiness above all
Competency: evoking trust and reliably delivering
Togetherness: family-level support and alignment
Impact: meaningful outcomes over activity
Commitment: ownership and follow-through
Equal Opportunity Statement
Weβre an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, or veteran status, or any other basis protected by law.