Data Pipeline & Ingestion Engineer
About the Program
The Operational Data Layer (ODL) program is a large-scale data-platform build for a leading benefitsadministration
platform. The platform ingests data from multiple legacy benefits systems, masters it into
golden records, transforms it into a canonical data model, and serves it through modern APIs — built on an
AWS / Java / Kafka stack in a HIPAA/SOX-regulated benefits domain spanning health, wealth/401(k), spending
accounts, and leaves.
About the Role
You will build and operate the data backbone of ODL: bulk and streaming ingestion from legacy source
systems, medallion-layered storage (Bronze/Silver/Gold), identity resolution and golden-record consolidation,
source-to-canonical mapping and crosswalks, and the data-quality and reconciliation gates that prove data is
complete and correct before it is published. This is the volume engine of the program — every new client
onboarded flows through the pipelines you build.
What You’ll Do
• Build batch-seed and event-tail ingestion per source system, including seed→tail watermark hand-off,
idempotent upserts, and dedup ledgers
• Build and operate medallion layers with reprocess-from-Bronze, pipeline orchestration (checkpoints,
retry/backoff, DLQ), and full observability
• Build data-quality gates (quarantine / pass-with-flag), quality scoring, and a reconciliation engine
covering count, record, and financial reconciliation — financial is zero-tolerance
• Build identity matching combining deterministic rules with probabilistic scoring and confidence bands;
deliver deduplication, golden-record materialization, and survivorship rules, calibrating match thresholds
with labelled data
• Author and maintain source→canonical structural mappings and value crosswalks (e.g., collapsing
1,800+ raw employment-status values to ~20 standard ones) as governed, versioned configuration
• Enforce data contracts at the boundary: schema registry, fail-fast validation, and semver-compatible
schema evolution
What We’re Looking For
• 5+ years building production data pipelines at scale
• Kafka depth: consumers/producers, replay, DLQ, exactly-once / idempotent processing patterns
• Strong SQL and solid ETL fundamentals
• Java and/or Python in production
• Medallion / lakehouse layering, CDC, watermark/checkpoint patterns, and batch–stream hand-off
• Data-quality frameworks: validation rules, quarantine and re-entry, quality scoring, reconciliation
• Entity resolution / MDM exposure: record matching, dedup, survivorship — via commercial tools
(Informatica MDM, Reltio) or custom builds
• Data mapping and crosswalk discipline: profiling messy datasets, authoring governed reference data,
config-as-code (YAML/JSON, Git)
Bonus Points
• Probabilistic record linkage at depth — blocking/candidate generation, scoring models, threshold
calibration (expected at senior level)
• Schema registry experience (Avro/Protobuf)
• Extracting from mainframe or older RDBMS sources with limited CDC support
• Financial reconciliation in finance-adjacent domains
• Benefits administration or healthcare domain knowledge