Data Platform Engineer
Cape Town
- Pay
- Salary not listed in the saved posting
- Work setup
- Unconfirmed
- Employment
- Unconfirmed
What you’ll work on
Full postingDesign and build scalable ETL/ELT and streaming pipelines using Spark/Dataproc, Kafka, and Pub/Sub.
Build modular, fault-tolerant components with graceful schema evolution.
Implement monitoring, alerting, and data quality checks.
From the employer’s posting
What You'll Do: End-to-end domain ownership: Own platform components and a pipeline domain end-to-end, including design, implementation, reliability, performance, and cost. Design and build scalable ETL/ELT and streaming pipelines using Spark/Dataproc, Kafka, and Pub/Sub. Integrate diverse sources into BigQuery and other stores. Build modular, fault-tolerant components with graceful schema evolution. Infrastructure & performance tuning: Configure and maintain core platform infrastructure as code, including Dataproc, Kafka, Airflow/Astronomer, BigQuery, and Cloud Storage. Tune performance and cost through query optimisation, partitioning, and resource configuration. Tune streaming configurations for throughput and reliability.
Infrastructure & performance tuning: Configure and maintain core platform infrastructure as code, including Dataproc, Kafka, Airflow/Astronomer, BigQuery, and Cloud Storage. Tune performance and cost through query optimisation, partitioning, and resource configuration. Tune streaming configurations for throughput and reliability. Reliability & incident response: Own and track SLOs for freshness, success, and latency. Implement monitoring, alerting, and data quality checks. Participate in on-call as a Tier 2 responder, troubleshooting and resolving incidents independently. Security, governance & self-service: Implement security and governance by design, including access control, secrets, encryption, audit logging, lineage, and retention. Build self-service capabilities and paved roads for analytics engineers, analysts, and data scientists.
What you’ll bring
All qualificationsPreferred experience
- 3 to 5 years of experience in data platform, data engineering, or backend/distributed systems engineering, including ownership of production systems at TB-scale, or systems carrying stringent SLAs, multiple downstream consumers, or compliance requirements
- Strong Python, plus production experience in a JVM language, with the ability to write production-quality code with proper error handling, logging, and testing.
- Experience with workflow orchestration such as Airflow: scheduling, dependencies, retry logic, and workflow coordination
- Experience migrating legacy big-data platforms onto cloud-native alternatives
Qualification wording
3 to 5 years of experience in data platform, data engineering, or backend/distributed systems engineering, including ownership of production systems at TB-scale, or systems carrying stringent SLAs, multiple downstream consumers, or compliance requirements
Strong Python, plus production experience in a JVM language, with the ability to write production-quality code with proper error handling, logging, and testing. Scala is our primary pipeline language: if you don't have it yet, you'll need a real appetite to become strong in it
Experience with workflow orchestration such as Airflow: scheduling, dependencies, retry logic, and workflow coordination
Experience migrating legacy big-data platforms onto cloud-native alternatives
Tools in this posting
- Python
- Scala
- SQL
- BigQuery
- Datadog
- dbt
- Google Cloud (GCP)
- Grafana
- Kafka
- Looker
- Spark
- Terraform
- Airflow
Source — Tool mentions in context
- 3 to 5 years of experience in data platform, data engineering, or backend/distributed systems engineering, including ownership of production systems at TB-scale, or systems carrying stringent SLAs, multiple downstream consumers, or compliance requirements - Strong Python, plus production experience in a JVM language, with the ability to write production-quality code with proper error handling, logging, and testing. Scala is our primary pipeline language: if you don't have it yet, you'll need a real appetite to become strong in it - Hands-on depth in either distributed batch processing (Spark: DataFrames, Spark SQL, partitioning, shuffles, fault tolerance) or streaming (Kafka or Pub/Sub: producers, consumers, topics, partitions, stream processing patterns), with working familiarity of the other
The platform: The platform has strong foundations: Scala and Spark pipelines processing TB-scale data, dozens of managed data connectors, Airflow on Astronomer for orchestration, a mature dbt transformation layer, and BigQuery as the core analytical warehouse. You'll also have room to shape the platform's next chapter: modernising parts of the compute stack, and building out governance frameworks, data contracts, SLOs, observability, cost management, and data quality monitoring within your domain. What You'll Do:
- Familiarity with dbt, Looker, or semantic layer concepts, and with dbt orchestration through Astronomer Cosmos - Functional Scala patterns (cats-effect and the Typelevel ecosystem), or JVM tuning exposure such as garbage collection, memory management, and thread pools - Observability tooling such as Grafana, Datadog, or Cloud Monitoring
- Strong Python, plus production experience in a JVM language, with the ability to write production-quality code with proper error handling, logging, and testing. Scala is our primary pipeline language: if you don't have it yet, you'll need a real appetite to become strong in it - Hands-on depth in either distributed batch processing (Spark: DataFrames, Spark SQL, partitioning, shuffles, fault tolerance) or streaming (Kafka or Pub/Sub: producers, consumers, topics, partitions, stream processing patterns), with working familiarity of the other - Advanced SQL, including complex queries, window functions, CTEs, and query optimisation
- Hands-on depth in either distributed batch processing (Spark: DataFrames, Spark SQL, partitioning, shuffles, fault tolerance) or streaming (Kafka or Pub/Sub: producers, consumers, topics, partitions, stream processing patterns), with working familiarity of the other - Advanced SQL, including complex queries, window functions, CTEs, and query optimisation - Experience with workflow orchestration such as Airflow: scheduling, dependencies, retry logic, and workflow coordination
What You'll Do: - End-to-end domain ownership: Own platform components and a pipeline domain end-to-end, including design, implementation, reliability, performance, and cost. Design and build scalable ETL/ELT and streaming pipelines using Spark/Dataproc, Kafka, and Pub/Sub. Integrate diverse sources into BigQuery and other stores. Build modular, fault-tolerant components with graceful schema evolution. - Infrastructure & performance tuning: Configure and maintain core platform infrastructure as code, including Dataproc, Kafka, Airflow/Astronomer, BigQuery, and Cloud Storage. Tune performance and cost through query optimisation, partitioning, and resource configuration. Tune streaming configurations for throughput and reliability.
- End-to-end domain ownership: Own platform components and a pipeline domain end-to-end, including design, implementation, reliability, performance, and cost. Design and build scalable ETL/ELT and streaming pipelines using Spark/Dataproc, Kafka, and Pub/Sub. Integrate diverse sources into BigQuery and other stores. Build modular, fault-tolerant components with graceful schema evolution. - Infrastructure & performance tuning: Configure and maintain core platform infrastructure as code, including Dataproc, Kafka, Airflow/Astronomer, BigQuery, and Cloud Storage. Tune performance and cost through query optimisation, partitioning, and resource configuration. Tune streaming configurations for throughput and reliability. - Reliability & incident response: Own and track SLOs for freshness, success, and latency. Implement monitoring, alerting, and data quality checks. Participate in on-call as a Tier 2 responder, troubleshooting and resolving incidents independently.
- Experience with workflow orchestration such as Airflow: scheduling, dependencies, retry logic, and workflow coordination - Working fluency with a major cloud platform, ideally GCP (BigQuery, Dataproc, Cloud Storage, Pub/Sub) - Solid engineering and operational practice: Git, CI/CD, automated testing, code review, monitoring, alerting, and incident response including on-call
- Functional Scala patterns (cats-effect and the Typelevel ecosystem), or JVM tuning exposure such as garbage collection, memory management, and thread pools - Observability tooling such as Grafana, Datadog, or Cloud Monitoring - Experience with Dataflow (Apache Beam), SingleStore, or BigTable
Impact's Data Platform Engineering team is looking for a Data Platform Engineer ready to independently own production pipelines and platform components at TB-scale, delivering data reliably, securely, and cost-effectively as the business grows. You'll design and build scalable batch and streaming pipelines, own the reliability, performance, and SLOs of your domain, and participate fully in on-call as a Tier 2 responder. You'll work within the Data Analytics Group (DAG), partnering closely with Analytics Engineers, Data Analysts, Data Scientists, and Data Product Managers, though you won't own dbt models, metric definitions, or reporting. We operate with a platform-as-a-product mindset: the platform is an internal product with clear interfaces, paved roads, and strong developer experience, and success is measured by the productivity of the teams who depend on it. Our ideal candidate combines solid distributed systems engineering with operational maturity, clear communication, and a habit of mentoring those earlier in their careers. The role reports to the Team Lead, Data Platform Engineering, and is based in Cape Town, hybrid, with two days per week in office. The platform:
- Experience migrating legacy big-data platforms onto cloud-native alternatives - Familiarity with dbt, Looker, or semantic layer concepts, and with dbt orchestration through Astronomer Cosmos - Functional Scala patterns (cats-effect and the Typelevel ecosystem), or JVM tuning exposure such as garbage collection, memory management, and thread pools
Nice to Have (Advantageous) Requirements: - Infrastructure as code (Terraform or equivalent) and automated deployment - Experience migrating legacy big-data platforms onto cloud-native alternatives
- Advanced SQL, including complex queries, window functions, CTEs, and query optimisation - Experience with workflow orchestration such as Airflow: scheduling, dependencies, retry logic, and workflow coordination - Working fluency with a major cloud platform, ideally GCP (BigQuery, Dataproc, Cloud Storage, Pub/Sub)
Benefits in the posting
Full benefits wording- Flexible Working: Our Responsible PTO policy means you can take the time off you need to rest and recharge. We're committed to a positive work-life balance and provide a flexible environment that allows you to be happy and fulfilled in both your career and your personal life.
- Health and Wellness: Your well-being is a priority. Our mental health and wellness benefit includes up to 12 fully covered therapy/coaching sessions per year, with additional dependent coverage. We also offer a monthly gym reimbursement policy to support your physical health.
- Parental Support: We offer a generous parental leave policy, 26 weeks of fully paid leave for the primary caregiver and 13 weeks fully paid leave for the secondary caregiver.
- Technology Financial Support: We provide a technology stipend to help you set up your home office and a monthly allowance to cover your internet expenses
- impact.com is proud to be an equal opportunity workplace. All employees and applicants for employment shall be given fair treatment and equal employment opportunity regardless of their race, ethnicity or ancestry, color or caste, religion or belief, age, sex (including gender identity, gender reassignment, sexual orientation, pregnancy/maternity), national origin, weight, neurodivergence, disability, marital and civil partnership status, caregiving status, veteran status, genetic information, political affiliation, or other prohibited non-merit factors.
From the employer’s posting.
About Impact.com
impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer journey.
In the employer’s words · Read in context
Job description
About impact.com
impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer journey. From affiliates and influencers to content publishers, brand ambassadors, and customer advocates, impact.com empowers brands to drive trusted, performance-based growth through authentic relationships. Its award-winning products—Performance (affiliate), Creator (influencer), and Advocate (customer referral)—unify every type of partner into one integrated platform. As consumers increasingly rely on recommendations from people and communities they trust, impact.com helps brands show up where it matters most. Today, over 5,000 global brands, including Walmart, Uber, Shopify, Lenovo, L’Oréal, and Fanatics, rely on impact.com to power more than 225,000 partnerships that deliver measurable business results.
Your Role at Impact:
Impact's Data Platform Engineering team is looking for a Data Platform Engineer ready to independently own production pipelines and platform components at TB-scale, delivering data reliably, securely, and cost-effectively as the business grows.
You'll design and build scalable batch and streaming pipelines, own the reliability, performance, and SLOs of your domain, and participate fully in on-call as a Tier 2 responder. You'll work within the Data Analytics Group (DAG), partnering closely with Analytics Engineers, Data Analysts, Data Scientists, and Data Product Managers, though you won't own dbt models, metric definitions, or reporting. We operate with a platform-as-a-product mindset: the platform is an internal product with clear interfaces, paved roads, and strong developer experience, and success is measured by the productivity of the teams who depend on it. Our ideal candidate combines solid distributed systems engineering with operational maturity, clear communication, and a habit of mentoring those earlier in their careers. The role reports to the Team Lead, Data Platform Engineering, and is based in Cape Town, hybrid, with two days per week in office.
The platform:
The platform has strong foundations: Scala and Spark pipelines processing TB-scale data, dozens of managed data connectors, Airflow on Astronomer for orchestration, a mature dbt transformation layer, and BigQuery as the core analytical warehouse. You'll also have room to shape the platform's next chapter: modernising parts of the compute stack, and building out governance frameworks, data contracts, SLOs, observability, cost management, and data quality monitoring within your domain.
What You'll Do:
- End-to-end domain ownership: Own platform components and a pipeline domain end-to-end, including design, implementation, reliability, performance, and cost. Design and build scalable ETL/ELT and streaming pipelines using Spark/Dataproc, Kafka, and Pub/Sub. Integrate diverse sources into BigQuery and other stores. Build modular, fault-tolerant components with graceful schema evolution.
- Infrastructure & performance tuning: Configure and maintain core platform infrastructure as code, including Dataproc, Kafka, Airflow/Astronomer, BigQuery, and Cloud Storage. Tune performance and cost through query optimisation, partitioning, and resource configuration. Tune streaming configurations for throughput and reliability.
- Reliability & incident response: Own and track SLOs for freshness, success, and latency. Implement monitoring, alerting, and data quality checks. Participate in on-call as a Tier 2 responder, troubleshooting and resolving incidents independently.
- Security, governance & self-service: Implement security and governance by design, including access control, secrets, encryption, audit logging, lineage, and retention. Build self-service capabilities and paved roads for analytics engineers, analysts, and data scientists.
- Engineering craft & documentation: Write clean, well-tested code with CI/CD and current runbooks and architecture docs. Use AI coding assistants while holding the same quality and review standards.
- Mentorship & communication: Mentor Associate Data Platform Engineers through pairing and code review. Communicate platform trade-offs clearly and contribute to platform standards and data contracts.
What You Have:
Preferred Requirements:
- 3 to 5 years of experience in data platform, data engineering, or backend/distributed systems engineering, including ownership of production systems at TB-scale, or systems carrying stringent SLAs, multiple downstream consumers, or compliance requirements
- Strong Python, plus production experience in a JVM language, with the ability to write production-quality code with proper error handling, logging, and testing. Scala is our primary pipeline language: if you don't have it yet, you'll need a real appetite to become strong in it
- Hands-on depth in either distributed batch processing (Spark: DataFrames, Spark SQL, partitioning, shuffles, fault tolerance) or streaming (Kafka or Pub/Sub: producers, consumers, topics, partitions, stream processing patterns), with working familiarity of the other
- Advanced SQL, including complex queries, window functions, CTEs, and query optimisation
- Experience with workflow orchestration such as Airflow: scheduling, dependencies, retry logic, and workflow coordination
- Working fluency with a major cloud platform, ideally GCP (BigQuery, Dataproc, Cloud Storage, Pub/Sub)
- Solid engineering and operational practice: Git, CI/CD, automated testing, code review, monitoring, alerting, and incident response including on-call
- Security and governance practice in data systems: access control, secrets management, encryption, and least-privilege design
- Clear written and verbal communication for a distributed team, plus some experience or aptitude for mentoring less experienced engineers
- Infrastructure as code (Terraform or equivalent) and automated deployment
- Experience migrating legacy big-data platforms onto cloud-native alternatives
- Familiarity with dbt, Looker, or semantic layer concepts, and with dbt orchestration through Astronomer Cosmos
- Functional Scala patterns (cats-effect and the Typelevel ecosystem), or JVM tuning exposure such as garbage collection, memory management, and thread pools
- Observability tooling such as Grafana, Datadog, or Cloud Monitoring
- Experience with Dataflow (Apache Beam), SingleStore, or BigTable
- Bachelor's degree in Computer Science, Engineering, Mathematics, or a related field is a plus
- Experience in the digital marketing technology industry is a plus
Benefits and Perks:
At impact.com, we believe that when you’re happy and fulfilled, you do your best work. That’s why we’ve built a benefits package that supports your well-being, growth, and work-life balance.
- Flexible Working: Our Responsible PTO policy means you can take the time off you need to rest and recharge. We're committed to a positive work-life balance and provide a flexible environment that allows you to be happy and fulfilled in both your career and your personal life.
- Health and Wellness: Your well-being is a priority. Our mental health and wellness benefit includes up to 12 fully covered therapy/coaching sessions per year, with additional dependent coverage. We also offer a monthly gym reimbursement policy to support your physical health.
- A Stake in Our Growth: We offer Restricted Stock Units (RSUs) as part of our total compensation, giving you a stake in the company's growth with a 3-year vesting schedule, pending Board approval.
- Investing in Your Growth: We’re committed to your continuous learning. Take advantage of our free Coursera subscription and our PXA courses.
- Parental Support: We offer a generous parental leave policy, 26 weeks of fully paid leave for the primary caregiver and 13 weeks fully paid leave for the secondary caregiver.
- Technology Financial Support: We provide a technology stipend to help you set up your home office and a monthly allowance to cover your internet expenses
impact.com is proud to be an equal opportunity workplace. All employees and applicants for employment shall be given fair treatment and equal employment opportunity regardless of their race, ethnicity or ancestry, color or caste, religion or belief, age, sex (including gender identity, gender reassignment, sexual orientation, pregnancy/maternity), national origin, weight, neurodivergence, disability, marital and civil partnership status, caregiving status, veteran status, genetic information, political affiliation, or other prohibited non-merit factors.
#LI-CT1
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
- Ask the employer about the salary range before committing time to the process.
Complete your application on job-boards.greenhouse.io. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
No pay amount identified in the saved description.
- Location & working pattern
Cape Town
Impact's Data Platform Engineering team is looking for a Data Platform Engineer ready to independently own production pipelines and platform components at TB-scale, delivering data reliably, securely, and cost-effectively as the business grows. You'll design and build scalable batch and streaming pipelines, own the reliability, performance, and SLOs of your domain, and participate fully in on-call as a Tier 2 responder. You'll work within the Data Analytics Group (DAG), partnering closely with Analytics Engineers, Data Analysts, Data Scientists, and Data Product Managers, though you won't own dbt models, metric definitions, or reporting. We operate with a platform-as-a-product mindset: the platform is an internal product with clear interfaces, paved roads, and strong developer experience, and success is measured by the productivity of the teams who depend on it. Our ideal candidate combines solid distributed systems engineering with operational maturity, clear communication, and a habit of mentoring those earlier in their careers. The role reports to the Team Lead, Data Platform Engineering, and is based in Cape Town, hybrid, with two days per week in office. The platform:
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Oct 7, 2026
- Recorded sightings
- 1
- Employer says posted
- Oct 5, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.