Data Platform Engineer – Data Operations (all genders)
Munich, Berlin
- Pay
- Salary not listed in the saved posting
- Work setup
- Unconfirmed
- Employment
- Unconfirmed
What you’ll work on
Full postingCreate Data Tooling: Build the programmatic tools that bridge raw data and downstream usage.
You will build the operational tooling and coordinate with external labeling subcontractors to ensure high-quality data deliveries, track progress, and run automated QA.
Drive Engineering Excellence: Establish and enforce good software engineering hygiene in a young codebase: rigorous testing, typing, documentation, logging, and CI/CD pipelines.
From the employer’s posting
Build the Data Backbone: Design, write, and maintain robust Python software and ETL pipelines that ingest, process, and merge complex field recordings (flight data, rosbags, video) and external deliveries into our GCP cloud storage. Create Data Tooling: Build the programmatic tools that bridge raw data and downstream usage. This includes writing services to automatically sub-sample video feeds, extract valuable frames, package datasets, and build self-serve data access tools. Database & Metadata Engineering: Design, implement, and maintain our metadata databases and data catalogs (tracking datasets, recordings, sensors, labels, and lineage) using solid SQL and schema design principles.
Database & Metadata Engineering: Design, implement, and maintain our metadata databases and data catalogs (tracking datasets, recordings, sensors, labels, and lineage) using solid SQL and schema design principles. Data Operations & Labeling Workflows: Own the end-to-end technical workflows for data curation and labeling. You will build the operational tooling and coordinate with external labeling subcontractors to ensure high-quality data deliveries, track progress, and run automated QA. Internal Tooling & APIs: Develop backend APIs, self-service dashboards, and data-access tools that allow the entire AI and engineering organization to quickly search, filter, and understand terabytes of multi-sensor data.
Internal Tooling & APIs: Develop backend APIs, self-service dashboards, and data-access tools that allow the entire AI and engineering organization to quickly search, filter, and understand terabytes of multi-sensor data. Drive Engineering Excellence: Establish and enforce good software engineering hygiene in a young codebase: rigorous testing, typing, documentation, logging, and CI/CD pipelines. Generalist Problem Solving: Thrive in an evolving startup environment. Take on ambiguous problems, migrate legacy data, handle access management, and aggressively automate away manual support tasks.
Tools in this posting
- Python
- Shell
- SQL
- BigQuery
- Docker
- Google Cloud (GCP)
- PostgreSQL
- S3
- Google Cloud Storage
Source — Tool mentions in context
Responsibilities - Build the Data Backbone: Design, write, and maintain robust Python software and ETL pipelines that ingest, process, and merge complex field recordings (flight data, rosbags, video) and external deliveries into our GCP cloud storage. - Create Data Tooling: Build the programmatic tools that bridge raw data and downstream usage. This includes writing services to automatically sub-sample video feeds, extract valuable frames, package datasets, and build self-serve data access tools.
Qualifications - Strong Software Engineering in Python: You write clean, typed, tested, and maintainable Python code. You approach data problems with a software developer's mindset. - Data Engineering & ETL: Proven experience designing, building, and operating robust ETL/data ingestion pipelines.
- Cloud & Infrastructure: Hands-on experience with object storage (GCP/GCS, S3) and the basics of containerization (Docker) and CI/CD. - Unix/Linux Environments: Strong familiarity with Unix-based systems, shell scripting, and standard Unix tools. You are comfortable working natively in a Linux environment. - Pragmatic & Adaptable: You know how to balance a quick, scrappy fix with a long-term architectural solution, and you are highly comfortable with the changing priorities of a startup environment.
- Create Data Tooling: Build the programmatic tools that bridge raw data and downstream usage. This includes writing services to automatically sub-sample video feeds, extract valuable frames, package datasets, and build self-serve data access tools. - Database & Metadata Engineering: Design, implement, and maintain our metadata databases and data catalogs (tracking datasets, recordings, sensors, labels, and lineage) using solid SQL and schema design principles. - Data Operations & Labeling Workflows: Own the end-to-end technical workflows for data curation and labeling. You will build the operational tooling and coordinate with external labeling subcontractors to ensure high-quality data deliveries, track progress, and run automated QA.
- Data Engineering & ETL: Proven experience designing, building, and operating robust ETL/data ingestion pipelines. - Database Mastery: Solid SQL (PostgreSQL preferred) with a strong grasp of schema design, data modeling, and metadata systems. - Cloud & Infrastructure: Hands-on experience with object storage (GCP/GCS, S3) and the basics of containerization (Docker) and CI/CD.
- Data Ecosystems: Familiarity with data versioning (DVC, LakeFS, FiftyOne), computer vision annotation formats (e.g., COCO), or large-scale data curation workflows. - Advanced GCP: Experience with GCP services beyond basic storage (BigQuery, Cloud Run, IAM). - Robotics / Complex Data: Exposure to robotics data formats (ROS bags, MCAP, PX4 logs) or handling heavy, multi-modal data streams (video, lidar).
- Database Mastery: Solid SQL (PostgreSQL preferred) with a strong grasp of schema design, data modeling, and metadata systems. - Cloud & Infrastructure: Hands-on experience with object storage (GCP/GCS, S3) and the basics of containerization (Docker) and CI/CD. - Unix/Linux Environments: Strong familiarity with Unix-based systems, shell scripting, and standard Unix tools. You are comfortable working natively in a Linux environment.
About SKD Se
STARK is a new kind of defence technology company revolutionizing the way autonomous systems are deployed across multiple domains.
In the employer’s words · Read in context
Job description
STARK is a new kind of defence technology company revolutionizing the way autonomous systems are deployed across multiple domains. We design, develop and manufacture high-performance unmanned systems that are software-defined, mass-scalable, and cost-effective. This provides our operators with a decisive edge in highly contested environments.
We're focused on delivering deployable, high-performance systems — not future promises. In a time of rising threats, STARK is bolstering the technological edge of NATO Allies and their Partners to deter aggression and defend Europe — today.
About the team
The Data Operations team owns the entire data lifecycle behind STARK's AI stack: collection, acquisition, generation, curation, and management. We run our own data-collection campaigns across Europe, evaluate new sensors and platforms, and build the internal data platform that turns raw recordings into ready-to-use datasets. Everything we produce feeds directly into the perception and autonomy systems deployed on STARK's platforms — a real data advantage is built, not bought. The team is scaling up right now: real scope, direct impact, no legacy.
Your mission
Data is the fuel of STARK’s AI stack — you build the engine that makes it usable. You own the software backbone of our data platform: the metadata systems, ETL pipelines, data contracts, catalogs, databases, and internal tools that let engineers find, understand, validate, and reuse terabytes of multi-sensor field data in minutes, not days. You treat data context as a product: structured, searchable, version-aware, documented, and traceable from raw recording to processed asset, annotation delivery, dataset, and downstream ML workflow. Success in this role requires seamless cross-functional collaboration—acting as the central hub between the teams collecting data in the field, our external annotation vendors, and the ML engineers training the models. Your job is to design the automated workflows and data tooling that unify these groups, transforming raw operational data into reliable, high-quality systems at scale.
Responsibilities
Build the Data Backbone: Design, write, and maintain robust Python software and ETL pipelines that ingest, process, and merge complex field recordings (flight data, rosbags, video) and external deliveries into our GCP cloud storage.
Create Data Tooling: Build the programmatic tools that bridge raw data and downstream usage. This includes writing services to automatically sub-sample video feeds, extract valuable frames, package datasets, and build self-serve data access tools.
Database & Metadata Engineering: Design, implement, and maintain our metadata databases and data catalogs (tracking datasets, recordings, sensors, labels, and lineage) using solid SQL and schema design principles.
Data Operations & Labeling Workflows: Own the end-to-end technical workflows for data curation and labeling. You will build the operational tooling and coordinate with external labeling subcontractors to ensure high-quality data deliveries, track progress, and run automated QA.
Internal Tooling & APIs: Develop backend APIs, self-service dashboards, and data-access tools that allow the entire AI and engineering organization to quickly search, filter, and understand terabytes of multi-sensor data.
Drive Engineering Excellence: Establish and enforce good software engineering hygiene in a young codebase: rigorous testing, typing, documentation, logging, and CI/CD pipelines.
Generalist Problem Solving: Thrive in an evolving startup environment. Take on ambiguous problems, migrate legacy data, handle access management, and aggressively automate away manual support tasks.
Qualifications
Strong Software Engineering in Python: You write clean, typed, tested, and maintainable Python code. You approach data problems with a software developer's mindset.
Data Engineering & ETL: Proven experience designing, building, and operating robust ETL/data ingestion pipelines.
Database Mastery: Solid SQL (PostgreSQL preferred) with a strong grasp of schema design, data modeling, and metadata systems.
Cloud & Infrastructure: Hands-on experience with object storage (GCP/GCS, S3) and the basics of containerization (Docker) and CI/CD.
Unix/Linux Environments: Strong familiarity with Unix-based systems, shell scripting, and standard Unix tools. You are comfortable working natively in a Linux environment.
Pragmatic & Adaptable: You know how to balance a quick, scrappy fix with a long-term architectural solution, and you are highly comfortable with the changing priorities of a startup environment.
Communicator & Coordinator: You are comfortable working cross-functionally and coordinating with external vendors and non-technical stakeholders to drive data labeling and curation efforts.
Automation Mindset: You aren't allergic to jumping in to do support tasks, but you have the technical chops to automate them away so you never have to do them twice.
Nice to have
GenAI / Synthetic Data: Experience with synthetic data generation or GenAI-assisted workflows (auto-labeling, data augmentation, foundation-model-based curation).
Data Ecosystems: Familiarity with data versioning (DVC, LakeFS, FiftyOne), computer vision annotation formats (e.g., COCO), or large-scale data curation workflows.
Advanced GCP: Experience with GCP services beyond basic storage (BigQuery, Cloud Run, IAM).
Robotics / Complex Data: Exposure to robotics data formats (ROS bags, MCAP, PX4 logs) or handling heavy, multi-modal data streams (video, lidar).
A note on our process: we value critical thinking, grit, and the ability to learn over a perfect checklist. If you don't hit every bullet point but you love building data systems that real engineers depend on every day, we still want to hear from you.
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
- Ask the employer about the salary range before committing time to the process.
Complete your application on stark.jobs.personio.com. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
No pay amount identified in the saved description.
- Location & working pattern
Munich, Berlin
Working pattern and location restrictions need checking in the full posting.
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Jul 24, 2026
- Recorded sightings
- 121
- Last seen by us
- Oct 2, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.