Data Engineer
Abu Dhabi
- Pay
- Salary not listed in the saved posting
- Work setup
- Unconfirmed
- Employment
Full-time — employment source
Employment type Full-time
Read the full posting
What you’ll work on
Full postingDevelop and maintain efficient web crawling solutions, APIs, and automated workflows to continuously improve data collection processes.
Implement scalable data pipelines, ensuring efficient data processing, storage, retrieval, and distribution to research teams.
Collaborate closely with researchers and engineers to ensure collected data meets specified quality and relevance criteria.
From the employer’s posting
Rapidly collect, curate, and preprocess datasets based on detailed specifications provided by NLP researchers, delivering data within tight timelines (typically within 1-2 days). Develop and maintain efficient web crawling solutions, APIs, and automated workflows to continuously improve data collection processes. Refine and evaluate outputs from Large Language Models (LLMs) to generate structured datasets suitable for model training and benchmarking.
Refine and evaluate outputs from Large Language Models (LLMs) to generate structured datasets suitable for model training and benchmarking. Implement scalable data pipelines, ensuring efficient data processing, storage, retrieval, and distribution to research teams. Collaborate closely with researchers and engineers to ensure collected data meets specified quality and relevance criteria.
Implement scalable data pipelines, ensuring efficient data processing, storage, retrieval, and distribution to research teams. Collaborate closely with researchers and engineers to ensure collected data meets specified quality and relevance criteria. Document data collection methodologies, dataset characteristics, and pipeline architecture clearly and effectively.
Tools in this posting
- Python
- SQL
- AWS
- Kafka
- Spark
- Kubernetes
Source — Tool mentions in context
The Role As a Data Engineer specializing in Natural Language Processing (NLP) and large-scale data processing, you will quickly and effectively gather, curate, and prepare high-quality datasets to support cutting-edge NLP research. Your role will be instrumental in enabling researchers by delivering essential data through efficient and scalable engineering practices, including web crawling, LLM-generated content refinement, and robust data pipelines, primarily leveraging Python and related technologies. Key Responsibilities
Professional Experience - Required - Extensive experience in data engineering, data processing, and automation using Python. - Demonstrated proficiency in designing and deploying web crawling solutions, automated data extraction, and processing pipelines.
- Demonstrated proficiency in designing and deploying web crawling solutions, automated data extraction, and processing pipelines. - Strong understanding of data structures, algorithms, databases, SQL, and performance optimization. - Experience working with cloud infrastructure and distributed data processing frameworks (e.g., AWS, Spark, Kafka, Kubernetes).
- Strong understanding of data structures, algorithms, databases, SQL, and performance optimization. - Experience working with cloud infrastructure and distributed data processing frameworks (e.g., AWS, Spark, Kafka, Kubernetes). - Excellent problem-solving abilities, attention to detail, and the capability to rapidly address technical challenges.
About Ifm-Us
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models.
In the employer’s words · Read in context
Job description
Key Responsibilities
- Rapidly collect, curate, and preprocess datasets based on detailed specifications provided by NLP researchers, delivering data within tight timelines (typically within 1-2 days).
- Develop and maintain efficient web crawling solutions, APIs, and automated workflows to continuously improve data collection processes.
- Refine and evaluate outputs from Large Language Models (LLMs) to generate structured datasets suitable for model training and benchmarking.
- Implement scalable data pipelines, ensuring efficient data processing, storage, retrieval, and distribution to research teams.
- Collaborate closely with researchers and engineers to ensure collected data meets specified quality and relevance criteria.
- Document data collection methodologies, dataset characteristics, and pipeline architecture clearly and effectively.
- Engage with peer teams and participate in technical reviews to uphold best practices and data quality standards.
- Represent MBZUAI at industry and research forums, showcasing technical capabilities in large-scale data processing and AI data infrastructure.
- Perform all other duties as reasonably directed by the line manager commensurate with these functional objectives.
Academic Qualifications
- Bachelor's degree in Computer Science, Data Science, Engineering, or a related technical field required
- Master’s degree or equivalent experience in Computer Science, Data Engineering, or related technical fields preferred.
Professional Experience - Required
- Extensive experience in data engineering, data processing, and automation using Python.
- Demonstrated proficiency in designing and deploying web crawling solutions, automated data extraction, and processing pipelines.
- Strong understanding of data structures, algorithms, databases, SQL, and performance optimization.
- Experience working with cloud infrastructure and distributed data processing frameworks (e.g., AWS, Spark, Kafka, Kubernetes).
- Excellent problem-solving abilities, attention to detail, and the capability to rapidly address technical challenges.
- Strong communication and collaboration skills with cross-functional teams.
Professional Experience - Preferred
- Proven track record of supporting NLP or AI research teams with rapid and reliable data delivery.
- Experience with refining outputs from large-scale AI models, such as LLM-generated data.
- Contributions to open-source projects, coding competitions, or high visibility in coding communities (e.g., GitHub, Stack Overflow).
- Familiarity with the latest advancements in NLP data processing and large language model technologies.
Employment type
Full-time
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
- Ask the employer about the salary range before committing time to the process.
Complete your application on jobs.lever.co. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
No pay amount identified in the saved description.
- Location & working pattern
Abu Dhabi
Working pattern and location restrictions need checking in the full posting.
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Apr 15, 2026
- Recorded sightings
- 112
- Last seen by us
- Oct 9, 2026
- Employer says posted
- Jun 17, 2025
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.