Threat Intelligence Data Engineer
Pune City, India
- Pay
USD 1,300–2,000/month — pay source
Location: Remote in India. Work from wherever you please! Your home, the beach, our offices, etc. Compensation: USD 1300-2000 monthly Professional Growth: Amazing upward mobility in a rapidly expanding company.
Read the full posting- Work setup
- Unconfirmed
- Employment
Full-time — employment source
Benefits Position Type: Full-time Location: Remote in India. Work from wherever you please! Your home, the beach, our offices, etc.
Read the full posting
What you’ll work on
Full postingBuild automated collectors and archivers for anonymized and decentralized networks including:
Develop automated scanning systems to identify:
Build ETL pipelines to clean, normalize, enrich, and index structured and unstructured data
From the employer’s posting
Dark Web, Anonymized & Decentralized Networks Build automated collectors and archivers for anonymized and decentralized networks including: Tor (.onion), I2P, ZeroNet, Freenet, IPFS, GNUnet, Lokinet, Yggdrasil, and similar systems
Infrastructure & Exposure Discovery Develop automated scanning systems to identify: Unsecured databases (Elasticsearch, MySQL, PostgreSQL, MongoDB, etc.)
Pipeline Engineering & Operations Build ETL pipelines to clean, normalize, enrich, and index structured and unstructured data Implement advanced anti-bot evasion strategies (proxy rotation, fingerprinting, CAPTCHA mitigation, session management)
What you’ll bring
All qualificationsCore experience
- Strong Python expertise and experience with frameworks such as Scrapy, Playwright, Selenium, or custom async systems
- Familiarity with Tor, I2P, underground forums, stealer logs, or credential ecosystems
- Proven experience operating high-volume, automated data collection systems in production
- Experience processing large breach datasets or stealer logs
- Deep understanding of web protocols, HTTP, DOM parsing, and adversarial scraping environments
- Experience with asynchronous, concurrent, and distributed architectures
Qualification wording
Strong Python expertise and experience with frameworks such as Scrapy, Playwright, Selenium, or custom async systems
Familiarity with Tor, I2P, underground forums, stealer logs, or credential ecosystems
Proven experience operating high-volume, automated data collection systems in production
Experience processing large breach datasets or stealer logs
Deep understanding of web protocols, HTTP, DOM parsing, and adversarial scraping environments
Experience with asynchronous, concurrent, and distributed architectures
Tools in this posting
- Python
- Shell
- SQL
- AWS
- Azure
- Docker
- Elasticsearch
- Google Cloud (GCP)
- Kubernetes
- MongoDB
- MySQL
- NoSQL
- PostgreSQL
- S3
Source — Tool mentions in context
- Minimum 4 years of hands-on experience in data engineering, intelligence collection, crawling, or distributed data pipelines - Strong Python expertise and experience with frameworks such as Scrapy, Playwright, Selenium, or custom async systems - Proven experience operating high-volume, automated data collection systems in production
- Familiarity with SQL and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, Cassandra) - Strong Linux/Unix, shell scripting, and Git-based workflows - Experience deploying and operating systems using Docker, Kubernetes, AWS, or GCP
- Experience with asynchronous, concurrent, and distributed architectures - Familiarity with SQL and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, Cassandra) - Strong Linux/Unix, shell scripting, and Git-based workflows
- Strong Linux/Unix, shell scripting, and Git-based workflows - Experience deploying and operating systems using Docker, Kubernetes, AWS, or GCP - Excellent analytical, debugging, and problem-solving skills
- Unsecured databases (Elasticsearch, MySQL, PostgreSQL, MongoDB, etc.) - Exposed cloud storage (S3, Azure, GCP, DigitalOcean Spaces) - Open FTP servers, backups, and misconfigured archives
- Implement advanced anti-bot, evasion, and resiliency techniques (proxy rotation, fingerprinting, CAPTCHA mitigation, session handling) - Automate deployment, scaling, and monitoring using Docker, Kubernetes, and cloud infrastructure - Continuously optimize performance, reliability, and cost efficiency of crawler clusters
- Develop automated scanning systems to identify: - Unsecured databases (Elasticsearch, MySQL, PostgreSQL, MongoDB, etc.) - Exposed cloud storage (S3, Azure, GCP, DigitalOcean Spaces)
Job description
Our Threat Research Team’s mission is aggressive: achieve near-total coverage of global breach and leak data with 99%+ automation. Your work directly enables HEROIC’s ability to identify exposures before they are weaponized.
Architect and operate large-scale, distributed crawling and discovery systems across:
Surface web, deep web, and dark web
Hacker forums, underground marketplaces, and breach communities
Chat platforms (Telegram, Discord, IRC, WhatsApp, etc.)
Paste sites, code repositories, and social platforms used for breach disclosure
Continuously discover, archive, and download newly released datasets, logs, credentials, and artifacts the moment they appear
Build automated collectors and archivers for anonymized and decentralized networks including:
Tor (.onion), I2P, ZeroNet, Freenet, IPFS, GNUnet, Lokinet, Yggdrasil, and similar systems
Design resilient workflows for unreliable, adversarial, or ephemeral data sources
Normalize and index data from non-traditional network protocols and formats
Develop automated scanning systems to identify:
Unsecured databases (Elasticsearch, MySQL, PostgreSQL, MongoDB, etc.)
Exposed cloud storage (S3, Azure, GCP, DigitalOcean Spaces)
Open FTP servers, backups, and misconfigured archives
Monitor and ingest data from file hosting and distribution platforms commonly used for breach dumps
Build ETL pipelines to clean, normalize, enrich, and index structured and unstructured data
Implement advanced anti-bot evasion strategies (proxy rotation, fingerprinting, CAPTCHA mitigation, session management)
Integrate collected intelligence into centralized databases and search systems
Design APIs and internal tooling to support downstream analysis and AI/ML workflows
Implement advanced anti-bot, evasion, and resiliency techniques (proxy rotation, fingerprinting, CAPTCHA mitigation, session handling)
Automate deployment, scaling, and monitoring using Docker, Kubernetes, and cloud infrastructure
Continuously optimize performance, reliability, and cost efficiency of crawler clusters
Requirements
Minimum 4 years of hands-on experience in data engineering, intelligence collection, crawling, or distributed data pipelines
Strong Python expertise and experience with frameworks such as Scrapy, Playwright, Selenium, or custom async systems
Proven experience operating high-volume, automated data collection systems in production
Deep understanding of web protocols, HTTP, DOM parsing, and adversarial scraping environments
Experience with asynchronous, concurrent, and distributed architectures
Familiarity with SQL and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, Cassandra)
Strong Linux/Unix, shell scripting, and Git-based workflows
Experience deploying and operating systems using Docker, Kubernetes, AWS, or GCP
Excellent analytical, debugging, and problem-solving skills
Strong written and verbal communication skills.
Direct experience with dark web intelligence, breach data, OSINT, or threat research
Familiarity with Tor, I2P, underground forums, stealer logs, or credential ecosystems
Experience processing large breach datasets or stealer logs
Background working in adversarial data environments
Exposure to AI/ML-driven intelligence platforms
Benefits
- Position
Type: Full-time
- Location: Remote in India. Work
from wherever you please! Your home, the beach, our
offices, etc.
- Compensation: USD 1300-2000 monthly
- Professional
Growth: Amazing upward
mobility in a rapidly expanding
company.
- Innovative
Culture: Be part of a team
that leverages AI and cutting-edge
technologies.
Your next step
- Have your CV and examples of relevant work ready.
- Check the listed location, eligibility and core experience before starting.
Complete your application on careers.heroic.com. The employer’s form will show what is required.
Already applied? Track this application
Source & posting history
Source notes
Source excerptsSelected passages from the saved posting. Check the full description for conditions and exceptions.
- Pay
Location: Remote in India. Work from wherever you please! Your home, the beach, our offices, etc. Compensation: USD 1300-2000 monthly Professional Growth: Amazing upward mobility in a rapidly expanding company.
- Location & working pattern
Pune City, India
- Position Type: Full-time - Location: Remote in India. Work from wherever you please! Your home, the beach, our offices, etc. - Compensation: USD 1300-2000 monthly
- Work authorization
No clear work-authorization passage found. Eligibility is unconfirmed.
- Status in our records
- Active
- First seen by us
- Jun 2, 2026
- Recorded sightings
- 43
- Last seen by us
- Sep 30, 2026
These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.
Report an errorSee how this role fits your experience
Add your resume to compare the role’s scope, tools and requirements with your experience.