Research Engineer - Web Crawlers
ABOUT THE ROLE
We are looking for a Research Engineer to join the research team at ElevenLabs, focused on large-scale web crawling for our frontier AI models. The quality of our models is bounded by the quality and scale of the data behind them, and you will own the crawling systems that source world-class data from the open web. You will thrive in this role if you enjoy:
- Building and operating large-scale, distributed web crawlers that discover, fetch, and extract data across billions of pages reliably and efficiently.
- Solving hard crawling problems such as content extraction from messy HTML, deduplication at web scale, freshness and recrawl strategies, and politeness and rate-limit handling.
- Designing targeted crawling pipelines that find high-value data sources, including audio, video, and multilingual content, and turn them into clean training-ready datasets.
- Creating tooling and infrastructure that lets researchers request, monitor, and explore newly crawled web data quickly and reliably.
REQUIREMENTS
We do not require any formal certifications or degrees. Instead, we are seeking enthusiastic engineers who can showcase solving impressively hard problems with artifacts such as past projects, designs, or GitHub contributions. Ideally, you bring:
- Hands-on experience building and scaling web crawlers or scraping systems, ideally in support of machine learning training data.
- Strong engineering skills in distributed systems at scale (e.g., Kubernetes, queue-based architectures, or custom pipelines processing billions of documents).
- The capacity to autonomously evaluate the quality, coverage, and compliance of crawled data, and to build the tooling to measure it.
LOCATION
This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw.
#LI-Remote