Staff Software Engineer - Curated Data
ABOUT THE ROLE
Data Products builds and owns datasets end to end: from raw chain data through decoding to the 3000+ models and 4 petabytes we curate, share directly with customers, and replicate into their warehouses.
The role will focus on the lifecycle of building high quality data: orchestrating thousands of interdependent models, propagating schema changes without breaking downstream consumers, propagating corrections.
That is a software architecture problem in a data domain. This role is a hybrid: a backend engineer who thinks in systems and contracts, working on data.
You will be the engineer we hand ambiguous product requirements to, and will come back with a design, a sequence, and work the team can pick up, while building the hardest parts yourself.
IN THIS ROLE YOU WILL
- Design and build the control plane for our curated data lifecycle: dependency-aware orchestration, backfills, restatements, retries, partial failure, and recovery
- Decide, dataset by dataset, whether the answer is a model, a service or a job, and own that architecture through production
- Design the contracts between ingestion and curation so a dataset can be reasoned about end to end
- Build alerting and data quality signals that catch real problems and stay quiet otherwise, so on-call is about incidents rather than noise
- Work across Go, Kotlin, Rust, Python and SQL, choosing the right tool rather than the familiar one
- Break large problems into work other engineers can own, and sequence it so we ship something useful early
YOU MIGHT BE A GREAT FIT IF
- You are a backend engineer who has gone deep on data systems, or a data engineer who became a strong software engineer. You ship production services, not only pipelines
- You have built or materially extended orchestration and scheduling systems, and can explain precisely what breaks at scale and why
- You have handled schema evolution and data correctness in a system with real consumers downstream, where a breaking change has a cost
- You have built or operated stateful stream processing in production (Flink,Kafka Streams, Spark Structured Streaming, RisingWave, Materialize, Feldera)
- You have strong SQL and modeling skills on large datasets, and an interest in how the query engine underneath actually executes your work
- You have solid computer science fundamentals and distributed systems understanding
- You debug independently and drive root cause analysis to a fix that holds
- You use AI tools well enough that they have changed how you work, you understand their failure modes and dislike ai-slop.
- You communicate clearly in writing and get the best out of a distributed team
NOT REQUIRED, BUT A PLUS
- Deep experience with a transformation framework such as dbt or SQLMesh: specifically, having hit its limits and built beyond them
- Data lake formats such as Parquet, Iceberg or Deltalog
- Stateful stream processing in production (Flink, Kafka Streams, Spark Structured Streaming)
- Experience at a company where the data is the product