Senior Data Engineer
The role
Nebius is looking for a Senior Data Engineer who in addition to building and owning data pipelines will also drive the design and technical leadership within the data engineering team.
This is a hands-on data engineering role, focused on designing, implementing, and maintaining reliable data flows for analytics and machine learning. Infrastructure, cloud, and Kubernetes are used only as tools to run pipelines reliably and cost-efficiently — this is not an SRE or platform engineering role.
You’re welcome to work in our offices in Tel Aviv, Israel.
Your responsibilities will include:
Core Responsibilities (Primary Focus)
Design, build, and own production-grade data pipelines using Python and SQL.
Develop stateless, idempotent pipelines that are resilient to retries, failures, and infrastructure interruptions.
Implement data transformations, validation, and data quality checks.
Optimize pipelines for performance, reliability, and cost efficiency.
Collaborate closely with Analytics, Data Science, and ML teams to deliver trusted datasets.
Supporting Infrastructure (Secondary Focus)
Orchestrate pipelines using a workflow orchestration framework (e.g., Airflow or equivalent).
Package and run data workloads using Docker and deploy them on Kubernetes.
Use autoscaling and Spot / Preemptible compute for efficient pipeline execution.
Build CI/CD automation for data pipelines.
Use Infrastructure as Code only to provision and manage the infrastructure required to run pipelines.
We expect you to have:
8+ years of experience as a Data Engineer, primarily focused on building data pipelines.
6+ years of hands-on experience with Python and SQL.
3+ years of experience running workloads on Kubernetes.
Strong understanding of stateless system design and idempotent data processing.
Experience building and operating data pipelines in cloud environments.
Experience with workflow orchestration frameworks.
Strong Linux fundamentals and production debugging skills.
Working knowledge of spoken and written English
It will be an added bonus if you have:
Experience contributing to or working extensively with open-source software.
Experience building data pipelines using Apache Spark or similar distributed processing frameworks.
Experience building data pipelines that support machine learning workflows.
Familiarity with cost-optimized data processing (e.g., Spot / Preemptible compute).
Experience with relational and non-relational data stores.
Experience working with large-scale or high-reliability data systems.
Experience collaborating with strong Data Science and ML teams.