Senior Production Engineer - Realtime Products

Databricks · United States · Engineering

Posted 2026-10-10

Apply for this role →

CSQ127R192

As a Senior Production Engineer working on Databricks’ realtime products you will be directly contributing to our customers’ success. You will build advanced monitoring and incident mitigation tooling, and drive changes across the stack to proactively make the realtime products like Lakebase, Neon and Model Serving reliable, secure, and scalable in production.

Our production engineers understand the Databricks platform from end to end, and partner across engineering teams, building durable solutions for a platform operating across AWS, Azure, and GCP.

The Impact You’ll Have

Observability

Build advanced monitoring and detection capabilities that give early warning of issues and deep insights into workload performance and customer experience.

Build reliable, observable automation for debugging and incident mitigation.

Reliability

Improve service reliability, scalability, security, and operational efficiency.

Develop dependable, safe, mitigations for production issues.

On-Call & Incident Response

Participate in a follow-the-sun on-call rotation and lead incident response and mitigation.

Perform root-cause analysis, identify and implement lasting corrective actions.

Partner with Product Engineering, Security, Support, and other infrastructure teams on follow up actions.

What We Look For

Experience

5+ years of experience in Production Engineering, Customer Reliability Engineering (CRE), Site Reliability Engineering (SRE), infrastructure engineering, backend software engineering, or a related field.

Experience in holistic monitoring and alerting of complex stateful systems, applying multiple strategies like workload alerting, anomaly detection and probing.

Experience with PostgreSQL, or related managed databases or distributed systems.

Experience in incident management and participating in on-call rotations for critical infrastructure.

Skillset

A mindset focused on automation, root-cause resolution, and continuous improvement.

Strong programming skills in one or more languages such as Python, Go, Java, Scala, or similar.

Proficiency with infrastructure automation and Infrastructure as Code.

Ability to work across system boundaries and collaborate effectively during complex incidents, and engage with customers’ infrastructure teams.

Experience in using AI to address production and operational challenges.

Bonus

Experience with AWS, Azure, or GCP.

Experience with Lakebase or Neon.

Experience with Kubernetes, Terraform.

Experience building internal platforms, operational tooling, or developer productivity systems.

Education

BS degree (or higher) in Computer Science, or a related field.

Pay Range Transparency

Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles.  Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here.

Zone 1 Pay Range

$164,200—$225,700 USD

Apply for this role →

← Back to all jobs