Release Engineer

Supabase · Remote · Engineering

Posted 2026-07-27

Apply for this role →

ABOUT SUPABASE

Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth.

ABOUT THE ROLE

We're looking for a Release Engineer (SRE) to join our Release Engineering team (part of EngOps) — a production-operations expert who brings an SRE mindset to how Supabase ships and runs, making deploys safe, observable, and recoverable at scale.

Release Engineering's scope has grown well beyond build-and-ship: we increasingly own the operational reliability of the systems that deploy and run Supabase. In this role you'll treat our deployment pipelines, pre-production signal, and the control plane itself as production systems — with SLOs, error budgets, and on-call ownership — and you'll be the person teams lean on when reliability is on the line.

This is not a "gatekeeper" role. You'll make the reliable path the easy path: standardising how we deploy, instrumenting what we ship, and ensuring that when something breaks, we detect it quickly and recover quickly.

WHAT YOU'LL BE RESPONSIBLE FOR

In this role, you'll:

- Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets

- Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows

- Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies)

- Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do

- Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load

- Participate in on-call, lead blameless postmortems, and turn findings into runbooks, alerting, and automation that remove toil

- Improve deployment observability and auditability — a clear record of what shipped where, when, and by whom

- Document operational procedures — break-glass paths, access models, and runbooks — so reliability knowledge isn't tribal

RELIABILITY & OPERATIONS

- Define and track SLAs, SLOs, error budgets, and DORA delivery metrics — with meaningful alerting over noise

- Ensure deployments fail fast and safely when health checks degrade

- Harden access and break-glass workflows (e.g. scoped self-service) so the right people can act in an incident without unsafe workarounds

- Partner with product engineering and platform teams to align release practices with reliability and availability targets

YOU MIGHT BE A GOOD FIT IF YOU

- Have 5+ years in SRE, production operations, platform engineering, or release engineering

- Have operated production systems at scale and carried on-call for them

- Are fluent in SLAs, SLOs, error budgets, DORA metrics, and operational KPIs — and the observability tooling behind them (Prometheus, Grafana, Alertmanager, or similar)

- Have led incident response with tooling like incident.io http://incident.io (or PagerDuty / Opsgenie), run blameless postmortems, and driven down MTTD/MTTR

- Operate confidently on AWS (multiple accounts, IAM, VPC) in production

- Are comfortable with infrastructure-as-code (Pulumi, Terraform) and Kubernetes

- Script and automate to eliminate toil rather than absorb it

- Communicate clearly with both infrastructure specialists and product engineers

- Thrive in async, globally distributed teams

- Are comfortable navigating ambiguity and iterating toward better systems over time

Apply for this role →

← Back to all jobs