Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US

Tech Holding · USA, Remote · Engineering

Posted 2026-08-15

Apply for this role →

The Role:

We are looking for a hands-on Lead Site Reliability Engineer for a project based assignment to establish and continuously improve the performance, reliability, and scalability of our platform.This role will define what the platform can reliably sustain today, identify where constraints will emerge, and ensure the organization is prepared to scale before demand arrives.

This is not a traditional DevOps role or an advisory architecture position. You will work directly across application services, infrastructure, databases, networking, caching, queues, external dependencies, and operational processes to identify bottlenecks, validate system limits, and lead remediation.

You will partner closely with engineering, product, and leadership to provide clear, evidence-based answers around capacity, performance, reliability, risk, and the cost of scaling.

Key Responsibilities:

Establish performance, throughput, latency, and capacity baselines for critical customer and platform workflows

Define and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds

Instrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services

Identify system bottlenecks and lead cross-functional remediation efforts with engineering teams

Build capacity models that show what the platform can sustain, where constraints will emerge, and what additional scale will cost

Lead load, stress, soak, spike, failure, and recovery testing in representative environments

Develop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events

Drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements

Partner with Test Automation and Scalability Engineering to establish automated performance testing, regression coverage, and production release gates

Own technical readiness assessments for major pilots, partnerships, and production launches

Create operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures

Lead performance and reliability investigations during incidents and ensure lessons are incorporated into future engineering work

Make infrastructure cost, performance, and reliability tradeoffs visible to engineering and executive leadership

Recommend capacity and reliability investments before they become production constraints

Required Skills:

Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline

Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements

Deep understanding of observability, performance analysis, capacity planning, and reliability engineering

Strong hands-on experience with cloud infrastructure and production distributed systems

Deep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes

Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics

Hands-on experience performing load, stress, soak, scalability, and resilience testing

Ability to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements

Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios

Strong incident management and root-cause analysis experience

Ability to translate technical performance and reliability risks into clear business implications for senior leadership

Strong judgment around when systems genuinely require optimization versus when additional complexity is premature

Nice to have:

Experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms

Experience creating capacity-cost models and forecasting infrastructure requirements

Experience building performance and reliability gates into CI/CD pipelines

Experience preparing platforms for significant increases in traffic associated with enterprise customers or strategic partnerships

Experience leading reliability or performance initiatives that span multiple engineering teams

What Success Looks Like

Within your first several months, you will have:

Established measurable throughput, latency, and capacity baselines for critical platform journeys

Defined initial SLOs, error budgets, dashboards, alerts, and performance thresholds

Identified the platform's most significant scalability and reliability constraints and created an actionable remediation roadmap

Validated representative high-scale scenarios through load, soak, stress, failure, and recovery testing

Developed a capacity and cost model showing how the platform can support significant increases in demand

Established production-readiness criteria and clear go/no-go evidence for major launches and partnerships

Created repeatable scale-up, incident, rollback, and dependency-failure runbooks

Given leadership a clear, evidence-based understanding of the platform's current capacity envelope and future scaling requirements

Employment type:

Contract

*Applicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time

#LI-Remote

Apply for this role →

← Back to all jobs