Staff Software Engineer - Test Systems & Tooling
SUMMARY
Temporal’s reliability is foundational to our value proposition, and as more customers run mission-critical workloads on Temporal, we need the ability to validate changes under production-like conditions and realistic failure modes before they reach production.
As a Staff Software Engineer on the Test Systems & Tooling team, you will lead technical initiatives that make sophisticated system-level testing practical across Engineering—spanning production-representative environments, workload replay/generation, and repeatable failure-mode testing. You’ll build platforms and tooling that enable teams to reproduce and prevent the kinds of issues that only emerge from complex interactions between workload, scale, configuration, and multi-tenant behavior.
WHAT YOU'LL DO
- Build production-like test environments: Design and evolve tooling that makes it easy for engineers to spin up test cells that resemble production topology, configuration, and behavior, in partnership with Cloud/Infrastructure.
- Enable production-representative workloads: Build capabilities for workload generation, multi-tenant testing, stress/load testing, and (where appropriate) traffic replay or replay-like systems to reproduce production behaviors in controlled environments.
- Make failure-mode testing repeatable: Create tooling for injecting and validating failure conditions (latency, dependency faults, resource exhaustion, degraded-state recovery, and region-level failure scenarios) so we test them deliberately—not just learn from production.
- Drive system-level test coverage: Identify high-impact system-level scenarios that are currently hard to test, and implement (or partner with the right teams to implement) repeatable coverage that’s incorporated into continuous and/or release validation.
- Build shared frameworks and paved paths: Provide opinionated libraries, harnesses, and patterns that make it easier for product teams to write high-quality system and integration tests consistently.
- Raise the technical bar through leadership: Mentor engineers and influence testing strategy across the org through design docs, code reviews, and cross-team collaboration.
- Partner across Engineering: Work closely with Reliability and Release Engineering to turn incident learnings into durable pre-production test coverage and to integrate validations into the release pipeline.
WHAT YOU'LL BRING
- Strong software engineering fundamentals and experience building and operating production-quality systems (platforms, infrastructure, developer tooling, or distributed systems).
- Demonstrated technical leadership: leading ambiguous, cross-team initiatives; driving alignment; and delivering durable systems that other teams rely on.
- Experience designing for observability, debuggability, and operational readiness (metrics, logs, tracing, safe rollouts, and failure analysis).
- Comfort working across organizational boundaries—collaborating with Cloud/Infrastructure, Reliability, Release Engineering, and product teams to land outcomes.
- Strong written communication skills (design docs, testing strategy proposals, and documentation that scales adoption).
NICE TO HAVES
- Experience with chaos engineering, fault injection, workload replay/shadowing, performance testing, or multi-tenant test strategy.
- Experience building internal platforms used broadly across an engineering org (paved roads, self-serve environments, shared frameworks).
- Familiarity with measuring engineering/system outcomes (e.g., signal quality, coverage gaps closed, incident-to-coverage time, adoption).
COMPENSATION
- The estimated pay range for this role is $212,000 - $278,250 depending on experience and location.
- Additionally, this role is eligible to participate in Temporal's equity plan.