Senior Principal Site Reliability Engineer
Core Responsibilities
Chaos Engineering Platform Architecture & Development (50%)
Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (K8s + EC2 hybrid), multi-region, and multi-environment (testnet/mainnet) deployments
Core capability development:
Fault Injection Engine: Pod-level / Node-level / AZ-level fault simulation, network latency / packet loss / partition, dependency timeout / error injection
Production Safety Assurance: Blast radius control, one-click Kill Switch, automatic rollback, real-time impact monitoring
Traffic Isolation: Experiment traffic tagging and isolation to ensure fault injection does not impact real users
Fault Isolation: Precise impact scoping at service / cluster / AZ granularity
Design experiment orchestration capabilities supporting complex fault scenario composition (e.g., simultaneous network latency + downstream timeout + cache invalidation)
Deep integration with existing monitoring, alerting, and SLO systems to achieve an automated closed loop: inject fault → observe impact → determine pass/fail
Production Resilience Validation Framework (30%)
Define safety standards and approval workflows for mainnet fault injection
Design and drive routine chaos experiments:
Daily patrol-level experiments: Low-risk experiments executed automatically on a daily/weekly basis
Periodic validation experiments: Monthly/quarterly resilience verification of critical paths
Large-scale drills: Cross-AZ / cross-region disaster recovery failover validation
Establish a resilience scoring system to quantify system health based on experiment results
Deliver improvement recommendations and drive business teams to remediate identified weaknesses
3. Technology Selection & Team Enablement (20%)
Evaluate and select the technology foundation (Chaos Mesh / Litmus / custom components — hybrid strategy)
Develop chaos engineering best practices and playbooks to enable SRE teams and application developers
Mentor and grow the team (2–3 engineers) in chaos engineering capabilities
Stay current with industry developments and introduce cutting-edge practices (e.g., AI-driven fault scenario discovery)
──────
Requirements
Must-Have:
8+ years of backend / infrastructure engineering experience, with 3+ years dedicated to chaos engineering or stability engineering
Hands-on experience with large-scale fault injection in production environments (not just test environments), with deep understanding of production safety constraints
Expert-level proficiency in Kubernetes fault injection (Chaos Mesh / Litmus / custom solutions), familiar with CRD / Operator development
Proficient in at least one backend language (Go preferred), with platform-level system architecture design capability
Deep understanding of distributed system failure modes (network partitions, split-brain, cascading failures, data inconsistency, etc.)
Familiarity with observability tech stack (Prometheus / Grafana / Thanos / OpenTelemetry)
Excellent technical documentation and solution design skills
Nice-to-Have:
Experience in financial / trading system stability (understanding of transaction consistency and fund safety constraints)
Experience building SLO / Error Budget frameworks
Experience building automated fault recovery (self-healing) systems
Familiarity with AWS infrastructure (EC2 / EKS / Multi-AZ / Multi-Region)
Knowledge of Netflix Chaos Engineering / AWS Fault Injection Simulator / Gremlin
Open-source community contributions (Chaos Mesh / Litmus or similar projects)
Soft Skills:
Ability to balance "safety" and "validation depth" — not afraid of production injection, while maintaining strict risk control
Strong cross-team collaboration and influence — chaos engineering requires buy-in from business teams; this role demands persuasion skills
Self-driven, capable of independently planning and executing in ambiguous situations