Principal Platform Engineer (High Availability & Disaster Recovery)
Principle Platform DR & Capacity Engineer
At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market.
What unites Anaplanners across teams and geographies is our collective commitment to our customers’ success and to our Winning Culture.
Our customers rank among the who’s who in the Fortune 50. Coca-Cola, LinkedIn, Adobe, LVMH and Bayer are just a few of the 2,400+ global companies who rely on our best-in-class platform.
Our Winning Culture is the engine that drives our teams of innovators. We champion diversity of thought and ideas, we behave like leaders regardless of title, we are committed to achieving ambitious goals, and we love celebrating our wins – big and small.
Supported by operating principles of being strategy-led, values-based and disciplined in execution, you’ll be inspired, connected, developed and rewarded here. Everything that makes you unique is welcome; join us and let’s build what’s next - together!
Your impact
As the Principal HA & DR Engineer, you will own the architectural strategy and technical execution for the platform's resilience, high availability, and disaster recovery lifecycle. You will design, build, and deliver technical solutions that support Platform availability, resiliency, and disaster response.
You will act as the principal authority for business continuity across our global infrastructure, defining the guardrails for system fault-tolerance and leading cross-functional coordination for multi-region disaster preparedness between both cloud and on prem. You will also design automated failover mechanisms, champion chaos engineering practices, and develop robust mitigation strategies to eliminate single points of failure across both cloud and on-premises infrastructure.
You will :
Drive the HA Strategy & Roadmap: Own the end-to-end design and technical execution of our global High Availability roadmap, ensuring continuous platform uptime across both cloud and on-premises environments.
Build & Implement HA Architectures: Hand-on code, configure, and engineer active-active clustering, global load balancing, and real-time database replication to completely eliminate single points of failure.
Establish the Resiliency Practice: Build and champion our foundational Resiliency Engineering framework from scratch, defining platform-wide standards for fault tolerance and self-healing.
Influence Engineering Teams: Partner with and influence cross-functional engineering squads to ensure self-healing mechanisms and automated recovery protocols are embedded directly into their core deliverables.
Bootstrap Chaos Engineering: Introduce and champion our first-ever Chaos Engineering program; you will define the strategy, establish the safe guardrails, and prepare the organization to execute its initial automated fault-injection and live-fire failover drills.
Build & Implement HA Architectures: Hand-on code, configure, and engineer active-active clustering, global load balancing, and real-time database replication to completely eliminate single points of failure.
Optimize Uptime & Manage Risks: Monitor infrastructure performance KPIs to identify potential bottlenecks, mitigate shortage or overload risks before they cause downtime, and ensure compliance with strict Customer SLA requirements.
Your skills and experience
We work with a broad range of technologies, and we don’t expect you to know everything on day one. You’ll have time to learn our tools and grow into the role. We're looking for diverse experiences to help strengthen our team.
Essential
8+ Years of HA Engineering: Extensive HA experience across cloud (AWS, GCP, or Azure) and on-premises environments.
Technical Influence: Proven ability to evangelize resiliency standards and influence cross-functional engineering teams to build self-healing features.
Failover & Replication: Hands-on expertise with real-time data replication, database clustering, and automated, zero-downtime traffic failover.
Infrastructure as Code (IaC): Advanced proficiency with Terraform or Ansible to build and replicate highly available environments programmatically.
Systems & Container Engineering: Deep technical knowledge of Linux internals and container orchestration using Kubernetes (K8s).
Hybrid Network Engineering: Hands-on experience managing complex hybrid networking topologies, including BGP routing, DNS management, Anycast, and CDNs.
Strategic Execution: Ability to translate high-level uptime requirements into a technical roadmap and personally execute the engineering work.
Desirable
Experience supporting SaaS products.
Experience with Incident Management, Post Mortems and related practices.
Knowledge of observability and monitoring best practices.
Experience operating within one or more public clouds (AWS, GCP, Azure).
Experience with configuration management, and infrastructure as code
Knowledge of observability and monitoring best practices