Staff Reliability Engineer
Okta’s TDI Network Engineering team is responsible for the global corporate network, building and supporting a high-performing, reliable network at scale. As a member of this team, you will have a direct impact on network design, deployment, and reliability, enabling our employees to work effectively from any location globally. Your role ensures the overall security and integrity of our corporate network by leveraging network security best practices, innovative products, and rigorous security validation.
Reporting to the Network Engineering Manager, this operations-focused role is distinct from core Network Engineering and Network Security, centering primarily on operational execution—including responding to alerts, maintaining service availability, and ensuring system health across our global enterprise network. You will drive the strategic reduction of systemic toil and technical debt across multiple teams, applying a systems-level perspective and leveraging deep expertise in Distributed Systems, Networking fundamentals, Infrastructure as Code, and observability to architect scalable platforms and lead technical efforts to ensure an "Always Secure. Always On." environment. You will own multi-quarter objectives and establish long-term strategies for network reliability.
What you'll be doing :
Design and Own the resilience, health and availability of our entire global corporate network domain, managing operational responsibilities such as responding to alerts, monitoring health indicators, and executing reliability projects to ensure an "Always Secure. Always On." environment.
Drive Strategic Reduction of systemic toil and technical debt across multiple teams by introducing process efficiencies, automating network operations, and building scalable self-service operational tooling.
Collaborate and Influence closely with cross-functional stakeholders—including Business Technology, Workplace, Security, and executive leaders—challenging assumptions with grace and cascading relevant information to project teams.
Make Critical Decisions and lead the resolution of complex network operations issues and alerts from a systems perspective, anticipating potential business challenges, monitoring leading indicators, and preventing future outages.
Foster Learning and Talent by defining success for the whole team, cultivating an open and transparent environment, and actively mentoring team members through the P4 level to develop their skills and operational engineering best practices.
What you'll bring to the role:
Typically requires 8+ years of related experience in a professional role with a Bachelor’s degree; or 6+ years with a Master’s degree; or 3+ years with a PhD; or equivalent experience.
Deep expertise in AWS Networking and Palo Alto Networks solutions as core required technical competencies.
Comprehensive operational experience in Distributed Systems & Networking fundamentals, including quick incident response to alerts, monitoring system health, and managing protocols such as WiFi, DNS, DHCP, VLANs, VPN, ACLs, Routing, and Firewall Policies.
Strong proficiency in core technical skills: Cloud Platforms, IaC (e.g., Terraform/Ansible), Observability tools (e.g., Prometheus/Grafana), Programming (Python/Go), and Service Reliability Management (SLOs/SLIs) for large-scale enterprise environments.
Proven track record of managing operational availability, delivering multi-quarter objectives, and executing technical projects within defined budgets and strategic VMTs (Vision, Mission, Targets).
Demonstrated ability to navigate high levels of ambiguity, establish credibility with executive stakeholders, and model resilience during major system transitions or production incidents.
Why join us:
Desirable Technical Skills: Experience with Juniper/JUNOS switching/routing, Palo Alto Networks NGFWs, and enterprise office build and construction processes is highly desired and will help you hit the ground running.
Direct Customer Impact: You will have the opportunity to showcase your strong focus on customer and technology experiences, ensuring a secure, consistent end-user experience across a global enterprise network.
Growth and Support: Join a collaborative environment that values systemic learning over personal errors, offering you the chance to eliminate manual toil through innovation, mentor peers, and occasionally travel to build impactful connections.
#P10445_3513586
#LI_Hybrid