Site Reliability Engineer
Cloud Operations Engineer III
Function: Engineering
Reports to: Manager, Cloud Operations
Location: Remote - United States
Position Summary
The Cloud Operations Engineer III is a senior member of the cloud operations team, responsible for the reliability, observability, performance, and operational security of our multi-product SaaS platform. This role owns our Datadog observability practice — instrumentation standards, dashboards, SLOs, monitors, and alert routing — and leads the migration off our legacy monitoring stack.
It is an engineering role, not a ticket-queue role: the expectation is that recurring operational work gets replaced with code. The Cloud Operations Engineer III participates in on-call, incident response and is measured on fewer customer-impacting incidents, faster detection and recovery, and less manual work year over year.
The ideal candidate is a proactive problem-solver who thrives in dynamic, evolving environments and works effectively across departments to address complex challenges. They have experience partnering with cross-functional teams to understand and document requirements, then translating those needs into meaningful dashboards that improve service visibility (Observability) and support informed decision-making. They are passionate about automation, process improvement, and eliminating unnecessary manual effort. They confidently propose better approaches when opportunities for improvement arise.
Key Responsibilities
Observability and Datadog Ownership
Own the Datadog platform across all products and environments, including agent lifecycle, instrumentation standards, unified service tagging, and per-cluster configuration.
Instrument services for APM and distributed tracing, log collection, and synthetic monitoring; partner with engineering teams to close instrumentation gaps in both legacy and modern codebases.
Build and maintain the dashboard, monitor, and SLO catalog; define SLIs and error budgets for critical user journeys and use them to drive prioritization with engineering and product.
Design high-signal alerting: reduce noise and duplicate alerts, tune thresholds, and ensure every alert has an owner and a runbook.
Automation and Toil Elimination
Develop and maintain automation in PowerShell, Python, and Bash for provisioning, configuration, diagnostics, remediation, and reporting.
Extend our infrastructure-as-code estate — Bicep modules, Kubernetes manifests, Helm releases, and Azure DevOps pipeline templates — so environments and regions are reproducible and drift-free.
Convert manual runbooks into automated or self-service workflows: cluster upgrades, secret and certificate rotation, tenant provisioning, data retention purges, and access provisioning.
Security, Documentation, and Mentorship
Implement and maintain platform security controls and audit-ready operational evidence: managed identities, secret and key rotation, least-privilege access, and image and dependency scanning.
Author and maintain runbooks, on-call guides, and architecture documentation, and provide technical leadership and mentorship to junior engineers on observability, automation, and incident response.
Skills and Experience Needed
Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
5-7 years of professional experience in cloud operations, site reliability, platform, or DevOps engineering for production SaaS systems.
Demonstrated hands-on depth with a modern observability platform — Datadog strongly preferred — including APM and distributed tracing, log pipelines and indexing controls, dashboards, monitors, and SLOs.
Strong scripting and automation ability in PowerShell, with the judgment to write tooling that other engineers can safely operate.
Strong knowledge of containers, container orchestration, and the Kubernetes ecosystem, including autoscaling, cluster upgrades, and diagnosing pod-level failures.
Production experience with Azure — Kubernetes Service, Azure SQL, Cosmos DB, Redis, Service Bus, Key Vault, and Entra ID — or equivalent depth in another major cloud.
Experience with infrastructure-as-code and CI/CD pipeline authoring (Bicep or Terraform; Helm or Kustomize; Azure DevOps preferred).
Proven incident response experience in a customer-facing production environment, including on-call participation and leading post-incident reviews.
Experience operating multi-region, multi-tenant systems.
Strong knowledge of platform security and operational best practices: secret and key rotation, least-privilege access, and vulnerability remediation.
Excellent problem-solving and analytical abilities, with strong written communication for runbooks, incident updates, and technical proposals.
Strong communication, and teamwork skills, including the ability to work effectively with legacy systems and their constraints.
Nice to have: experience migrating from a legacy monitoring stack to a consolidated observability platform; relevant Azure, Kubernetes, or Datadog certifications.
Competencies
Accountability
Adaptability
AI Curiosity/Innovation
Applied Learning
Business Acumen
Collaboration
Customer Focus
Dealing w/Ambiguity
Decision Making
Driving for Results
Initiating Action
Planning and Organizing
Technical/Professional Knowledge