Senior Cloud Operations Engineer
About the Role
We are seeking an experienced Senior Cloud Operations Engineer to join our Cloud Skill Community. This role owns the operational health of the Dragos customer cloud fleet across Azure, AWS, and GCP -- both Dragos-managed and customer-managed environments. You will drive fleet reliability, deployment automation, and day-to-day operations at scale, bringing a strong infrastructure-as-code mindset and a bias toward automation and repeatability.
Responsibilities
Operate, maintain, and improve the Dragos cloud fleet across Azure, AWS, and GCP
Own the full customer environment lifecycle -- onboarding, configuration, upgrades, and off boarding
Build and maintain Terraform-based infrastructure-as-code for customer environment provisioning and fleet standardization
Manage fleet health, drift detection, patching, and version lifecycle management at customer scale
Design and enforce multi-tenant isolation patterns -- blast radius containment, RBAC at scale, and cross-account access controls
Configure and maintain cloud networking components across all three providers (VPCs/VNets, peering, transit gateways, DNS, firewalls, load balancers)
Manage cloud-to-OT/on-prem connectivity for customer environments
Implement and maintain IAM, secrets management, and compliance posture across cloud providers
Build and maintain Datadog observability -- monitors, dashboards, log pipelines, and SLOs
Participate in an on-call rotation (PagerDuty) for the Dragos cloud fleet -- triage, respond to, and resolve production incidents across customer environments
Drive SRE practices: define SLOs, manage error budgets, maintain runbooks, and lead post-incident reviews
Support audit and compliance activities including evidence collection for FedRAMP, SOC2, and customer-specific requirements
Identify and eliminate toil through automation and process improvement
Qualifications
4+ years of hands-on cloud operations experience across one or more of Azure, AWS, or GCP
Cybersecurity Experience
Strong proficiency with Terraform for infrastructure-as-code (primary IaC tool at Dragos)
Experience operating cloud environments at scale -- fleet management, patching, upgrades, drift detection
Hands-on experience with multi-tenant cloud architectures and customer-facing environment management
Solid knowledge of cloud networking: VPCs/VNets, peering, transit gateways, DNS, firewalls, and hybrid connectivity
Experience with IAM across AWS, Azure Entra ID, and/or GCP -- roles, policies, federation, SSO
Proficiency with Datadog for monitoring, alerting, dashboards, and log pipelines
Comfort with on-call responsibilities -- this role participates in a PagerDuty rotation for the customer cloud fleet
Experience with SRE practices: SLO definition, error budget management, incident response, and blameless post-mortems
Strong scripting skills in Python and/or Bash
Familiarity with compliance frameworks (FedRAMP, SOC2, NIST CSF) and audit evidence collection
Strong ownership mindset and accountability for production stability
Excellent communication, documentation, and collaboration skills
Preferred Qualifications
Experience with all three major cloud providers: Azure, AWS, and GCP
Experience managing cloud environments for external customers (customer-managed and vendor-managed models)
Familiarity with cloud-to-OT/ICS network connectivity and segmentation
Experience with secrets management tooling: HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, GCP Secret Manager
Experience with cloud-native fleet tools: AWS Systems Manager, Azure Arc, GCP Fleet Management
Proficiency with policy-as-code tools (OPA, Sentinel, AWS SCPs, Azure Policy)
Experience with FinOps practices and cloud cost optimization
Familiarity with Cloud Security Posture Management (CSPM) tooling
Experience with CI/CD pipeline authoring (GitHub Actions)
Passion for automation and continuously improving operational efficiency
Compensation:
Salary: $165,000
Competitive Equity Package
Comprehensive Benefits Plan
#LI-NH1 #LI-REMOTE