Senior Software Engineer - Backend & Cloud Infrastructure

Rubrik Job Board · Palo Alto, CA · Engineering

Posted 2026-09-22

Apply for this role →

About the Team & Role:

We're building Rubrik Agent Cloud, the enterprise platform to monitor, govern, and remediate AI agents. Our team operates at the intersection of AI, distributed systems, and enterprise security, creating solutions that make it possible for organizations to operate production-grade AI agents at scale.

As a member of the Rubrik Agent Cloud team, you'll work with cutting-edge AI technologies while solving complex challenges in scaling, security, and performance. We're a collaborative team passionate about pushing the boundaries of what's possible with enterprise AI infrastructure.

Nature of the Specialized Duties

➢ Architecting and Governing Secure Multi-Cloud Foundations (30% of time)

Leading the end-to-end architectural design and onboarding of complex business units into a multi-cloud ecosystem (AWS, Azure, GCP, OCI), ensuring structural parity across disparate environments.

Engineering secure "Landing Zones" and multi-tenant structures using advanced Identity and Access Management (IAM) hierarchies and automated policy enforcement.

Developing and implementing global Role-Based Access Control (RBAC) models and Attribute-Based Access Control (ABAC) frameworks to manage granular permissions for AI agent workloads.

Designing and governing resource hierarchy models and tagging standards that utilize metadata for automated lifecycle management and complex cost-attribution algorithms.

➢ Engineering Security and Compliance at Scale for AI Workloads (25% of time)

Designing and implementing cryptographic controls, including Key Management Service (KMS) integration, encryption-at-rest/transit protocols, and hardware security module (HSM) configurations.

Architecting network security perimeters using VPC Service Controls, Privileged Identity Management (PIM), and deep packet inspection logging.

Conducting technical root-cause analysis of security vulnerabilities and engineering automated remediation scripts to close compliance gaps (SOC 2, ISO 27001, FedRAMP).

Designing "Zero Trust" networking architectures utilizing Service Mesh (Istio/Envoy) and eBPF for deep observability into AI agent inter-process communication.

➢ Developing Infrastructure-as-Code (IaC) and Automation Pipelines (20% of time)

Authoring complex, reusable Infrastructure-as-Code (IaC) templates using Terraform, CloudFormation, and Pulumi to ensure deterministic environment provisioning.

Building and managing sophisticated CI/CD pipelines using GitHub Actions, FluxCD, or ArgoCD to implement GitOps workflows for infrastructure state management.

Developing custom automation tooling in Python, Go, or Rust to bridge gaps between cloud-native services and Rubrik’s proprietary Agent Cloud platform.

Engineering automated "drift detection" systems to ensure production environments remain in sync with version-controlled architecture definitions.

➢ Cloud Financial Engineering and Performance Optimization (15% of time)

Designing and implementing automated cost-optimization engines that leverage machine learning to perform resource rightsizing, storage tiering, and idle resource elimination.

Conducting deep-dive architecture reviews to identify performance bottlenecks at the infrastructure layer, optimizing for low-latency AI inference and data throughput.

Modeling cloud consumption patterns to forecast infrastructure scaling requirements and engineering auto-scaling logic based on real-time telemetry.

➢ Technical Leadership and Cross-Functional Systems Integration (10% of time)

Leading collaboration with Backend Engineers and AI Researchers to design "environment-agnostic" solutions that function across on-prem, customer-managed cloud, and sovereign cloud environments.

Mentoring junior and mid-level engineers on complex systems design, cloud-native patterns, and advanced troubleshooting of distributed networking issues.

Redesigning core infrastructure components to simplify customer onboarding while strictly adhering to external security requirements (e.g., FedRAMP operational controls).

Minimum Requirements for the Position

Education: A Bachelor’s degree (or higher) in Computer Science, Computer Engineering, Information Systems, or a closely related technical field. A Master’s degree is preferred given the seniority and the requirement for advanced knowledge in distributed systems and cloud security.

Specialized Experience:

7+ years of progressive experience in Cloud Engineering, Systems Architecture, or Site Reliability Engineering (SRE).

Expert-level proficiency in Infrastructure-as-Code (Terraform) and at least one high-level programming language (Python or Go).

Deep technical expertise in Multi-Cloud IAM, Network Virtualization, and Cloud Security governance.

Proven experience with Container Orchestration (Kubernetes) and modern observability/networking stacks (Service Mesh, eBPF).

Demonstrated ability to engineer solutions for highly regulated environments (FedRAMP, SOC2).

The minimum and maximum base salaries for this role are posted below; additionally, the role is eligible for bonus potential, equity and benefits. The range displayed reflects the minimum and maximum target for new hire salaries for the role based on U.S. location. Within the range, the salary offered will be determined by work location and additional factors, including job-related skills, experience, and relevant education or training.

US Pay Range

$188,500—$282,700 USD

Join Us in Securing and Accelerating the World's AI Transformation

Apply for this role →

← Back to all jobs