Principal Infrastructure Architect - Cloud Platform

ZoomInfo Technologies LLC · Remote · Engineering

Posted 2026-08-03

Apply for this role →

As a Principal Infrastructure Architect on our Infrastructure Architecture team, you'll set technical direction that spans every part of the platform: cloud, networking, security posture, and delivery. You will join a team of architects that governs standards, leads architecture reviews, and partners with engineering teams to move the platform forward.

This is a generalist architect seat with room to lean into your specialty. The baseline is Kubernetes, cloud infrastructure, standard networking practice, GitOps, and the ability to write an ADR and defend it. Depth areas are CI/CD and delivery, AI/ML infrastructure, enterprise networking, and service mesh. We don't expect deep experience across all of them; we expect real depth in a few and the ability to dive into the others as needed.

We are optimizing for the judgment and breadth to see a pattern from prototype to adoption: setting vision, running POCs, and partnering with runtime owners to land them. Leadership and communication rather than positional authority turn direction into adopted patterns here. We value a willingness to dive in as a hands-on technical contributor when a critical initiative calls for it.

What You'll Do

Core platform responsibilities:

Runtime patterns. Design and maintain reference patterns for GKE (cluster and namespace topology, workload identity, autoscaling, upgrade strategy, multi-tenancy), shipped as Terraform modules and GitOps workflows.

Network, Identity and security baseline. Provide the standard cloud network patterns (shared VPC topology, IAM inheritance, private connectivity to managed services, load balancing, DNS, egress and cost modeling) and baseline security controls in partnership with security architects.

Delivery and reliability guardrails. Partner with platform engineering on IaC standards, policy gates, SLO practices, multi-region and DR patterns with RTO/RPO targets, capacity planning, and observability defaults.

Cost and performance. Establish patterns for bin-packing, right-sizing, and egress tracking. Partner with FinOps on runtime and networking spend, and review workloads with teams against those patterns.

Architecture review and enablement. Host design reviews, write ADRs, expand the standards library, and run structured POCs of new technologies and vendor offerings as part of the technology-lifecycle and procurement process.

Nice to to Have:

CI/CD and delivery. Multi-tenant progressive delivery (canary, blue/green at scale), cross-environment promotion, policy-as-code (OPA, Kyverno), and paved-road templates.

AI/ML infrastructure. Accelerator infrastructure on GCP (GPU and TPU node pools and scheduling on GKE), training job infrastructure (storage throughput, networking, checkpointing, spot and reservation strategy), and model serving runtimes, in partnership with the ML engineering teams that own the models and pipelines.

Enterprise networking. Multi-cloud and multi-region connectivity and the perimeter around it: VPC Service Controls, Private Service Connect and PrivateLink, interconnect and transit, DNS and routing.

IAM. Workload identity and CMEK across trust boundaries

Service mesh. Istio deployment topology, mTLS, traffic management (canary, retries, timeouts, circuit breaking), gateways, and multi-cluster mesh with its observability.

What We're Looking For

You've gotten patterns adopted without owning the teams. A track record of landing platform standards across an engineering organization through ADRs, prototypes, and benchmarks rather than mandate.

You've run the baseline at scale. Hands-on architecture and operation of Kubernetes (GKE preferred) and GCP infrastructure for many teams, with Terraform and GitOps modules other teams consumed. You know what broke in production and what pattern came out of it.

You ship code, not only diagrams. You've written the code or IaC to prove a pattern under load, and turned a production fix (over-provisioned nodes, idle accelerators, a runaway egress bill) into something reusable.

You have depth in a few of the areas above. Enough fluency in the others to evaluate and review designs across all of them.

Bonus Points

Built or applied AI-augmented developer or architecture workflows (agent-driven provisioning, MCP servers, LLM-based rubric evaluation).

Designed privacy and compliance controls (data residency, GDPR/CCPA) into infrastructure.

Run or participated in an Architecture Review Council or equivalent governance body.

Familiar with data platform infrastructure (Kafka/Confluent, BigQuery, Snowflake) as a consumer of the runtime.

Education and Experience

Bachelor's degree in Computer Science, related technical field, or equivalent practical experience.

10–15+ years of cloud infrastructure, platform engineering, or SRE experience, including several years establishing shared platform standards.

#LI- RA1

#LI-Remote

Actual compensation offered will be based on factors such as the candidate’s work location, qualifications, skills, experience and/or training. Your recruiter can share more information about the specific salary range for your desired work location during the hiring process. We want our employees and their families to thrive.

In addition to comprehensive benefits we offer holistic mind, body and lifestyle programs designed for overall well-being. Learn more about ZoomInfo benefits here.

Below is the US base salary for this position. Additional compensation such as Bonus, Commission, Equity and other benefits may also apply.

$157,500—$247,500 USD

Apply for this role →

← Back to all jobs