Principal Engineer, Distributed Systems

CoreWeave · New York, NY / Sunnyvale, CA · Engineering

Posted 2026-09-15

Apply for this role →

About the role

We are looking for a Principal Engineer to provide technical leadership across Security Products. This is a senior individual-contributor role for an engineer who can define architecture, guide execution across multiple teams, and solve complex distributed-systems problems in security-critical infrastructure.

The systems you help design will support demanding production requirements: four nines of availability, high scalability, consistently low latency, strong security boundaries, and safe behavior under partial failure in a multi-region setup. You will work across the full lifecycle of these systems, from architecture and technical strategy through implementation guidance, operational readiness, incident learning, and long-term evolution.

Your work will span across high-scale authorization systems, anomaly detection systems, bot-defense and security token service (STS), API authentication gateway, and the shared infrastructure required to operate these capabilities reliably across regions and deployment environments.

You will partner with engineering, security, infrastructure, networking, platform, product, and customer-facing teams. Success in this role requires both deep technical judgment and the ability to create alignment, raise engineering standards, and make complex architecture understandable and actionable for others.

What you will do

Establish architectural approaches for distributed systems - multi-region operation, including service placement, failover, replication, traffic management, disaster recovery, and regional independence.

Drive decisions around consistency models, caching, invalidation, propagation, revocation, idempotency, concurrency, and ordering where correctness and security are critical.

Design for data residency, tenant isolation, trust boundaries, blast-radius reduction, and controlled handling of sensitive security data.

Define fault-tolerance strategies for dependencies, networks, regions, storage systems, and control-plane components, including graceful degradation and safe recovery.

Design systems that meet four-nines availability goals while maintaining predictable low latency and high throughput under normal operation, traffic spikes, and partial failures.

Improve the reliability and operability of critical services through SLOs, error budgets, metrics, logs, traces, audit events, alerting, incident response, and post-incident learning.

Guide teams through architecture reviews, design reviews, implementation tradeoffs, capacity planning, load testing, performance analysis, and production readiness assessments.

Mentor senior and staff engineers, develop technical talent, and raise the quality of engineering practice across the organization.

Communicate architecture, tradeoffs, risks, and recommendations clearly to technical and executive audiences.

Lead the design of high-scale authorization systems that support policy authoring, policy evaluation, access-control enforcement, auditability, and integration across many services and tenants.

Provide technical direction for a Security Token Service, including token issuance, validation, lifecycle management, trust relationships, key rotation, revocation, and secure service-to-service access.

Guide the design and implementation of an API Authentication Gateway that provides consistent, secure, observable authentication and authorization for CoreWeave APIs.

Set technical standards for API design, service contracts, threat modeling, cryptographic key management, secrets handling, observability, and operational readiness.

Partner with service teams to make authorization and authentication capabilities easy to adopt through clear interfaces, SDKs, reference implementations, documentation, and reliable integration patterns.

What you bring

12+ years of experience designing and building production software, including substantial experience with distributed systems and platform infrastructure.

A track record of serving as a principal, staff-plus, distinguished, or equivalent technical leader across multiple teams or a broad technical domain.

Deep expertise in distributed-systems design, including scalability, availability, latency, consistency, concurrency, partition tolerance, caching, replication, and failure recovery.

Experience designing and operating security-sensitive or reliability-critical services in production.

Strong software engineering experience with one or more systems languages or backend ecosystems, such as Go, Java, Rust, C++, Python, or equivalent.

Demonstrated ability to move from ambiguous requirements to clear architecture, sequenced execution plans, and durable engineering outcomes.

Experience making architecture decisions that balance security, correctness, performance, operational simplicity, and time to value.

Strong written and verbal communication skills, including the ability to influence without direct authority and build alignment across organizational boundaries.

A track record of improving engineering standards, mentoring senior engineers, and helping teams deliver complex systems safely.

Preferred qualifications

Experience designing and implementing authorization systems at high scale, including policy engines, permission models, policy distribution, decision services, enforcement points, or authorization observability.

Experience designing or operating a Security Token Service, identity platform, credential service, or comparable security-critical control plane.

Experience designing or operating an API Authentication Gateway, service-mesh authorization layer, or centralized authentication and authorization platform.

Experience operating systems with four-nines availability requirements, high request volume, strict latency objectives, and demanding reliability expectations.

Experience with multi-region architecture, regional failover, active-active or active-passive operation, replication, disaster recovery, and regional isolation.

Experience reasoning about strong, eventual, and bounded-staleness consistency models, especially for security policy, credentials, tokens, permissions, and revocation.

Experience designing for data residency, regional data controls, tenant isolation, trust-domain separation, and compliance-sensitive deployment models.

Experience building systems that remain secure and useful during dependency failures, network partitions, regional outages, stale data, degraded capacity, or partial compromise.

Experience with authentication and identity standards such as OAuth 2.0, OpenID Connect, SAML, SCIM, JWT, mTLS, SPIFFE/SPIRE, or related protocols.

Experience with cryptography, key management, HSMs, secrets management, certificate authorities, signing keys, token validation, or secure credential lifecycle management.

Experience with Kubernetes, cloud infrastructure, service networking, distributed storage, messaging systems, and multi-cluster operations.

Experience building security and platform services with strong observability, including metrics, logs, traces, audit records, security events, and actionable alerting.

Experience working with enterprise customers, regulated workloads, or hybrid and multi-cloud environments.

How you will be successful

In your first year, you will:

Establish a clear technical north star for Security Products’ authorization, authentication, and security-infrastructure capabilities.

Build strong working relationships with engineering and product leaders, staff-plus engineers, security teams, and the service teams that depend on these platforms.

Advance the architecture and delivery plan for high-scale authorization systems, the Security Token Service, and the API Authentication Gateway.

Improve the reliability, performance, observability, and operational maturity of critical security services against four-nines availability goals.

Help teams make explicit, durable decisions about multi-region design, data consistency, data residency, tenant isolation, fault tolerance, and security boundaries.

Reduce duplicated access-control and authentication patterns by enabling consistent, well-documented platform capabilities.

Raise the quality of architecture reviews, threat modeling, performance testing, capacity planning, and production readiness across Security Products.

Mentor senior engineers and help grow a strong technical leadership community.

Why join Security Products

Security Products is building the systems that make CoreWeave trustworthy at scale. You will work on foundational infrastructure that protects access to cloud resources, enables secure product experiences, and supports customers running critical AI workloads.

This is an opportunity to shape the architecture of a modern cloud security platform while solving problems at the intersection of distributed systems, security engineering, reliability, performance, and global-scale infrastructure.

Location requirement

This position is based in Sunnyvale, California or New York, New York and requires working on-site. Remote work is not available for this role.

CoreWeave is an equal opportunity employer. We evaluate qualified applicants without regard to legally protected characteristics.

The base salary range for this role is $227,000 to $303,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

Apply for this role →

← Back to all jobs