Software Engineer, Compute Infrastructure

Render · Remote: United States · $170K – $290K · Engineering

Posted 2026-06-30

Apply for this role →

ABOUT THE ROLE

Render's mission is to eliminate the undifferentiated work that goes into building software products by offering an easy-to-use, powerful cloud platform for developer teams of all sizes.

We are scaling rapidly. Our customers have created millions of services on our platform, and the numbers continue to accelerate. Our customers trust us to deliver a secure, reliable and performant cloud — this is our top priority as a company.

By joining us at an early stage, you will design and build the cloud platform you've always wanted for yourself, and make decisions that will shape our product and company, directly impacting developers around the globe.

Render builds and orchestrates a growing number of kubernetes clusters on different hyperscalers and recently, our own hardware. We leverage open source and industry-standard technologies, modifying and extending the core components to meet the unique reliability, performance, and scaling demands of our platform.

We are looking for engineers with deep specialization in cloud and compute infrastructure. We are particularly interested in experience with Kubernetes and container orchestration, micro VMs, controllers, operators, and the automation and self-healing of large scale distributed systems.

Areas of focus for the team this year will unlocking bottlenecks to scale our clusters while also automating the multi-step creation, configuration, testing, and tuning of Render clusters. We will expand Render to new regions, our own bare metal, and start orchestrating micro VMs outside of Kubernetes entirely.

WHAT YOU'LL DO

- Own Render's core compute infrastructure across multiple cloud providers, regions, and data centers. You'll shape how we evolve our compute platform as we rapidly scale.

- Design and build capabilities that give users greater performance and flexibility in how their services are built, deployed, perform, and stay available even when underlying resources go down.

- Investigate challenging cloud and compute issues across the stack, from the kernel and data plane to our kubernetes cluster, control plane, and other orchestration mechanisms.

- Improve the performance and reliability of our infrastructure through systematic profiling, experimentation, and tuning.

- Partner with engineers across the company to build a platform that is stable, predictable, and secure.

- Participate in our on-call rotation. Help continuously improve how we detect, respond to, and learn from incidents.

WHAT WE'RE LOOKING FOR

- At least 7 years of experience building and operating large-scale platform or compute infrastructure.

- Deep expertise in operating, scaling, and enhancing Kubernetes clusters or similar resource/container orchestration system.

- Experience developing in Go, Rust, or similar languages to develop custom infrastructure components, scheduling, controllers, that apply business logic to resource management.

- Comfort going broad and deep in a complex systems, making tradeoffs to improve performance and efficiency without sacrificing reliability.

- Strong experience designing, debugging, and operating distributed systems.

- Experience planning and executing rapid, high-risk upgrades, changes with minimal downtime to user services.

NICE-TO-HAVES

- Background in virtualization technologies like Firecracker, gVisor, Kata, or similar.

- Experience optimizing the performance of node, pod, and container spin-up times.

- Familiarity with eBPF, Linux kernel internals, resource management.

- Comfort securing and isolating workloads in multi-tenant execution environments.

RELEVANT LINKS FROM THE RENDER INFRASTRUCTURE TEAM

- The Namespaces Scaling Trap https://www.youtube.com/watch?v=NCKhoYsNsmE (video podcast interview) and related article How We Found 7 TiB of Memory Just Sitting Around https://render.com/blog/how-we-found-7-tib-of-memory-just-sitting-around

- SEV0 SF 2025 | Boundary Cases: Technical and social challenges in cross-system debugging https://www.youtube.com/watch?v=d4hAyzmBNXk

- Distributing Global State to Serve over 1 Billion Daily Requests https://render.com/blog/distributing-global-state-to-serve-over-1-billion-daily-requests

- Breaking down OpenAI's outage: How to avoid a hidden DNS dependency in Kubernetes https://render.com/blog/a-hidden-dns-dependency-in-kubernetes

- Kubernetes Informers are so easy... to misuse! https://render.com/blog/kubernetes-informers

Apply for this role →

← Back to all jobs