AI Infrastructure Engineer, Sandbox Platform

Scale AI · London, UK · Engineering

Posted 2026-07-23

Apply for this role →

As a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer powering our agentic workflows, deployed across both internal and customer-managed environments. This is a role for someone who cares as much about the experience of the engineers and researchers using this system as they do about the kernel internals underneath it.

You'll combine deep systems expertise (isolation, virtualisation, performance) with an obsession for developer experience: clean APIs, clear error messages, good docs, and a client library that feels well-crafted. You'll partner closely with internal teams to understand how they use the platform, debug their issues, and shape a roadmap that balances immediate needs with long-term architecture.

You will:

Design and build the sandboxing platform, client library, and API surface for secure code execution across containerized and virtualized environments

Ensure strong isolation, security, and reproducibility of execution across user sessions and workloads

Optimise for cold-start latency, memory footprint, and resource utilisation at scale

Drive down error rates through systematic debugging, monitoring, and proactive fixes

Partner closely with internal teams using the platform to understand their needs, debug issues, and build tooling that serves their use cases

Respond to incidents and production issues with urgency, conducting root cause analysis and implementing preventive fixes

Help develop and maintain a product roadmap for sandboxing, balancing immediate needs against long-term architectural investment

Lead architecture reviews and own projects end-to-end, from design through deployment, in fast-paced cross-functional settings

Ideally you'd have:

4+ years of experience building high-performance systems software, with meaningful time spent maintaining libraries, SDKs, or developer-facing APIs

Deep understanding of Linux internals: process isolation, memory management, cgroups, namespaces, etc.

Experience with containerisation and virtualisation technologies (e.g., Docker, Firecracker, gVisor, QEMU, Kata Containers)

Proficiency in a systems programming language such as Go, Rust, or C/C++

A track record of obsessing over developer experience — API design, error propagation, documentation, and the small details that make a library feel well-crafted

Comfort working across infrastructure layers, from kernel modules to orchestration frameworks (e.g., Kubernetes)

Strong debugging skills and the ability to navigate performance/security tradeoffs in production systems

Comfort with ambiguity, and the ability to context-switch between reactive incident work and proactive product development

Nice to haves:

Experience as a founder or early engineer at an infrastructure-focused startup, owning a product end-to-end

Familiarity with LLM agents and agent frameworks (e.g., OpenHands, Agent2Agent, MCP)

Experience running secure workloads in multi-tenant or untrusted environments (e.g., FaaS, CI sandboxes, remote notebooks)

Exposure to snapshotting and restore techniques (e.g., CRIU, VM snapshots, overlays)

Open-source contributions to systems or developer-tools projects

History of on-call/incident response for production systems

Apply for this role →

← Back to all jobs