Senior Runtime Engineer
About the role
The Runtime Team builds a high-performance, distributed and scalable software execution environment for SambaNova SambaStack and Cloud platforms to support data-flow applications such as ML training and inference and HPC applications. We are searching for a software engineer who will work on all parts of the runtime stacks, supporting AI, ML, and scientific applications in high-performance distributed systems. You will participate in building, testing and deploying next-generation high-performance compute systems for AI applications at scale.
Responsibilities
div]:bg-bg-000/50 [&_pre>div]:border-0.5 [&_pre>div]:border-border-400 [&_.ignore-pre-bg>div]:bg-transparent [&_.standard-markdown_:is(p,blockquote,h1,h2,h3,h4,h5,h6)]:pl-2 [&_.standard-markdown_:is(p,blockquote,ul,ol,h1,h2,h3,h4,h5,h6)]:pr-8 [&_.progressive-markdown_:is(p,blockquote,h1,h2,h3,h4,h5,h6)]:pl-2 [&_.progressive-markdown_:is(p,blockquote,ul,ol,h1,h2,h3,h4,h5,h6)]:pr-8">
_*]:min-w-0 gap-3 [&_>_*:last-child]:mb-0 print:block print:[&_>_*_+_*]:mt-3 standard-markdown">
In this role, you'll design and implement new features across the runtime stack including:
New and enhanced features to support high-performance, scalable ML inference and training applications
Drivers and kernels for next generation silicon
Eliminate networking bottlenecks to enable high performance distributed systems
User-space libraries for high performance and high utilization of HW resources
User-facing tools (analysis, job and HW management, profiling, debugging, etc) for Datascale systems
Cross-functional collaboration including Hardware, ML Application, Compiler, and DevOps
Required Qualifications
Bachelor's or higher degree in Computer Science, Engineering, or related field (or equivalent experience)
5+ years of software engineering experience, often with emphasis on distributed systems, networking, or cloud infrastructure
Proven experience building, testing, and tuning software for distributed, high-performance systems. In-depth knowledge of user libraries, and runtime stacks
Solid understanding of Switching and Routing and ability to configure and debug switches and routers for maximum application performance
Significant experience with RDMA and RoCE networking stacks, such as RDMA based verbs and congestion management
Hands-on experience with kernel drivers and system-level software that directly interfaces with hardware
Expertise in designing and optimizing systems that handle massive parallel workloads, including machine learning training and inference tasks that involve billions of operations per second
Deep understanding of hardware-software interaction, including registers, device memory management, and the intricacies of accelerator design. Experience working with ASIC accelerators is highly desirable
Familiarity with distributed systems architecture, including networking, communication protocols, and the challenges of scaling compute resources efficiently
Hands-on experience with software development tools such as Git, Jenkins, and Jira, with an ability to drive automation and continuous integration efforts
Ability to work at the intersection of hardware and software, designing systems that optimize both performance and reliability
Preferred Qualifications
Experience designing or working closely with custom hardware accelerators (ASICs, FPGAs, etc.) and understanding low-level interactions
Knowledge of SDN, DPDK, SONiC, or modern networking frameworks and background with distributed communication libraries (e.g., UCX, MPI, NCCL).
Familiarity with deploying high-performance systems in distributed, cloud, or data center environment