Senior ML Infrastructure Engineer
About the team
The ML Infrastructure team builds and operates the inference stack that serves SambaNova's models on RDU accelerators, from request scheduling and caching through the public APIs in SambaStack and SambaCloud. We take inference techniques like speculative decoding, constrained decoding, and long-context serving from prototype to production, and own the accuracy infrastructure that gates every feature we ship. We work alongside the ML, compiler, runtime, and product teams, since most of what we build touches all four.
About the role
The Senior Software Engineer, ML Infrastructure will be responsible for designing, building, and operating the production-grade inference infrastructure that powers SambaNova's serving stack on our Reconfigurable Dataflow Unit (RDU) architecture. SambaNova is an inference-first company, and this role sits at the heart of that mission: turning state-of-the-art inference techniques into reliable, high-throughput, low-latency services exposed to customers through SambaStack and SambaCloud. The engineer will own end-to-end systems spanning request scheduling, advanced decoding algorithms, caching layers, API surfaces, and the accuracy infrastructure that keeps the stack trustworthy. This role partners closely with ML, compiler, runtime, and product teams to ship inference features from prototype to production.
Responsibilities
Design and productionize advanced inference techniques on RDU to optimize for performance and cost. Key areas include speculative decoding, constrained decoding, function/tool calling, prompt caching, and long-context inference.
Own SambaNova's integration with vLLM and adjacent serving frameworks, adapting them to RDU's architecture.
Own the public inference API surface exposed through SambaStack and SambaCloud.
Build and maintain the accuracy verification and regression infrastructure that gates every inference feature shipped to customers.
Partner with ML, compiler, runtime, and product teams to take inference features from prototype to production.
Contribute to technical design discussions, code reviews, and architectural decisions as a senior individual contributor.
Required Qualifications
B.S. in Computer Science, Electrical Engineering, or related field
5+ years of industry experience building and operating large-scale distributed systems, ideally in ML serving
Strong software engineering fundamentals: algorithms, data structures, concurrency, and systems design
Experience designing and maintaining production services with strict latency, throughput, and availability requirements
Working knowledge of modern LLM inference techniques and familiarity with open-source serving stacks such as vLLM, TensorRT-LLM, or SGLang
Proficiency in Python
Experience collaborating across teams to deliver complex, system-level engineering solutions
Base Salary Range:
Base Pay Range
$200,000—$275,000 USD