Senior Software Engineer - ML Infrastructure
Overview
The Senior Software Engineer, ML Infrastructure will be responsible for designing, building, and operating the production-grade inference infrastructure that powers SambaNova's serving stack on our Reconfigurable Dataflow Unit (RDU) architecture. SambaNova is an inference-first company, and this role sits at the heart of that mission: turning state-of-the-art inference techniques into reliable, high-throughput, low-latency services exposed to customers through SambaStack and SambaCloud. The engineer will own end-to-end systems spanning request scheduling, advanced decoding algorithms, caching layers, API surfaces, and the accuracy infrastructure that keeps the stack trustworthy. This role partners closely with ML, compiler, runtime, and product teams to ship inference features from prototype to production.
Qualifications
Bachelor's degree in Computer Science, Electrical Engineering, or related field
5+ years of industry experience building and operating large-scale distributed systems, ideally in ML serving
Strong software engineering fundamentals: algorithms, data structures, concurrency, and systems design
Experience designing and maintaining production services with strict latency, throughput, and availability requirements
Working knowledge of modern LLM inference techniques and familiarity with open-source serving stacks such as vLLM, TensorRT-LLM, or SGLang
Proficiency in Python
Experience collaborating across teams to deliver complex, system-level engineering solutions
Key responsibilities
Design and productionize advanced inference techniques on RDU to optimize for performance and cost. Key areas include speculative decoding, constrained decoding, function/tool calling, prompt caching, and long-context inference.
Own SambaNova's integration with vLLM and adjacent serving frameworks, adapting them to RDU's architecture.
Own the public inference API surface exposed through SambaStack and SambaCloud.
Build and maintain the accuracy verification and regression infrastructure that gates every inference feature shipped to customers.
Partner with ML, compiler, runtime, and product teams to take inference features from prototype to production.
Contribute to technical design discussions, code reviews, and architectural decisions as a senior individual contributor.
Base Salary Range:
Base Pay Range
$200,000—$275,000 USD