Senior Software Engineer, Cloud Image/ML Compute Architecture
About the role
Muon Space is expanding its High Performance Compute (HPC) team to accelerate mission-critical data products and unlock a growing pipeline of scientific imaging and ML workloads across ground and flight systems.
This role works directly with the team that produces our image data products. As the architect for our cloud image processing and ML infrastructure, you'll design, build, and deploy the serving platform that runs Muon's image algorithms and ML models in production. Architecture and performance optimization are the core of the job: you'll set the patterns the rest of the team uses to take research models to a low-latency production service, and you'll own the interfaces by which our data pipelines invoke that serving layer.
This position is hybrid and requires working on-site in our San Jose, CA office three days per week.
Responsibilities
Architect, develop, align and deploy the cloud image processing and ML serving stack for Muon's ground data products, including runtime integration, request batching, and serving multiple customer data feeds.
Design distributed, low-latency image processing systems that scale in the cloud.
Define the patterns and interfaces by which data pipelines call into performant native/accelerated code.
Set up a production-grade image processing and ML serving path for the Image Data Product pipeline, and extend the architecture to support upcoming imaging and onboard-compute missions and future onboard accelerator opportunities.
Deliver high-performance reference algorithm implementations and benchmark them.
Mentor other engineers on modern C++, image and ML runtimes, and distributed-systems design.
Own verification and validation, including unit, integration, and performance regression tests.
Produce clear designs, trade studies, and interface documentation.
Required Qualifications
Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Math, or a related technical field.
6+ years of professional software engineering experience, including technical leadership on distributed computation or data-intensive systems.
Expert-level C++, with deep familiarity with memory management, concurrency, and performance-oriented programming in production systems.
Strong PyTorch experience and production experience running models in C++ runtimes in the cloud using LibTorch or equivalent.
Demonstrated experience architecting distributed systems for low-latency image and ML, data analysis, or computer vision.
Strong Python skills, including interaction with native optimized implementations.
Strong written and verbal communication and ability to lead and conclude cross-functional technical discussions.
Ability and willingness to obtain and maintain a U.S. security clearance. Active clearance is a plus.
Nice-to-Have Skills
Experience with accelerators beyond GPUs (e.g., TPUs) and reasoning about accelerator tradeoffs.
Model optimization for inference (quantization, graph compilation, kernel fusion).
Production ownership for latency-critical production systems.
GPU acceleration (CUDA, HIP, or similar).
CI/CD for containerized workloads, including performance regression testing.
Direct exposure to space, aerospace, or other mission-critical software environments.
Salary
The salary range for this role is $232,000 - $261,000, plus a competitive equity grant and comprehensive benefits package. Final compensation will be determined based on skills, qualifications, experience, and geographic location as assessed during the interview process.