Principal Cloud Platform Engineer

SambaNova Systems · Austin, Texas, United States; San Jose, California, United States · Engineering

Posted 2026-08-19

Apply for this role →

About the team

The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.

About the role

As a Principal Cloud Platform Engineer, you will be specializing in our AI Inferencing Service and will be the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability.

Responsibilities

Some of your responsibilities will include:

Shared ownership of the production inferencing service across regions, covering availability, latency, performance, change management, and capacity planning

Standing-up and automating AI infrastructure in new regions

Participating in a shared primary/secondary on-call rotation, and leading incident response

Building monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog for service health, model latency and throughput, and accelerator utilization

Finding and eliminating performance bottlenecks

Designing auto-scaling policies that handle variable inference loads

Managing cloud and on-prem infrastructure as code in Terraform and Ansible

Building CI/CD pipelines that safely deploy new model versions and service updates

Forecasting infrastructure needs against the product roadmap and usage trends, and working with finance to manage cloud spend

Defining and reporting on SLOs and SLIs for the inferencing platform, using that data to prioritize reliability work

Required Qualifications

B.S. in Computer Science, Computer Engineering, or related field

3+ years of experience in a Site Reliability Engineering, DevOps

Experience supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure)

Strong programming and scripting skills in languages like Python, Go, Rust, or Java

Proven experience with containerization and orchestration technologies (Docker and Kubernetes)

Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog)

Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)

Experience with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD)

Strong Linux/Unix system administration fundamentals

Preferred Qualifications

Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure.

Direct experience supporting ML/AI inferencing services in production.

Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs.

Knowledge of model serving frameworks like vLLM, SGLang or Ray.

Understanding of MLOps principles and practices.

Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached).

Base Salary Range:

Base Pay Range

$144,000—$189,000 USD

Apply for this role →

← Back to all jobs