Principal Cloud Platform Engineer

SambaNova Systems · Bengaluru, Karnataka, India · Engineering

Posted 2026-06-30

Apply for this role →

About the team

The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.

About the role

As a Cloud Platform Engineer you'll keep our AI inferencing platform reliable, fast, and scalable. Your focus is uptime, latency, and resource utilization on the inference endpoints customers depend on. The work spans monitoring, deployment, capacity planning, and incident response, and includes a shared on-call rotation covering 24/7 coverage.

Responsibilities

Some of your responsibilities will include:

Owning the availability, latency, and efficiency of the production inferencing service, including change management, emergency response, and standing up AI infrastructure in new regions

Building and maintaining monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog covering service health, model latency and throughput, and RDU utilization

Leading incident response, running blameless post-mortems, and automating the toil those incidents expose

Finding and fixing performance bottlenecks, and designing auto-scaling policies that handle variable inference load without overspending

Managing cloud and on-prem infrastructure as code in Terraform and Ansible, and building the CI/CD pipelines that deploy new model versions and service updates

Required Qualifications

B.S. in Computer Science, Computer Engineering, or related field, or equivalent practical experience

3+ years in a Site Reliability Engineering, DevOps, or related role supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure)

Strong programming and scripting skills in Python, Go, or Java

Production experience with Docker and Kubernetes

Deep understanding of monitoring and observability tooling such as Prometheus, Grafana, Datadog, or the ELK Stack

Experience with infrastructure as code using Terraform or CloudFormation

Experience with CI/CD tooling such as Jenkins, GitHub Actions, or ArgoCD

Strong Linux system administration fundamentals

Preferred Qualifications

Experience in a hybrid environment spanning cloud and on-premise data center infrastructure

Experience supporting ML or AI inferencing services in production

Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs, for purposes of mapping to RDUs

Knowledge of model serving frameworks such as vLLM, SGLang, or Ray

Experience managing and tuning databases (SQL or NoSQL) and caching systems such as Redis or Memcached

Base Salary Range:

Base Pay Range

₹6,000,000—₹8,000,000 INR

Apply for this role →

← Back to all jobs