Principal Cloud Platform Engineer
About the team
The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.
About the role
As a Cloud Platform Engineer you'll keep our AI inferencing platform reliable, fast, and scalable. Your focus is uptime, latency, and resource utilization on the inference endpoints customers depend on. The work spans monitoring, deployment, capacity planning, and incident response, and includes a shared on-call rotation covering 24/7 coverage.
Responsibilities
Some of your responsibilities will include:
Owning the availability, latency, and efficiency of the production inferencing service, including change management, emergency response, and standing up AI infrastructure in new regions
Building and maintaining monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog covering service health, model latency and throughput, and RDU utilization
Leading incident response, running blameless post-mortems, and automating the toil those incidents expose
Finding and fixing performance bottlenecks, and designing auto-scaling policies that handle variable inference load without overspending
Managing cloud and on-prem infrastructure as code in Terraform and Ansible, and building the CI/CD pipelines that deploy new model versions and service updates
Required Qualifications
B.S. in Computer Science, Computer Engineering, or related field, or equivalent practical experience
3+ years in a Site Reliability Engineering, DevOps, or related role supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure)
Strong programming and scripting skills in Python, Go, or Java
Production experience with Docker and Kubernetes
Deep understanding of monitoring and observability tooling such as Prometheus, Grafana, Datadog, or the ELK Stack
Experience with infrastructure as code using Terraform or CloudFormation
Experience with CI/CD tooling such as Jenkins, GitHub Actions, or ArgoCD
Strong Linux system administration fundamentals
Preferred Qualifications
Experience in a hybrid environment spanning cloud and on-premise data center infrastructure
Experience supporting ML or AI inferencing services in production
Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs, for purposes of mapping to RDUs
Knowledge of model serving frameworks such as vLLM, SGLang, or Ray
Experience managing and tuning databases (SQL or NoSQL) and caching systems such as Redis or Memcached
Base Salary Range:
Base Pay Range
₹6,000,000—₹8,000,000 INR