Sr. Cloud Ops Engineer
You are an experienced Cloud Ops Engineer who thrives in a fast-paced and AI-forward environment. You have a passion for innovation, solid design principles, and high-quality development. You excel at designing, improving, and maintaining secure, performant infrastructure and enjoy automating processes to ensure efficiency and scalability. You are a proactive problem solver with a strong understanding of system and networking concepts.
What You'll Do:
Infrastructure Design and Maintenance:
Design, improve, and maintain secure, durable, and performant infrastructure to power APIs, AI operations, web applications, and data mining/ETL workflows to meet established SLAs.
Collaborate with developers to bring new products and services into production.
Automation and Monitoring:
Automate testing, deployment, and monitoring of all products and services throughout the software development lifecycle.
Continuously improve operational processes and apply best practices to ensure scalability, security, and availability.
Security and Compliance:
Proactively meet standards for information security and compliance, such as SOC 2/ISO27001/CMMC.
Implement and uphold security measures across all infrastructure components.
Requirements:
Professional Experience:
At least 5 years of professional experience in a Cloud Ops / Platform Ops / DevOps role maintaining production infrastructure, preferably supporting a highly available environment for a SaaS or cloud service provider.
Technical Proficiency:
Strong working knowledge of AWS services such as EC2, ECS or EKS, Lambda, API Gateway, RDS, DynamoDB, Cloudwatch, S3, Code/Build/Pipeline/Deploy, VPC Lattic etc.
Strong working knowledge of Terraform or similar tools, Ansible, AWS CLI/SDK, Boto.
Proficiency with scripting languages such as Python, Bash, etc., and Linux environments.
Strong understanding of system and networking concepts and troubleshooting techniques for bare metal and containerized workloads.
Experience supporting AI and ML systems, agent orchestration frameworks (e.g., AgentCore or similar), and experience integrating LLMs into production systems.
Additional Skills:
Experience with release automation, system administration and configuration, and system debugging.
Nice to Have
Databricks or Snowflake infrastructure experience.
Cost optimization for AI and LLM workloads.
Internal developer platform experience.
Base Salary Range: $146,000 – $190,000