Platform Engineer II
Overview:
We are looking for a Platform Engineer who will work with senior management, Platform Engineering, QA and development teams to continuously improve the stability, reliability and efficiency of our global SaaS platform through automation solutions and platform tools. This individual will contribute to building the software, automation frameworks, and systems that make our production environment more resilient, enable self-service capabilities, and reduce operational toil through code.
What you will do:
Design, develop, and maintain automation tools and platform capabilities using Go, Python, or similar languages to help eliminate manual toil and build self-service components across infrastructure, observability, reliability, and security domains.
Build and enhance internal tools, integrations, and workflows that improve operational efficiency and developer productivity within the SaaS platform.
Develop and support automation frameworks for incident response, infrastructure testing, monitoring optimization, and compliance validation, including CI/CD pipeline enhancements.
Help develop observability and monitoring solutions including setting up alerts, contributing to noise reduction systems, automated SLO management, and cost optimization.
Partner closely with SaaS Engineering to deploy application and infrastructure updates using automation tools.
Available for on-call duties for incident response, blameless postmortems, and helping translate learnings into automated solutions.
Contribute to build playbook documentation and procedures.
What we are looking for:
Bachelor’s degree in Computer Science, Engineering, or a related field (equivalent work experience will be considered in lieu of a degree).
3-5 years of experience in software development, platform engineering, Site Reliability Engineering (SRE), or DevOps, with hands-on experience supporting production environments.
Strong technical expertise in Kubernetes, cloud platforms (AWS, Azure, or GCP), Go/Python development, Linux, Infrastructure as Code (Terraform, CloudFormation, Ansible), CI/CD, container technologies, APIs, microservices, distributed systems, and SaaS production operations.
AI Skills/Knowledge:
Familiarity with AI-assisted tools such as Microsoft Copilot, GitHub Copilot, Claude, Cursor, or similar tools to support troubleshooting, root cause analysis (RCA), coding, code review, and operational problem solving.
Ability to leverage AI tools to improve infrastructure automation, scripting, documentation, test generation, and engineering productivity.
Familiarity with LLM APIs and agentic AI workflows, with an interest in applying AI capabilities to platform engineering, DevOps, and reliability engineering workflows.
Awareness of responsible AI principles, including security, bias, explainability, and appropriate use of AI-generated outputs.
Preferred Skills:
Experience contributing to automation projects, internal tools, or monitoring implementations (i.e. Datadog, Prometheus, Grafana, Backstage) is a plus.
Experience with Kubernetes operators, controllers, or orchestration concepts is a plus.
FedRAMP, SOC2, or compliance framework exposure is a plus.
#LI-SR1