Site Reliability Engineer
Title: Site Reliability Engineer, Senior Site Reliability Engineer
Department: Platform Engineering
Reports To: Director, Platform Engineering
FLSA Status: Exempt
Location: US Remote
Job Summary
We are seeking Mid to Senior level Site Reliability Engineers to join our Platform Engineering team. In this role, you will be instrumental in designing, implementing, maintaining and deploying highly available complex, scalable and reliable systems leveraging automation, effective monitoring and infrastructure-as code. Working closely with our application engineering teams to ensure our services meet the highest standards of reliability, performance, and security.
Key Responsibilities for Senior level
The Site Reliability Engineering teams at Prove are responsible for driving maximum uptime for existing and developing products. Qualified candidates will be well versed in the difference between methods and ownership of outcomes and be able to demonstrate and document their relevant experience.
Observability Leadership
Design and implement comprehensive observability solutions across our infrastructure and within applications
Establish metrics, logging, and tracing systems that enable quick identification and resolution of issues
Create alerting thresholds and automated responses based on service level objectives (SLOs)
Provide actionable insights into service to service communications
Infrastructure Management
Design, build, and maintain scalable cloud infrastructure on AWS
Implement infrastructure-as-code using tools such as Terraform
Automate routine operational tasks to reduce toil and improve efficiency
Ensure infrastructure security compliance and implement least-privilege access controls
Design and implement infrastructure-as-code deployments for container based applications
Scale containers based on custom metrics for applications and critical observability infrastructure
Incident Response
Conduct thorough post-incident reviews and implement preventative measures
Use observability data to perform root cause analysis and system improvements
Participate in a 24/7 on call rotation to achieve 99.999% system availability.
Required Qualifications for Senior level
5+ years of experience in Site Reliability Engineering, Platform Engineering or equivalent experience. Software Engineering roles with a strong infrastructure and production engineering aspect also qualify.
Expert knowledge of observability platforms and practices (OpenTelemetry, Prometheus, Grafana, Jaeger, ELK stack / Splunk, etc)
Experience with Kubernetes and container orchestration
Strong experience with infrastructure-as-code tools (Terraform, Spacelift, Pulumi)
Proficiency in at least one programming language ( Go, Python )
Deep understanding of cloud platforms, preferably AWS
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
Key Responsibilities for IC3
The Site Reliability Engineering teams at Prove are responsible for driving maximum uptime for existing and developing products. Qualified candidates will be well versed in the difference between methods and ownership of outcomes and be able to demonstrate and document their relevant experience.
Preferred Qualifications
Experience with distributed systems and microservice architectures
Experience working in a high compliance environment
Hand-on experience instrumenting code with OpenTelemetry
Familiarity with service mesh technologies
Contributions to open-source projects
Experience in the identity verification or financial technology industry
Application development experience
Optimize
Improve new and existing systems by increasing reliability, performance, and scalability
Automate routine operational tasks to reduce toil and improve efficiency
Ensure infrastructure security compliance and implement least-privilege access controls
Implement efficient infrastructure that balances rapid development and cost
Embrace technological changes and development practices while maintaining reliability
Respond
Participate in a 24/7 on-call rotation
Conduct thorough post-incident reviews and implement preventative measures
Use observability data to identify system improvements
Run
Implement infrastructure as code in a myriad of high compliance development, production, and other environments
Scale developer experiences by being the standard bearer of an opinionated platform approach
Required Qualifications for IC3
3+ years of experience in Site Reliability or Platform Engineering teams
Deep understanding of cloud platforms, particularly AWS
Strong experience with Kubernetes and container orchestration
Experience withTerraform and infrastructure-as-code tools
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
Preferred Qualifications
Experience with distributed systems and microservice architectures
Experience working in a high compliance environment
Experience with holistic monitoring and alerting for developing platforms
Skilled proficiency in at least one programming language (Go, Python)
Benefits & Perks for FTE Provers:
Competitive salaries & Bonus Plan (for eligible roles) and Equity Plan
Modern Health for financial, mental, and physical wellness
401(k) Retirement Plan & Match (US Offices) and Local Country Pension (International Offices)
Unlimited Vacation and Flexible hours
Comprehensive medical benefits for you and your family ❤️
Emotional & Physical Wellness – Access to wellness services (EAP & Prove Well-Being Reimbursement)
Bottomless snacks & beverages for certain office locations
Daily GrubHub stipend for lunch if coming into the office (US Offices)
A great place to work and connect with other talented Provers like yourself!
This position description should not be considered the final description of the position. The position description is not intended to be an all-inclusive list of duties and standards of the positions. It should be assumed that we would, to some extent, structure responsibilities in accordance with the successful candidate’s capabilities and changing business conditions. Incumbents will follow any other instructions, and perform any other related duties, as assigned by their supervisor.
Site Reliability Engineer:
Metro 2: $130,000 - 150,000
Metro 3: $120,000 - 135,000
Senior Site Reliability Engineer:
Metro 2: $166,000 - 185,000
Metro 3: $153,000 - 171,000
Plus variable commission / company bonus. Offered salary will be determined by the applicant’s education, experience, knowledge, skills, geo-location and abilities, as well as internal equity and alignment with market data.