Customer Reliability Engineer
Overview of Role
66degrees is looking for smart, energetic, and personable technology business professionals for a Site Reliability Engineer position. We are looking for someone with a broad range of skill sets with 4+ years of relevant experience.
As a Site Reliability Engineer you will be a part of data management and engineering projects within GCP environments, including migrating, analyzing, and managing data structures within GCP. Assist in the design, definition, development, and testing of GCP data solution components. Capable of quickly adapting and building plans around migrating data from legacy applications. Work closely with the GCP Security Team and other SME roles to ensure datastores are protected. Liaise between clients and developers to ensure that all data requirements are met.
Responsibilities
System Reliability & Automation: Ensure near-zero downtime by building monitoring, alerting, and self-healing automation. Create automated, available, and scalable systems using modern software and infrastructure principles.
Architecture & Documentation: Deliver comprehensive Cloud Architecture reviews, technical network diagrams, and thorough system documentation for internal tools.
Process Standardization: Actively aiming to eliminate operational toil and reduce manual administrative effort.
Client Operations & DevOps Advisory: Advise clients on deployment pipelines, high availability (HA), service reliability and best practices..
Cross-Functional Collaboration: Partner closely with clients, internal engineering, and Google engineers to rapidly investigate and resolve complex infrastructure issues.
Qualifications
Experience: Minimum 3+ years of cloud and infrastructure experience, featuring demonstrated expertise with Linux, Windows, Kubernetes (k8s), databases, and networking services.
Cloud Expertise: 2+ years of dedicated Google Cloud Platform (GCP) experience is highly preferred. Additional full-time experience with AWS and/or Azure is a strong plus. Relevant cloud certifications are preferred.
Automation & Provisioning: Strong proficiency with Python is required (other programming languages are a plus). Proven hands-on infrastructure provisioning and configuration experience using Terraform.
SRE Principles: Strong background in determining and negotiating Error budgets, SLIs, SLOs, and SLAs. Proven experience balancing service reliability metrics against operational toil.
Methodologies & Tools: Experience working within Agile Scrum and Kanban frameworks inside the SDLC. Familiarity with 24x7x365 monitoring, incident response, and troubleshooting across systems, networks, and code.
Communication: Exceptional written and verbal communication skills, with a proven track record in heavily customer-facing environments.
Education: Bachelor’s degree in Computer Science, Electrical Engineering, or an equivalent technical field.
Shift Flexibility: Ability to work designated shift hours on weekdays during UK business hours.