Engineering Manager, SRE
What this job can offer you
Remote's SRE team exists so that our engineers can move quickly and our customers get a product that stays up. The team owns Kubernetes, AWS, PostgreSQL, CI infrastructure, our observability stack and the reliability practices that sits on top of all of it.
We are looking for a Team Leader to run that team. This is a 60% IC, 40% leadership role. You will own the career development of your reports, steer the teams focus using judgment against the company goals, and you will be the spokesperson for the team across engineering. You will also stay close enough to the technical work to set direction with credibility and to know when something is going wrong before it is escalated to you.
Reliability practice at Remote is maturing rather than mature. Our SLO framework is live on its first few teams and needs to reach the rest, there is real work to do on how we balance operational load against project delivery. If you want a team where the foundations are in place and the interesting problems are still open, this is that team.
What you bring
People leadership
You have led an SRE, infrastructure or platform engineering team, and you have owned your reports' growth, performance and career progression rather than just their sprints.
You coach both craft and the soft skills, you can point to people who grew because of it.
You handle underperformance directly and early, with clarity and empathy.
You have hired engineers, and you can tell the difference between a good interview and a good engineer.
You read team dynamics well and you resolve conflict rather than routing around it.
You have a natural talent fostering commitment to the goals of the company.
Technical depth
A hands-on background in site reliability, DevOps or cloud infrastructure engineering, deep enough that you can review your team's work, challenge a design and be taken seriously in an incident.
Kubernetes in production, including the operational reality of it rather than the happy path.
AWS at meaningful scale.
Hands on AI building, enablement, scaling AI infrastructure.
Solid o11y practices and principles,
Infrastructure as code with Terraform.
CI/CD systems such as GitLab CI, GitHub Actions or Jenkins.
Docker and shell scripting.
You have run a reliability practice: incident response, on-call, SLOs and error budgets, and the discipline of turning incidents into changes that stick.
Understanding and history of working in regulated environments.
Ways of working
You prioritise exceptionally well when operational load and project work compete, and you protect your team's focus without dropping the operational commitment.
You write clearly. Remote is fully distributed and async, so most of your leadership will happen in writing.
You build relationships across teams. A lot of SRE's value comes from being the team others bring problems to early.
Nice to have
Working knowledge of a backend language, ideally Elixir, or otherwise Java, Clojure, Node.js, Python or similar.
Depth in modern observability: OpenTelemetry, distributed tracing, and tools such as Honeycomb.
Database operations experience, particularly PostgreSQL or Aurora performance, connection pool health and query tuning.
Running and configuring Linux systems outside a cloud environment.
Security capability from both a defensive and an offensive standpoint.
Cloud cost management and FinOps.
Experience growing a team from a small base, including building the hiring bar as you go.
Key Responsibilities
Your people
The full career lifecycle of your reports: onboarding, feedback, performance assessment, progression and hiring.
Team health, dynamics and the retrospective habit that keeps them honest.
Being the team's spokesperson to the rest of engineering and to senior leadership.
Delivery
The SRE goals: what the team commits to, in what order, and why.
The support rotation and on-call model.
The platform
Remote's core infrastructure: Kubernetes, AWS, PostgreSQL, DNS and TLS, CI infrastructure.
The reliability practice: SLOs, error budgets, incident response and the observability stack.
The partnership with our Security team on threats, patching and infrastructure controls, including our audit and compliance obligations.
The vendor relationships that sit behind the platform, including renewals and commercial conversations with support from your Director.
Practicals
You will report to: Director of Engineering, Platform
Team: Site Reliability Engineering, part of Platform Engineering
Direct reports: 4
Location: Anywhere in the world. Our current coverage is strongest in EMEA and APAC, so candidates who overlap with either, or who can help us close the Americas gap, are especially welcome.
Start date: As soon as possible
Application process
Interview with recruiter
Interview with hiring manager
Scenario interview with a peer team leader
Executive interview
Bar Raiser interview
Offer + Prior employment verification check
#LI-DNP
Remote's Total Rewards philosophy is to ensure fair, unbiased compensation and fair equity pay along with competitive benefits in all locations in which we operate. We do not agree to or encourage cheap-labor practices and therefore we ensure to pay above in-location rates. We hope to inspire other companies to support global talent-hiring and bring local wealth to developing countries.
At first glance our salary bands seem quite wide - here is some context. At Remote we have international operations and a globally distributed workforce. We use geo ranges to consider geographic pay differentials as part of our global compensation strategy to remain competitive in various markets while we hiring globally.
Our salary ranges are determined by role, level and location, and our job titles may span more than one career level. The actual base pay for the successful candidate in this role is dependent upon many factors such as location, transferable or job-related skills, work experience, relevant training, business needs, and market demands. The base salary range may be subject to change.
At Remote, we foster internal mobility as a key element of our culture of employee growth and development, supported by a compensation philosophy that guarantees pay equity and fairness. Therefore, all compensation changes associated with an internal move will be reviewed by the Total Rewards & People Enablement team on a case by case basis.
The annual salary range for this full-time position is
$75,450—$169,700 USD