Senior Site Reliability Engineer, AI Platform

Algolia · Paris, France · Engineering

Posted 2026-09-18

Apply for this role →

Algolia was built to help users deliver intuitive search experiences across websites and mobile applications. Our Search API serves thousands of customers in more than 100 countries, answering billions of queries every month.

Join AI Platform: Powering AI in Production

AI Platform builds and operates the shared production foundations supporting Algolia's evolving AI ecosystem.

The team works at the intersection of Site Reliability Engineering, cloud infrastructure, software engineering and AI, helping engineering teams bring AI-powered capabilities to production reliably, securely and efficiently. Our scope includes Kubernetes, cloud infrastructure, CI/CD, networking, databases, observability, reliability, FinOps and production operations.

We are looking for a Senior Site Reliability Engineer who can independently own complex production systems, drive technical decisions across teams, and help shape reliable and efficient infrastructure at scale.

YOU WILL:

Own and evolve production infrastructure supporting AI-related workloads and services at scale

Design and operate highly available Kubernetes-based platforms

Drive reliability through SLOs, observability, capacity planning and production guardrails

Lead complex production investigations and turn findings into durable architectural improvements

Improve shared infrastructure across networking, databases, service communication and compute

Build better CI/CD, progressive delivery, automation and developer experience

Drive cloud infrastructure efficiency and FinOps initiatives

Participate in and improve on-call and incident response

Mentor engineers and raise the technical bar for reliability and production engineering

YOU MIGHT BE A FIT IF YOU HAVE:

Strong hands-on production experience with at least one major cloud provider: GCP, AWS or Azure

Strong experience designing and operating Kubernetes and cloud-native production systems at scale

Strong understanding of distributed systems, networking and reliability engineering

Experience operating business-critical systems with strong availability, scalability and operational requirements

Ability to independently own ambiguous, cross-team technical problems and drive them to measurable outcomes

Strong automation mindset and ability to balance reliability, engineering velocity and cost

Excellent written and spoken English

NICE TO HAVE:

Go and/or Python engineering experience

Experience with infrastructure supporting AI/ML workloads, model serving, GPUs or other compute-intensive systems

Comfortable working AI-first, using coding agents, agentic development workflows, AI-assisted debugging and automation to accelerate engineering and operations

Algolia does not discriminate on the basis of race, color, religion, sex, age, national origin, military status, veteran status, disability status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.

The annual base salary compensation range for this role reflects market pay data within this location. The exact compensation offered for this role may vary depending on specific location and job-related knowledge, technical skills, and experience; and is only one part of our Total Rewards philosophy to compensate and recognize employees for their work.

Base Salary Pay Range

€69.768—€96.900 EUR

Apply for this role →

← Back to all jobs