Site Reliability Engineer, AI Platform
Algolia was built to help users deliver intuitive search experiences across websites and mobile applications. Our Search API serves thousands of customers in more than 100 countries, answering billions of queries every month.
Join AI Platform: Powering AI in Production
AI Platform builds and operates the shared production foundations supporting Algolia's evolving AI ecosystem.
The team works at the intersection of Site Reliability Engineering, cloud infrastructure, software engineering and AI, helping engineering teams bring AI-powered capabilities to production reliably, securely and efficiently. Our scope includes Kubernetes, cloud infrastructure, CI/CD, networking, databases, observability, reliability, FinOps and production operations.
We are looking for a Site Reliability Engineer with strong production fundamentals who enjoys solving operational problems, automating repetitive work and progressively taking ownership of complex systems at scale.
YOU WILL:
Build and operate production infrastructure supporting AI-related workloads and services
Operate and improve highly available Kubernetes-based platforms
Improve reliability through SLOs, observability, alerting and capacity management
Investigate production issues and turn findings into durable fixes and improvements
Work across networking, databases, compute and service infrastructure
Improve CI/CD pipelines, deployment automation and developer experience
Build and maintain infrastructure using Infrastructure as Code
Participate in on-call, incident response and operational improvements
Collaborate with experienced engineers across AI Platform and progressively take ownership of broader production areas
YOU MIGHT BE A FIT IF YOU HAVE:
Solid hands-on Kubernetes knowledge, including workloads, resource management, and production operations
Strong experience with Infrastructure as Code, and the lifecycle of cloud infrastructure
Solid experience building and operating CI/CD pipelines and automated deployment workflows
Hands-on experience with at least one major cloud provider: GCP, AWS or Azure
Good understanding of networking, distributed systems and reliability engineering
Experience with monitoring, observability and troubleshooting production systems
Strong automation mindset and the ability to take ownership of well-defined production systems and progressively tackle more complex problems
Excellent written and spoken English
NICE TO HAVE:
Go and/or Python engineering experience
Exposure to AI/ML infrastructure and inferences
Comfortable working AI-first, using coding agents, agentic workflows and AI-assisted debugging to accelerate engineering and operations
Algolia does not discriminate on the basis of race, color, religion, sex, age, national origin, military status, veteran status, disability status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.
The annual base salary compensation range for this role reflects market pay data within this location. The exact compensation offered for this role may vary depending on specific location and job-related knowledge, technical skills, and experience; and is only one part of our Total Rewards philosophy to compensate and recognize employees for their work.
Base Salary Pay Range
€69.768—€96.900 EUR