Senior Site Reliability Engineer
We are looking for a Senior Site Reliability Engineer (SRE) to join the Site Reliability Engineering team at SentinelOne. This organization’s mission is to keep our uptime promise to our customers by ensuring we meet our SLOs/SLAs, help our engineering teams ship software to our customers fast and with quality, and ensure our customers are successful. We are looking to add a Senior SRE who has experience running incident post-mortems, automating repetitive operational tasks, improving alerting accuracy, and building and refining processes that reduce downtime. You will work closely with cross-functional teams to lead reliability initiatives and bring best practices to our team.
We value good written communication skills, data-driven decisions, and a keen eye for continuous improvements. You’ll help simplify, have a passion for new ideas and know how to execute iteratively toward the final goal. We value candor and collaboration.
What Will You Do?
Primary responsibilities include:
Participate in and help drive incident management for production issues, ensuring rapid recovery and root cause analysis; may lead postmortems
Improve and optimize the observability strategy
Collaborate with application engineering teams to design and implement monitoring solutions that enhance our alerting capabilities and reduce noise
Help define and refine SLOs, SLIs, and SLAs that align with business objectives and customer expectations
Conduct post-incident review, documenting findings and driving follow-up actions to prevent recurrence
Mentor peers and junior engineers in incident response, troubleshooting techniques, and reliability best practices
What Skills and Knowledge Will You Bring?
Ideal candidates will have:
5+ years of experience in Site Reliability Engineering, DevOps, or a related field in cloud native environments
Solid experience troubleshooting complex issues under pressure, using and improving established runbooks
Experience with Kubernetes and container orchestration
Experience with industry standard observability stacks (Prometheus, Grafana, ELK, OpenTelemetry, etc.)
Proficiency in Python and Bash scripting to improve operational workflows and incident response
Familiarity with modern CI/CD pipelines and DevOps practices
Excellent communication skills with demonstrated ability to mentor peers in reliability practices
Why SentinelOne?
AI is redefining how the world operates and rewriting the rules of security in real time, and SentinelOne was built for this moment. From day one, we architected an AI-native platform designed to operate at machine speed, not as an add-on to legacy systems but as the foundation itself. If you want to build where innovation and impact move together, this is that place.
We invest in our Sentinels with comprehensive, competitive benefits designed to support you and your family:
Equity & Rewards
Restricted Stock Units (RSUs)
Employee Stock Purchase Plan (ESPP)
Performance-based bonus
Time Off & Wellbeing
Flexible time off, on top of the standard 5 weeks PTO
Paid company wellness days and fully paid short-term sick/nursing leave
Gender-neutral parental leave
Insurance & Financial Security
Private medical care benefit
Life and disability insurance
Pension Insurance contribution
Global business travel medical insurance
Employee Assistance Program (EAP)
Work Perks & Flexibility
Global home office allowance
Meal allowance
Hybrid work in Prague (Karlin), Brno (Clubco) or remote in CZ/SK. Only Prague-based employees are required to work from the office at least 2 days//week.
Wellness & Lifestyle
Wellbeing allowance
MultiSport benefit program
Wellness Coach app