Principal DevOps Lead
Location: Berlin, Germany Remote Status: Fully Remote (#LI-Remote)
LivePerson (NASDAQ: LPSN) is a leader in trusted enterprise conversational AI and digital transformation. The world's leading brands use our award-winning Conversational Cloud platform to connect with millions of consumers. We power nearly a billion conversational interactions every month, providing uniquely rich data analytics and safety tools to unlock the power of conversational AI for better business outcomes. Fast Company named LivePerson the #1 Most Innovative AI Company in the world.
Position Overview
The Observability Platform team is building a state-of-the-art observability ecosystem for logging, monitoring, and tracing across cloud and on-premises data centers. We’re looking for an experienced Senior DevOps Lead to lead our Logging and Monitoring initiatives and drive the development of robust, scalable observability solutions within Google Cloud Platform (GCP).
In this role, you’ll help build systems that give software engineers greater visibility into the health, performance, and reliability of their applications and services. You’ll oversee a modern observability technology stack that includes Elastic Cloud, Loki, Grafana Labs, and the ELK stack for logging, as well as Zabbix, Captain Hook, and Anodot for metrics, monitoring, and anomaly detection.
You Will: Key Responsibilities & Impact
Lead the design, implementation, operation, and continuous improvement of LivePerson’s observability platforms across logs, metrics, traces, alerting, and synthetic monitoring.
Own and optimize large-scale observability pipelines processing high volumes of telemetry data daily, leveraging technologies such as Filebeat, Kafka, Logstash, Elastic Cloud, Prometheus, OpenTelemetry, Grafana Labs, Zabbix, Anodot, and related tools.
Design, build, and maintain scalable, Kubernetes-based observability services using Helm, CI/CD pipelines, GCP, GKE, Docker, and cloud-native best practices.
Define and establish observability standards, dashboards, alerting frameworks, best practices, and onboarding resources for hundreds of engineering users.
Partner closely with DevOps, SRE, Engineering, NOC, Security, and vendor teams to deliver reliable, scalable, and actionable observability solutions.
Evaluate emerging observability technologies and drive adoption of modern practices in areas such as OpenTelemetry, distributed tracing, anomaly detection, and proactive monitoring.
Lead the design and development of an end-to-end Synthetic Monitoring platform, running hundreds of daily synthetic tests on GCP Spot VMs to enable proactive service validation and help engineering teams identify potential issues before they impact customers.
Provide technical leadership and mentorship, driving best practices across observability, DevOps, and cloud engineering.
You Have: Required Skills & Qualifications
5+ years of experience in software engineering, DevOps, SRE, or a related discipline, with a strong background in application development and cloud engineering.
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
Strong experience with Kubernetes and containerization technologies, including Docker.
Extensive experience with observability and monitoring technologies such as Grafana Labs, Captain Hook, Zabbix, Fluentd, ELK, Kafka, and Prometheus.
Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, or CloudFormation.
Experience working with cloud platforms such as GCP, AWS, or Azure, including compute, storage, networking, and related cloud services.
Strong programming and scripting skills in one or more languages, such as Go, Java, JavaScript, or Python.
Experience designing and operating highly available, scalable, production-grade systems.
Strong understanding of CI/CD practices and cloud-native development methodologies.
Experience with OpenTelemetry Collector and Grafana Agent is highly preferred.
Excellent problem-solving, communication, and collaboration skills, with the ability to influence technical direction across engineering teams.
Our Benefits & Perks
We are committed to supporting the complete well-being, health, financial security, family, and professional growth of our permanent employees.
💰 Financial Security & Growth
Additional Pension scheme: deferred pension scheme (LivePerson contributes 20%), ESPP and annual bonus depending on achievements
Equipment: Internet and Mobile reimbursement
Development: Access to internal professional development resources.
👨👩👧👦 Time Away & Family Support
Flexible Paid Time Off (PTO): Personal time off 33 days , vacation (28 days) and care days (5)
Paid Public Holidays.
Volunteering Days: Use them to make a difference in your community. Whether it's a cleanup, supporting a local initiative or holding an “Ehrenamt”, just let us know! We're here to help you create a meaningful impact.
💻 Workplace Flexibility
Remote-First Model: Be flexible, work flexibly and from anywhere - no defined working hours or boundaries with offices in Mannheim and Berlin
Monthly Connection: Join our monthly connection pizza day in the office!