AI/ML Engineer

PagerDuty · Lisbon · Engineering

Posted 2026-08-19

Apply for this role →

About the role

PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for an early-career AI/ML Engineer who is excited to grow at the intersection of two disciplines: large-scale distributed systems and machine learning.

In this role you will help build and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll work alongside senior engineers on real production problems, learning how AI features go from a prototype to something that serves reliably at scale.

We are looking for a candidate who is genuinely excited about building with modern AI — LLMs, agents, and retrieval — eager to learn how resilient, high-throughput systems are built, and motivated to grow into an engineer who is strong in both.

What you’ll do

Contribute to AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time data, with support and guidance from senior engineers.

Help build and maintain the systems behind them — prompt and agent orchestration, retrieval pipelines, tool/API integrations, and inference services — writing code, adding tests, and improving observability.

Learn to reason about latency, throughput, cost, and reliability of LLM-powered services, and apply those lessons in the code you write.

Help take AI features from prototype toward production, and support the evaluation and monitoring loops that keep them accurate and trustworthy over time.

Partner with platform, product, and applied-research teams, asking good questions and turning requirements into working code.

Grow through code review, pairing, and mentorship, and steadily take on more ownership as you develop.

What you’ll bring

2+ years of software engineering experience building and shipping production software.

Degree in CS/a related field, or equivalent practical experience.

Solid programming fundamentals and comfort moving between application code and AI/model code.

Hands-on experience with modern AI — building something real with LLMs, prompting, retrieval, or agent frameworks.

Some exposure to distributed systems and how software runs reliably at scale, and eagerness to deepen it.

Strong communication and collaboration skills.

Nice to have

Personal, academic, or internship projects involving LLM apps, agents, RAG, or backend services.

Exposure to cloud infrastructure (AWS, GCP, or Azure), containers, or Kubernetes.

Familiarity with the ecosystem — e.g. LLM APIs and frameworks such as LangChain or LlamaIndex, vector databases, and orchestration/streaming tools like Kafka or Airflow.

Interest in agentic systems, evaluation and guardrails for LLMs, or applied problems like anomaly detection and event correlation.

Contributions to open-source projects.

Why PagerDuty

At PagerDuty, AI it’s core to how we help the world’s teams keep their digital services running. This is a place to start your career on problems where scale, latency, and correctness genuinely matter, surrounded by engineers who will invest in helping you grow and who build systems that people depend on in their most critical moments.

Apply for this role →

← Back to all jobs