Engineering Manager - Performance & Reliability

Feedzai · Portugal · Engineering

Posted 2026-08-05

Apply for this role →

The Engineering (Tech) Team is responsible for all Feedzai product development. Together with Product Management and Data Science, we build the next generation of tools to catch fraud in real-time with a machine learning first approach. Formed by engineers and managed by engineers, at Feedzai, you will find one of the most talented teams out there, from junior to senior engineers.

Driven to build the best value for our customers, our work involves a wide range of technical challenges. Such as building distributed systems that need to operate 24/7 with ultra-low latencies and solving UI/UX problems to help fraud analysts to fight fraud more efficiently. In addition, designing extensive databases from relational, NoSQL and graphs, validate and develop new data science techniques and algorithms.

With Cloud at its core, Platform Engineering powers the entire product development lifecycle. From initial testing to global deployment and 24/7 operations, we enable a true DevOps culture. We provide a fast-paced, open, and collaborative environment that encourages team members to lean in, experiment, and discover their full potential through continuous learning.

You:

We are looking for an experienced Engineering Manager to lead this team. You've built and operated distributed systems at scale — you know the difference between theoretical reliability and keeping a platform running 24/7 with strict latency budgets. You're not satisfied with "good enough" — you proactively identify reliability risks, push for automation over manual intervention, and raise the engineering bar for your team and adjacent teams.

Processing 10B+ events/month with strict latency and availability requirements, in our "you build it, you run it" culture, every engineering team owns the reliability of their services. The Performance & Reliability team exists to raise the bar across all of them — building the platform capabilities, tooling, and operational standards that make it easy for teams to ship reliable software and hard to ship fragile software.

Your Day to Day

Team Leadership: Lead and mentor a team of highly skilled Engineers, fostering a culture of technical excellence and continuous growth.

Reliability Ownership: Own the reliability roadmap — identifying, prioritizing, and executing on technical debt that threatens system health. Drive a "you build it, you run it" mindset focused on automation and reducing operational friction.

Performance at Scale: Champion performance profiling, capacity planning, and latency optimization. Ensure the team is proactively addressing bottlenecks before they become incidents.

Incident Culture: Build and improve incident response processes, postmortem practices, and on-call rotations that teams actually respect. Use data (SLOs/SLIs) to measure and improve system health.

Strategic Collaboration: Partner with Product, Engineering, Security, and Customer Success to define, enforce, and evolve engineering best practices. Balance delivery speed with sustainability.

Predictable Delivery: Assess team capacity realistically to ensure consistent delivery of platform commitments. Leverage team members according to seniority and strengths.

You Have and You Know-How

Distributed Systems Expertise: Hands-on experience building and operating distributed systems at scale — you've diagnosed cascading failures, optimized P99 latencies, and designed for fault tolerance and graceful degradation.

Proven Leadership: Extensive experience managing software engineering teams within environments that demand high levels of scalability, reliability, and operational excellence. You drive hiring aligned with execution needs and future roadmap demands.

Cloud & Orchestration: Deep understanding of major cloud providers (AWS, GCP, or Azure), Kubernetes, container ecosystems, and Infrastructure-as-Code practices.

Mission-Critical Mindset: A strong background in performance-sensitive systems where high availability is non-negotiable. You've operated systems with strict SLAs and know what it takes to meet them consistently.

Bar-Raising Drive: You're driven to eliminate entire classes of failures, not just patch symptoms. You bring a passion for reliability that elevates everyone around you.

Candid Mentorship: Proven ability to grow talent by providing candid, continuous feedback aligned with career frameworks, fostering a high-trust and psychologically safe environment.

Strategic Proximity: Ability to remain technically engaged enough to be accountable for architectural quality, security, and reliability decisions, without becoming a bottleneck in day-to-day operations.

The Product Team builds our product to disrupt the financial crime industry from a data-led approach. We partner with our clients using a holistic lens and have result-driven solutions to manage financial risk with a cloud-first platform and a world-class UX interface. Being part of this team, you have a voice in planning, strategising, and challenging the status quo. Your thoughts and ideas are valued. Our fast-paced and open environment encourages us to lean in, try new things, and discover our potential. We define and act on what could be in tomorrow's world, not on what is today. Join Us!

#LI-Remote #LI-MG3

Apply for this role →

← Back to all jobs