Head of Engineering, Infrastructure & SRE
The Opportunity
Postman is seeking a strategic and results-driven engineering leader who is passionate about cloud agnostic infrastructure, operational excellence, and enabling engineering teams to operate autonomously and build with confidence.
As Head of Infrastructure, you'll lead a talented and geographically distributed team of engineers across the SF Bay Area, India, and Europe, fostering a culture of collaboration, ownership, and continuous improvement. You'll own the infrastructure that underpins one of the world's most widely used API platforms, an environment handling ~80,000 requests per second at the front door, and be responsible for its reliability, scalability, and evolution. In addition to infrastructure, you'll own the Site Reliability Engineering (SRE) function at Postman, setting the standards and practices that keep the platform reliable at scale. You'll work closely with engineering managers, product managers, and platform teams to drive the technical roadmap for our cloud agnostic infrastructure and reliability practices, ensuring we can support a large and rapidly growing engineering organization.
If you're passionate about building resilient, scalable infrastructure, have a proven track record of leading distributed engineering teams, and thrive in fast-paced environments where your decisions have company-wide impact, we want you on our team!
What You'll Do
Team Leadership & Development
Hire, manage, mentor, and coach a geographically distributed team of infrastructure and SRE engineers across SF Bay Area, India, and Europe, helping them grow both technically and professionally.
Drive and uphold a culture of respect, integrity, inclusion, ownership, and accountability within the team.
Set clear goals and provide regular feedback to ensure your team is motivated and aligned with the platform and company vision.
Build a high-performing team with the skills and practices to reliably operate and evolve cloud agnostic infrastructure at scale.
Foster psychological safety and cross-regional collaboration across time zones.
Technical Leadership
Own the architecture and evolution of Postman's cloud agnostic infrastructure, driving improvements that increase reliability, performance, and cost efficiency at scale.
Lead the design and implementation of infrastructure improvements across Kubernetes, Cluster API, Argo, Helm, Crossplane, service mesh (Istio), AWS, and Azure environments.
Own the SRE function end to end: SLIs/SLOs, error budgets, capacity planning, incident management, and reliability engineering practices across the platform.
Set the technical direction for how Postman's infrastructure and reliability practices evolve to support a large engineering organization with hundreds of services and dozens of teams, with a focus on enabling product teams to operate autonomously.
Partner with platform, security, and product engineering teams to ensure infrastructure and reliability decisions align with broader company goals.
Ensure infrastructure and reliability best practices are upheld across the organization, including GitOps, CI/CD, observability, on-call, and incident response.
Project Management
Own the infrastructure and reliability roadmap, balancing operational reliability with longer-term architectural investments.
Break down complex infrastructure and reliability initiatives into clear, actionable milestones and manage delivery on time and at high quality.
Proactively identify and resolve roadblocks, working across teams to unblock engineering work and minimize customer impact.
Collaboration & Communication
Work closely with stakeholders across engineering, product, and security teams to align on infrastructure and reliability priorities and constraints.
Foster open communication within and across teams, promoting transparency on system health, risk, error budgets, and roadmap.
Represent infrastructure and SRE in leadership forums, advocating for technical needs and communicating clearly on trade-offs.
Operational Excellence & Site Reliability Engineering
Own Postman's SRE function, establishing and evolving on-call practices, escalation policies, monitoring, and incident management across the company.
Define and track SLIs/SLOs and error budgets for critical services, using them to guide investment decisions and prioritization.
Drive a blameless postmortem culture, ensuring incidents produce durable learnings, clear action items, and measurable follow-through.
Drive a culture of continuous improvement, learning from incidents, automating toil, and reducing operational burden so that product engineering teams can ship independently.
Champion proactive reliability practices such as load testing and capacity planning to stay ahead of scale.
Maintain high standards for security, cost management, and infrastructure quality across all environments.
About You
You are a seasoned infrastructure and reliability leader with deep technical expertise and a track record of building and managing high-performing, distributed infrastructure and SRE teams. You've owned large-scale cloud agnostic environments and the reliability practices that keep them running, and know how to balance the demands of keeping the lights on with investing in architectural improvements that compound over time. You understand that great infrastructure and reliability engineering are ultimately about enabling teams to move fast independently and with confidence.
Must Have Qualifications
15+ years of experience in infrastructure, platform, or site reliability engineering, with 7+ years in an engineering management or leadership role.
Proven experience managing geographically distributed teams across multiple time zones.
Hands-on experience with Kubernetes, Cluster API, Argo, Helm, and Crossplane in large-scale production environments.
Experience with service mesh technologies, including Istio.
Deep experience with both AWS and Azure cloud infrastructure.
Demonstrated experience building or leading a Site Reliability Engineering function, including on-call, incident management, SLIs/SLOs, and error budgets.
Demonstrated ability to design and implement cloud agnostic infrastructure that enables engineering teams to be autonomous and self-sufficient.
Experience supporting a large engineering organization: multiple product teams, hundreds of services, high deployment frequency, and significant traffic scale.
Excellent communication skills, with the ability to translate complex infrastructure and reliability topics for both technical and non-technical audiences.
Nice-to-Have
Experience with GitOps workflows and infrastructure-as-code at scale.
Familiarity with FinOps / cloud cost optimization practices.
Prior experience in a SaaS or API-focused product company.
A passion for developer experience: building infrastructure and reliability practices that make product engineers faster and more confident.