Senior Platform Software Engineer I

Braze · London · Engineering

Posted 2026-10-09

Apply for this role →

Braze is the leading customer engagement platform, helping brands build personalized messaging experiences across push, email, in-app, and more at massive scale. Our technology stack is rooted in Ruby on Rails, MongoDB, Redis, Kafka, Kubernetes and more.

The In-Memory Databases & Observability team provides the foundational infrastructure and tools that keep Braze fast, reliable, and observable. Valkey sits at the heart of Braze: every message we send flows through it and our fleet serves around 36 million operations per second while holding terabytes of data in memory. We operate Valkey, Memcached, and Elasticsearch, along with the observability stack (Prometheus, Thanos, and related tooling) that gives engineering teams visibility into their services. These systems handle enormous volumes of traffic and have a direct impact on the reliability of the Braze platform.

As a Platform Software Engineer, you'll write code and run infrastructure at scale. You'll build the automation, tooling, and services that make these systems reliable, efficient, and easy for other engineers to use.

WHAT YOU’LL DO

Design, build, and operate Redis/Valkey, Memcached, Elasticsearch and Observability systems on Kubernetes with a focus on reliability, performance and cost efficiency at scale

Write the software that manages these systems: Kubernetes operators, automation, CLIs, and self-serve tooling that replace manual operations

Own and evolve Braze's observability platform, including metrics collection, long-term storage, alerting, and dashboards, so teams can detect and diagnose issues early

Lead capacity planning, upgrades, and migrations for large, stateful systems with minimal customer impact

Partner with engineering teams across Braze to improve how they use these systems, from data modeling and access patterns to query performance and monitoring

Investigate complex production issues across the stack and turn what you learn into durable fixes and better tooling

Contribute to the team's on-call rotation and help improve runbooks and write tooling that would help incident response

Document your designs and decisions clearly for a global team

WHO YOU ARE

5+ years of experience in software engineering, platform engineering, or SRE, with a strong record of writing production code (not only operating systems)

Hands-on experience running stateful, distributed systems in production at scale

Strong working knowledge of Kubernetes and infrastructure-as-code (Terraform, Pulumi) and comfort building on top of them

Proficiency in at least one general-purpose language such as Go, Python, or Ruby. We care more about design judgment than language choice

A systems mindset: you think about interfaces, failure modes, and edge cases, and you debug methodically when things break

A strong problem-solving mindset and a track record of working with other teams to untangle complex engineering problems

Clear written and verbal communication, and comfort collaborating across time zones

A bias toward fixing what's broken instead of working around it

NICE TO HAVE

Deep experience with Redis/Valkey, Elasticsearch, or similar systems, including performance tuning, replication, and failure handling

Experience with OpenTelemetry and observability tools such as Prometheus, Thanos, Grafana, or Datadog

Experience with Kafka or other streaming systems

Experience building developer-facing tooling or Kubernetes operators

A thoughtful approach to using AI tools in your own engineering workflow

Apply for this role →

← Back to all jobs