Senior Platform Software Engineer I
Braze is the leading customer engagement platform, helping brands build personalized messaging experiences across push, email, in-app, and more at massive scale. Our technology stack is rooted in Ruby on Rails, MongoDB, Redis, Kafka, Kubernetes and more.
The In-Memory Databases & Observability team provides the foundational infrastructure and tools that keep Braze fast, reliable, and observable. Valkey sits at the heart of Braze: every message we send flows through it and our fleet serves around 36 million operations per second while holding terabytes of data in memory. We operate Valkey, Memcached, and Elasticsearch, along with the observability stack (Prometheus, Thanos, and related tooling) that gives engineering teams visibility into their services. These systems handle enormous volumes of traffic and have a direct impact on the reliability of the Braze platform.
As a Platform Software Engineer, you'll write code and run infrastructure at scale. You'll build the automation, tooling, and services that make these systems reliable, efficient, and easy for other engineers to use.
WHAT YOU’LL DO
Design, build, and operate Redis/Valkey, Memcached, Elasticsearch and Observability systems on Kubernetes with a focus on reliability, performance and cost efficiency at scale
Write the software that manages these systems: Kubernetes operators, automation, CLIs, and self-serve tooling that replace manual operations
Own and evolve Braze's observability platform, including metrics collection, long-term storage, alerting, and dashboards, so teams can detect and diagnose issues early
Lead capacity planning, upgrades, and migrations for large, stateful systems with minimal customer impact
Partner with engineering teams across Braze to improve how they use these systems, from data modeling and access patterns to query performance and monitoring
Investigate complex production issues across the stack and turn what you learn into durable fixes and better tooling
Contribute to the team's on-call rotation and help improve runbooks and write tooling that would help incident response
Document your designs and decisions clearly for a global team
WHO YOU ARE
5+ years of experience in software engineering, platform engineering, or SRE, with a strong record of writing production code (not only operating systems)
Hands-on experience running stateful, distributed systems in production at scale
Strong working knowledge of Kubernetes and infrastructure-as-code (Terraform, Pulumi) and comfort building on top of them
Proficiency in at least one general-purpose language such as Go, Python, or Ruby. We care more about design judgment than language choice
A systems mindset: you think about interfaces, failure modes, and edge cases, and you debug methodically when things break
A strong problem-solving mindset and a track record of working with other teams to untangle complex engineering problems
Clear written and verbal communication, and comfort collaborating across time zones
A bias toward fixing what's broken instead of working around it
NICE TO HAVE
Deep experience with Redis/Valkey, Elasticsearch, or similar systems, including performance tuning, replication, and failure handling
Experience with OpenTelemetry and observability tools such as Prometheus, Thanos, Grafana, or Datadog
Experience with Kafka or other streaming systems
Experience building developer-facing tooling or Kubernetes operators
A thoughtful approach to using AI tools in your own engineering workflow