Staff Machine Learning Engineer, ML Platform
WHAT YOU'LL DO
Braze is seeking a Staff Machine Learning Engineer to join our Predictive and Generative AI (PGAI) team. The team's mission is to deliver a truly engaging and personalized customer experience through the creation of ML and AI enhanced marketing solutions. We run those solutions as production systems at global scale, from the distributed pipelines that train models for each customer to the high-throughput APIs that serve predictions into our messaging systems across multiple regions. You will own the platform underneath, and you will make deploying, operating, and scaling ML at Braze fast, safe, and efficient.
As the Staff Engineer on the team, you will:
Identify and drive the transformative initiatives that change how the team runs ML in production, whether that's replatforming our queueing and orchestration, overhauling deployment and cloud identity, or retiring a generation of infrastructure
Build and ship at high velocity. Staff at Braze is a hands-on delivery role; you carry the most complex infrastructure initiatives yourself from design through production. Current examples include multi-region model serving fleets, the pipelines that keep hundreds of customer-specific models healthy, and the CI and deployment tooling that moves it all safely
Own the platform's technical vision and production quality bar. Set direction for how models are trained, deployed, served, and observed; lead incident response for ML systems; and drive the reliability and cost work that keeps the platform efficient at scale
Drive initiatives that span teams. Our platform builds on shared infrastructure, deployment tooling, and data systems owned with partner teams, and you carry the technical relationships with those teams
Raise the team's engineering quality through design review, code review, and production readiness for ML systems, and mentor other senior engineers and data scientists
Connect technical decisions to customer and business outcomes, and represent the team's technical perspective to product and engineering leadership
WHO YOU ARE
8+ years building and operating distributed systems in production, with depth in deployment and operations. You have designed services for scale and reliability, owned CI/CD and infrastructure as code, and run what you built under production load
Hands-on experience with ML workloads in production. Training pipelines, model serving, feature systems, or ML platform tooling all count; deep modeling experience is a plus rather than a requirement
A technical leader who has owned direction for a team, led multi-quarter initiatives across team boundaries, and grown senior engineers, all while keeping a high personal output
Deep working knowledge of Kubernetes and cloud infrastructure, including identity and access management, networking, and the cost profile of what you run
An effective communicator, both verbal and written, whose designs and recommendations build consensus and drive forward decision making
Bonus:
Queueing and orchestration systems such as Celery, RabbitMQ, Kafka, or Ray
ML platform tooling such as MLflow or another model registry, feature stores, or ML observability
Experience in our stack (Python, Ruby on Rails, MongoDB, Redis, Kubernetes)
Operating under compliance regimes such as SOX or HIPAA
Customer engagement, personalization, or marketing technology domain experience
For candidates based in the United States, the pay range for this position at the start of employment is expected to be between $184,000 and $314,000/year, with an expected On Target Earnings (OTE) between $204,000 and $348,000/year (including bonus or commission). Your exact offer may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience. In addition to cash compensation, this role qualifies for a comprehensive Total Rewards package that includes equity grants of restricted stock (RSUs) so that you will own a piece of our company.
#LI-Hybrid