Senior Software Engineer, Machine Learning Platform
About the role
Chime’s Machine Learning Platform (MLP) team builds and operates the infrastructure, tooling, and developer experience that powers machine learning across the company. We enable data scientists and ML engineers to develop, train, deploy, and monitor models reliably and efficiently.
As a Senior Software Engineer on the Machine Learning Platform team, you will design and build scalable systems spanning traditional machine learning and emerging AI workloads, including model training, feature computation, real-time inference, foundation-model access, evaluation, and agentic orchestration. You’ll work at the intersection of distributed systems, cloud infrastructure, applied machine learning, and AI product engineering.
This role focuses on creating secure, reliable, and reusable platform capabilities that help teams choose the right approach, from conventional predictive models to LLM-powered and multi-step agentic systems, while maintaining strong standards for evaluation, observability, governance, privacy, and cost efficiency.
The base salary offered for this role and level of experience will begin at $187,000.00 and up to $259,000.00. Full-time employees are also eligible for a bonus, competitive equity package, and benefits. The actual base salary offered may be higher, depending on your location, skills, qualifications, and experience.
In this role, you can expect to
Design, build, and operate scalable ML and AI infrastructure on AWS.
Design and operate shared platform capabilities for LLM and agentic workloads, including model access, prompt and configuration lifecycle, retrieval, tool integration, state management, and workflow orchestration.
Build evaluation frameworks for non-deterministic AI systems, including offline benchmarks, regression testing, online quality signals, human feedback, and failure analysis.
Establish observability, reliability, and governance for models and agents, covering traces, model and prompt versions, tool calls, latency, token usage, quality, safety, privacy, and cost.
Help teams make principled architecture decisions across traditional ML, LLM-powered applications, and agentic workflows, and contribute to the platform’s technical roadmap.
Build distributed training, batch inference, and large-scale processing systems using frameworks such as Ray or Spark.
Build and maintain infrastructure as code using Terraform.
Support and evolve the feature store and feature pipelines.
Develop data ingestion and streaming systems using technologies such as Kinesis, Kafka, Flink, or Spark.
Improve CI/CD workflows for ML models, AI applications, and platform components.
Partner closely with Data Science and ML Engineering teams to improve developer experience.
Participate in on-call rotations to support production systems.
To thrive in this role, you have
Knowledge of the machine learning development lifecycle, including data preprocessing, model training, evaluation, deployment, and monitoring.
Experience designing distributed systems and large-scale data or compute platforms on AWS using frameworks such as Spark or Ray.
5+ years of experience in ML or AI infrastructure, platform engineering, distributed systems, or production ML systems.
Working knowledge of LLM application patterns such as retrieval-augmented generation, structured outputs, tool calling, agent orchestration, and evaluation of non-deterministic systems.
Experience designing production systems that integrate ML or foundation models through reliable APIs, workflows, and data contracts.
Hands-on experience with CI/CD pipelines, DevOps practices, and infrastructure as code.
Experience with containerization and orchestration technologies such as Docker and Kubernetes.
Strong programming skills in Python, Go, Scala, Java, or similar languages.
Solid understanding of software engineering fundamentals, including testing, version control, code review, and observability.
Nice-to-have
Experience shipping LLM-powered or agentic systems to production.
Experience with one or more of the following: model gateways, prompt lifecycle management, retrieval or vector search, tool execution, and agent orchestration frameworks.
Experience building evaluation, tracing, and observability capabilities for non-deterministic AI systems.
Familiarity with managed or self-hosted foundation model infrastructure, such as Amazon Bedrock, SageMaker, or equivalent platforms.
Experience operating GPU-based workloads and optimizing training or inference performance and cost; CUDA experience is a plus.
#LI-Onsite #LI-WW1