Software Engineer, Site Reliability

Upstart · United States | Remote · Engineering

Posted 2026-08-22

Apply for this role →

The Team

Upstart’s Site Reliability Engineering team enables engineers to operate reliable, resilient, and observable production systems at scale. We build the platforms, tooling, automation, and operational practices that help teams understand system health, respond effectively when things go wrong, and continuously improve the reliability of the services they own.

Our goal is to make reliability an integrated part of how software operates at Upstart. We provide shared observability and reliability capabilities, improve incident response and operational readiness, automate recurring operational work, and use system and customer data to identify where reliability investments will have the greatest impact.

SRE partners closely with product engineering, Cloud Platform, Delivery, Developer Platform, Security, and other infrastructure teams. SRE provides shared reliability capabilities and operational practices, while engineering teams remain accountable for the reliability and operation of the services they build.

The Role

As a Software Engineer on the Site Reliability Engineering team, you will build and operate systems that improve the reliability, resiliency, and observability of Upstart’s production environment.

You will work across observability platforms, reliability tooling, incident response systems, operational automation, and resiliency capabilities. You will independently deliver well scoped engineering projects, contribute to technical design, and use production data and operational experience to improve systems used across engineering.

We are looking for engineers who are thoughtful and intentional about how AI changes software development and operations. You should be comfortable using AI throughout the engineering lifecycle, including understanding unfamiliar systems, investigating production behavior, planning implementation, accelerating development, validating changes, and automating repetitive work. We expect engineers to continually develop more effective ways of working with increasingly capable AI tools and to apply sound engineering judgment to where they provide the most leverage.

How you’ll make an impact

Build and improve the tooling, services, and automation that help engineers understand and improve the reliability of production systems

Develop shared observability capabilities that make metrics, logs, traces, service health, and customer impact easier to understand and act on

Improve incident response and operational readiness through better tooling, automation, standards, and actionable production signals

Build resiliency capabilities that help teams identify failure modes, reduce operational risk, and recover effectively from infrastructure or application failures

Identify recurring operational toil and reliability problems and replace manual processes with durable software and automation

Use AI as an integrated part of software development and operational problem solving, while identifying opportunities for AI enabled capabilities that improve incident investigation, observability, reliability, and engineering efficiency

Minimum Qualifications

3+ years of professional experience in software engineering, site reliability engineering, or a related engineering discipline

Strong software development skills in one or more general purpose programming languages such as Python, Go, JavaScript, or TypeScript

Experience designing, building, testing, and operating production software, internal tooling, or infrastructure

Experience with cloud infrastructure, distributed systems, observability, monitoring, or production operations

Experience participating in on call or incident response for production systems

Demonstrated ability to independently deliver well scoped engineering projects, navigate technical ambiguity, and collaborate effectively across engineering teams

Demonstrated experience using AI assisted development tools across multiple stages of the software engineering lifecycle, with an interest in continually evolving how you use these tools as their capabilities advance

Preferred Qualifications

Experience with Kubernetes, AWS, infrastructure as code, and cloud native production environments

Experience building internal reliability, observability, incident management, or operational automation tools

Experience with observability platforms such as Datadog, Sumo Logic, CloudWatch, or similar technologies

Experience with reliability practices such as service level objectives, capacity planning, resiliency testing, disaster recovery, or operational readiness

Experience operating distributed applications with complex dependencies and high availability requirements

Experience building AI enabled operational workflows, tools, or automation that extend beyond individual code generation

Position location This role is available in the following locations: Remote

Travel requirements As a digital first company, the majority of your work can be accomplished remotely. The majority of our employees can live and work anywhere in the U.S but are encouraged to to still spend high quality time in-person collaborating via regular onsites. The in-person sessions’ cadence varies depending on the team and role; most teams meet once or twice per quarter for 2-4 consecutive days at a time.

#LI-REMOTE

#LI-Associate

At Upstart, your base pay is one part of your total compensation package.  The anticipated base salary for this position is expected to be within the below range. Your actual base pay will depend on your geographic location–with our “digital first” philosophy, Upstart uses compensation regions that vary depending on location. Individual pay is also determined by job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your preferred location during the hiring process.

In addition, Upstart provides employees with target bonuses, equity compensation, and generous benefits packages (including medical, dental, vision, and 401k).

United States | Remote - Anticipated Base Salary Range

$142,000—$196,600 USD

Apply for this role →

← Back to all jobs