Research Engineer, Takeoff Intel

Anthropic · Remote-Friendly (Travel Required) | San Francisco, CA · Engineering

Posted 2026-09-08

Apply for this role →

About the team

At Anthropic, we are delegating a growing share of AI development to AI systems themselves. Takeoff Intel is the team that measures this recursion from the inside. We're part of the Anthropic Institute. We design evaluations of AI R&D capabilities, build the internal telemetry Anthropic uses to track how much of its own model development is becoming AI-assisted, and develop the quantitative methods that turn those signals into a calibrated picture of where capability growth is heading, so that Anthropic and the wider world have accurate situational awareness on this acceleration.

Our work appears in Anthropic's model system cards (we own the AI R&D capability assessments and adapted Epoch's Capabilities Index to our evals); all the data in When AI Builds Itself comes from our team. Internally, our measurements shape research priorities and safety planning; externally, they contribute to Anthropic's public reporting on the pace of AI progress and to collaborations with third-party evaluators. We're a small team that works closely with pretraining, RL, economics, and policy researchers across the company. If you're passionate about measurement accuracy, and feel urgency about safety and situational awareness, you should consider joining us.

About the role

As a Research Engineer on Takeoff Intel you'll build and run the evaluation and measurement instruments that make this research possible. This is a generalist role on a small team: you'll work across evals infrastructure, large-scale data processing, and analysis tooling, and you'll prioritize shipping. We build instruments that answer real questions and help set priorities, not dashboards that surface noise. We value working prototypes, rapid iteration, accuracy and good prioritization. We often need to go from a vague research question to a running instrument quickly.

We're hiring at both junior and senior levels.

Responsibilities

Design, build, and run capability evaluations and measurement instruments at scale

Build the data and analysis pipelines that turn large volumes of model outputs and telemetry into reliable metrics

Prototype new instruments fast, validate them, and decide what to keep

Review and supervise AI-written code as a normal part of the workflow

Work closely with research scientists on the team and with partner teams to define what's worth measuring

Contribute to internal write-ups and public reporting

You may be a good fit if you

Have shipped an evaluation, data product, or research library end to end

Prototype fast and are comfortable throwing code away

Handle messy, large-volume data without over-engineering

Have run experiments on large language models, not just moved their outputs around

Can work from a vague question rather than a spec

Communicate results clearly and collaborate closely with the researchers whose questions your instruments answer

Strong candidates may also have

Built evaluation harnesses or benchmark infrastructure for LLMs

Experience with large-scale ML or data infrastructure (self-driving, observability, or similar) alongside ML exposure

Built tools or libraries that other researchers rely on

A track record of catching what AI-written code gets wrong

Some examples of our work

Anthropic ECI:  our adaptation of Epoch Capabilities Index published in all recent system cards to measure capability acceleration

AI R&D capability assessments in the Claude system cards

When AI Builds Itself: all data in the article comes from our team

The annual compensation range for this role is listed below.

For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.

Annual Salary:

$350,000—$850,000 USD

Apply for this role →

← Back to all jobs