Staff Research Scientist

Turing · United States · Engineering

Posted 2026-09-24

Apply for this role →

The Role

Turing is seeking exceptional Research Scientists to join our research organization and develop new ways to evaluate, train, and improve frontier AI systems.

This is a research-first role focused on problems where the right benchmark, dataset, or methodology often does not yet exist. You will identify important gaps in the literature, propose ambitious new research directions, and take projects from initial hypothesis through experimentation, benchmark construction, and publication.

Our research is deliberately focused on frontier evaluation, synthetic data, hallucination and reliability, and agentic science. We are looking for scientists who can recognize important problems early, formulate them precisely, and design rigorous research programs to answer them.

What You'll Do

Frontier benchmarks and evaluation

Identify high-impact gaps in existing benchmark and evaluation literature.

Design novel benchmarks in and across STEM fields and on general model functionality.

Develop evaluations for emerging model capabilities that are poorly captured by traditional static benchmarks.

Design rigorous task-generation, grading, contamination-control, difficulty-calibration, and validation methodologies.

Build benchmarks that can become both valuable research contributions and meaningful standards for evaluating frontier models.

Synthetic data and post-training

Develop methods for generating high-quality synthetic STEM training data.

Study how task selection, difficulty, diversity, verification, filtering, and data quality affect downstream performance.

Explore methods for generating useful training signal in domains where expert human data is scarce or expensive.

Design experiments that determine when synthetic data genuinely improves capabilities rather than simply increasing training volume.

Hallucination, reliability, and verification

Study hallucination, uncertainty, calibration, and epistemic failure in technical domains.

Develop evaluations and methods for improving factual reliability, self-correction, verification, citation, and appropriate abstention.

Investigate when models should reason internally, invoke tools, seek external evidence, or recognize that they do not know.

Agentic science

Research AI systems capable of performing extended scientific and technical work.

Develop workflows involving literature search, coding, simulation, tool use, experimentation, verification, and iterative reasoning.

Evaluate long-horizon scientific agents and identify the bottlenecks preventing them from reliably performing real research.

Explore new approaches to human-AI and multi-agent scientific collaboration.

New research directions

The areas above are our core focus, not an exhaustive list. Researchers will also have significant latitude to propose new programs in areas such as reasoning, model evaluation, AI-for-science, data generation, and emerging capabilities.

What We’re Looking For

PhD or equivalent research experience in machine learning, computer science, mathematics, physics, chemistry, biology, engineering, statistics, or another highly technical field.

Demonstrated ability to formulate and execute original research.

Strong understanding of modern LLMs and the frontier AI research landscape.

Excellent experimental design, quantitative reasoning, and scientific judgment.

Ability to rapidly understand unfamiliar technical literature and develop expertise in new areas.

Strong Python skills and the ability to independently build research prototypes and evaluation pipelines.

Excellent technical writing and communication.

Comfort working in a fast-moving environment where the research agenda evolves with the frontier.

A strong publication record is valuable, but we care most about whether you can identify important questions, design rigorous ways to answer them, and execute quickly enough for the results to matter.

A note from our CTO, Ece Kamar

Some of the most important questions in AI can't be answered from inside a single lab: how to evaluate frontier models, how data and RL truly drive capability, and what happens when AI meets the real world. At Turing, we work directly with frontier labs, the academic community, and enterprises deploying AI at scale. That creates a feedback loop where deployment shapes our research and our research shapes the field, and we publish our findings, datasets, and benchmarks openly. Join our talented team at Turing to push the frontier of AI with rigorous science that matters.

Why Turing

Work directly with leading AI labs and enterprises at the frontier of post-training and RL environment design.

Build datasets and environments that directly improve the capabilities of advanced AI systems.

Help advance coding agents’ ability to understand, plan, and execute complex software-engineering tasks.

Apply frontier AI innovations to high-value enterprise workflows.

Operate with high autonomy, rapid iteration, and meaningful commercial impact.

Collaborate with exceptional colleagues from organizations including Google, Meta, Amazon, and other leading technology companies.

Contribute to research that may be shared through technical publications and leading conferences such as ICLR, ICML, and NeurIPS.

Compensation: $250,000 to $400,000 OTE + Equity

Values

Apply for this role →

← Back to all jobs