Software Engineer, Tokens and Prompt Structures
About the Role:
The Encodings Infra team maintains the libraries that engineers and researchers across Anthropic use to encode text and multimodal data into a form that Claude can consume. It also determines Claude’s prompt shape: how a user’s turn is represented to the model, how Claude calls tools and receives tool results, and so on.
As a Software Engineer on this team, you'll own the design and maintenance of these libraries—keeping their APIs intuitive, their performance sharp, and their abstractions solid enough that most of the org never has to think about encodings or prompt structures at all. You’ll have the satisfaction of knowing that your work enabled Claude to learn new ways of understanding the world.
This role is unusually broad: your work will touch systems across the codebase, from pretraining to finetuning to the API, and you'll collaborate closely with both researchers and engineers to make sure new encoding ideas can move quickly from experiment to production.
Responsibilities:
Maintain and improve the encoding libraries used by engineers and researchers across Anthropic
Run experiments to determine the optimal way to feed structured data into Claude without confusing it
Design data structures and abstractions that shield most of the organization from the details of how encoded data works while enabling “power users”
Adapt the encoding libraries to support new research directions as they emerge, and make sure that we can ship these research ideas to production
Optimize encoding performance across the systems that depend on these libraries
You may be a good fit if you:
Have 5+ years of software engineering experience, with meaningful time spent maintaining libraries, SDKs, or developer-facing APIs
Have familiarity with ML terminology and LLM architecture — you don't need to be an ML expert, but enough understanding to work effectively alongside researchers and understand the results of experiments
Have experience carrying out complex refactors in large codebases
Have strong communication skills and enjoy working closely with researchers and engineers to understand what they need
Are results-oriented, with a bias towards flexibility and impact
Pick up slack, even if it goes outside your job description
Care about the societal impacts of your work
Strong candidates may also have experience with:
Tokenizers or other text/data encoding systems
Maintaining a widely-used library over a long period of time
Performance optimization
Python and/or Rust
Reinforcement learning or model training infrastructure
Representative projects:
Working with a research team to ship a new multimodal data type (audio, video, etc) to production
Redesigning a core abstraction so that we can change how data is encoded into Claude without breaking downstream teams
The annual compensation range for this role is listed below.
For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.
Annual Salary:
$320,000—$405,000 USD