Research Engineer, RL Scaling Science

London, UKOn-site 1w ago

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role

Anthropic's RL Scaling Science team studies how reinforcement learning behaves as we scale it (across model size, compute, and task horizon) and turns that understanding into the training recipes behind our frontier models. As a Research Engineer on this team, you'll design and run large-scale experiments to understand and resolve bottlenecks, build the benchmarks that make long-horizon progress measurable, and ship validated findings directly into production training.

This role lives at the boundary between research and engineering. The problems are open, the experiments run at frontier scale, and the path from a robust result to production is short.

Key responsibilities

Design, run, and interpret large-scale RL experiments, reasoning rigorously about what the data does and doesn't show
Investigate how RL improves as horizon, compute, and model size grow
Build and maintain benchmarks for long-horizon RL so progress is measurable and reproducible
Translate validated findings into production training recipes, exercising judgment about when a result is robust enough to ship
Debug complex issues at the seam where research meets infrastructure - failures that only appear at scale
Partner closely with adjacent RL teams across research and engineering and advance our overall RL stack

Minimum qualifications

Strong empirical research skills in Reinforcement Learning, large-scale ML training, or a closely adjacent area
Demonstrated ability to own large experiments end-to-end, from design through interpretation
Proficiency in Python and experience working with large-scale or distributed ML systems
Comfort operating at the research/systems boundary, including debugging where the two meet
Care about the societal impacts of AI and responsible scaling

Preferred qualifications

Published or shipped work in long-horizon RL or RL fundamentals

Apply now

Opens the company's application page

About the company

Anthropic

AI safety company.

All open roles Visit website

Listed via

Greenhouse

Similar roles

Sr. Customer Support Engineer, Raipur

Danaher

IndiaRemote

Collibra Platform Developer (Mid to Senior)

Arch Capital Group Ltd.

PhilippinesRemote

Scheduling Director (Renewables Construction)

MasTec Industrial

United StatesRemote

Mom and Baby Care Manager - RN - Must reside in Nevada

CareSource

United StatesRemote

Design & Tech

Related reads from TCHNX

View all →

Technology

The Quiet Revolution in Local-First Software

As major platforms face outages and data breaches, a new generation of developers is building applications that prioritise local data storage and peer-to-peer sync, challenging the cloud-first orthodoxy that's dominated tech for two decades.

tchnx.com

Products

The Return of Physical Controls: Why Haptic Feedback Is Reshaping Digital Interfaces

After years of pursuing flat, buttonless designs, tech companies are rediscovering the value of tactile interaction. A new wave of products proves that touching isn't just feeling it's understanding.

tchnx.com

Design

The Quiet Revolution of Parametric Design Tools in Everyday Products

Parametric design is migrating from architecture studios to consumer products. As tools democratize and manufacturers adopt flexible production, we're entering an era of mass customization that challenges fundamental assumptions about design.

tchnx.com