Thinking Machines Lab · Research · Unspecified · Posted 2026-08-21
Research, RL Scaling
Thinking Machines Lab · San Francisco · $350k–475k base
This range's midpoint is above 85% of posted research ranges at AI companies right now. See the salary index.
Apply on Thinking Machines Lab's site Watch Thinking Machines Lab for new roles
ABOUT THINKING MACHINES
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
ABOUT THE ROLE
Our team scales reinforcement learning for frontier models. Progress in RL is increasingly set by how well it scales: more rollouts, larger models, and training loops that keep large fleets of accelerators doing useful work. We are particularly interested in people working on high-training-compute, long-horizon RL. We believe the biggest gains come from designing the training recipe and the infrastructure together rather than separately, and we are hiring a researcher who wants to own that boundary.
A center of gravity for this role is asynchronous RL. Decoupling generation from training changes both the systems design and the learning problem, and doing it well requires a deep understanding of async RL algorithms, design choices, and trade-offs on both the ML and the systems sides. We expect much of the headroom in RL scaling to come from here.
Because generation dominates the cost of RL at scale, good knowledge of inference systems, low-precision numerics, and quantization is recommended: you should be able to reason quantitatively about rollout throughput and cost (batching, KV cache, MoE serving, speculative decoding) and about how inference constraints shape training design.
This is a research role with full-stack ownership, from the algorithms to the parallelism plan to the health of the run.
WHAT YOU’LL DO
- Co-design the RL recipe and the systems that run it: make recipe-level choices jointly with systems-level ones and validate them at frontier scale.
- Advance asynchronous RL algorithms.
- Improve the efficiency of rollout generation and its integration with training, treating inference as a first-class part of the RL loop.
- Run frontier-scale RL end to end: bring up new models and training setups, keep large runs stable and healthy.
- Jointly optimize the compute and training efficiency of RL: accelerator utilization, memory, communication, and low-precision numerics.
- Do careful empirical science: ablations and scaling studies backed by instrumentation you can trust, written up clearly.
SKILLS AND QUALIFICATIONS
Minimum qualifications:
- Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX). Comfortable with debugging distributed training and writing code that scales.
- Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
- Clarity in communication, an ability to explain complex technical concepts in writing.
- Strong research judgment: clean ablations, honest baselines, and clear technical writing.
Preferred qualifications — we encourage you to apply if you meet some but not all of these:
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.
- Strong grounding in RL for large language models, such as modern policy optimization methods and their behavior at scale.
- Deep understanding of asynchronous RL: the algorithms, design choices, and trade-offs, on both the ML and the systems sides.
- Experience training large models across many accelerators, with comfort inside the distributed stack (parallelism strategies, memory, communication).
- Good working knowledge of inference systems: able to reason quantitatively about rollout generation throughput and cost.
- Experience building or operating decoupled generation/training RL systems at scale.
- Experience with RL on verifiable and agent …
More research roles at Thinking Machines Lab
-
Web Crawling - Research Engineer
ResearchRemote US$350k–475k4d
-
Research, Mid Training
ResearchRemote US$350k–475k7d
-
Research, Tinker, RL Systems
ResearchRemote US$350k–475k12d
-
Research Lead, Tinker, Fine-tuning Science
ResearchLead / ManagerRemote US$350k–475k12d
-
Research, Finetuning Science
ResearchRemote US$350k–475k12d
See also: AI jobs in San Francisco Bay Area · Thinking Machines Lab salaries · Python jobs · PyTorch jobs · JAX jobs.
This listing is reproduced from Thinking Machines Lab's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All Thinking Machines Lab roles · AI salaries.