← People

Jan Leike

Source-listed: Lead · Anthropic

Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.

Selected work

5
  1. Deep reinforcement learning from human preferencesSaved credit: Contributor · 2017Explore people connected to this work →
  2. Scalable agent alignment via reward modeling: a research directionSaved credit: First author · 2018Explore people connected to this work →
  3. Training language models to follow instructions with human feedbackSaved credit: Contributor · 2022Explore people connected to this work →
  4. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionSaved credit: Contributor · 2023Explore people connected to this work →
  5. Automated Weak-to-Strong ResearcherSaved credit: Contributor · 2026Explore people connected to this work →