Jonathan Uesato
Source-listed: Anthropic
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- Natural Emergent Misalignment from Reward Hacking in Production RLSaved credit: Contributor · 2025Explore people connected to this work →
- Reasoning Models Don't Always Say What They ThinkSaved credit: Contributor · 2025Explore people connected to this work →
- Alignment faking in large language modelsSaved credit: Contributor · 2024Explore people connected to this work →
- Gemini: A Family of Highly Capable Multimodal ModelsSaved credit: Contributor · 2023Explore people connected to this work →
- Solving math word problems with process- and outcome-based feedbackSaved credit: First author · 2022Explore people connected to this work →