Jacob Hilton
Source-listed: VP of Research · Alignment Research Center (ARC)
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- Training language models to follow instructions with human feedbackSaved credit: Contributor · 2022Explore people connected to this work →
- WebGPT: Browser-assisted question-answering with human feedbackSaved credit: Contributor · 2021Explore people connected to this work →
- Backdoor defense, learnability and obfuscationSaved credit: Contributor · 2024Explore people connected to this work →
- Language models transmit behavioural traits through hidden signals in dataSaved credit: Contributor · 2026Explore people connected to this work →
- Estimating the expected output of wide random MLPs more efficiently than samplingSaved credit: Contributor · 2026Explore people connected to this work →