Elizabeth Barnes
Source-listed: Founder, CEO · Model Evaluation and Threat Research (METR)
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- Evaluating Large Language Models Trained on CodeSaved credit: Contributor · 2021Explore people connected to this work →
- Evaluating Language-Model Agents on Realistic Autonomous TasksSaved credit: Contributor · 2023Explore people connected to this work →
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human expertsSaved credit: Contributor · 2024Explore people connected to this work →
- Measuring AI Ability to Complete Long Software TasksSaved credit: Contributor · 2025Explore people connected to this work →
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer ProductivitySaved credit: Contributor · 2025Explore people connected to this work →