Meg Tong
Source-listed: Works on research infrastructure; formal job title unknown · Anthropic
- Language models
Selected work
5- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingAuthor · 2025
- Forecasting Rare Language Model BehaviorsContributor · 2025
- Auditing language models for hidden objectivesContributor · 2025
- Towards Understanding Sycophancy in Language ModelsContributor · 2023
- Steering Llama 2 via Contrastive Activation AdditionContributor · 2024