← People

Meg Tong

Source-listed: Works on research infrastructure; formal job title unknown · Anthropic

  • Language models

Selected work

5
  1. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingAuthor · 2025
  2. Forecasting Rare Language Model BehaviorsContributor · 2025
  3. Auditing language models for hidden objectivesContributor · 2025
  4. Towards Understanding Sycophancy in Language ModelsContributor · 2023
  5. Steering Llama 2 via Contrastive Activation AdditionContributor · 2024

Source-linked involvement

1
  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingAuthor · 2024