← People

Logan Graham

Source-listed: Head of the Frontier Red Team · Anthropic

  • Language models

Selected work

5
  1. Assessing Claude Mythos Preview’s cybersecurity capabilitiesContributor · 2026
  2. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingAuthor · 2025
  3. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingAuthor · 2024
  4. MultiVerse: Causal Reasoning using Importance Sampling in Probabilistic ProgrammingContributor · 2020
  5. Inferring Work Task Automatability from AI Expert EvidenceContributor · 2019