← People

Cem Anil

Source-listed: Research Scientist · Anthropic

  • Language models
  • Robotics
  • Evaluation

Selected work

5
  1. Modular Pretraining Enables Access ControlContributor · 2026
  2. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingAuthor · 2025
  3. Many-shot Jailbreaking: A Large-context Attack Surface for Language ModelsContributor · 2024
  4. Sabotage Evaluations for Frontier ModelsContributor · 2024
  5. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingAuthor · 2024