Cem Anil
Source-listed: Research Scientist · Anthropic
- Language models
- Robotics
- Evaluation
Selected work
5- Modular Pretraining Enables Access ControlContributor · 2026
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingAuthor · 2025
- Many-shot Jailbreaking: A Large-context Attack Surface for Language ModelsContributor · 2024
- Sabotage Evaluations for Frontier ModelsContributor · 2024
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingAuthor · 2024