← People

Hoagy Cunningham

Source-listed: Anthropic

  • Language models
  • Interpretability
  • Efficient ML

Selected work

5
  1. Sparse Autoencoders Find Highly Interpretable Features in Language ModelsContributor · 2023
  2. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 SonnetContributor · 2024
  3. Auditing Language Models for Hidden ObjectivesContributor · 2025
  4. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingAuthor · 2025
  5. Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal JailbreaksAuthor · 2026