Euan Ong
Source-listed: Anthropic
- Multimodal AI
- Language models
- Robotics
Selected work
5- Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual ExperimentsContributor · 2026
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation ExplainersContributor · 2025
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingAuthor · 2025
- Image Hijacks: Adversarial Images can Control Generative Models at RuntimeContributor · 2024
- Successor Heads: Recurring, Interpretable Attention Heads In The WildContributor · 2024