← People

Euan Ong

Source-listed: Anthropic

  • Multimodal AI
  • Language models
  • Robotics

Selected work

5
  1. Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual ExperimentsContributor · 2026
  2. Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation ExplainersContributor · 2025
  3. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingAuthor · 2025
  4. Image Hijacks: Adversarial Images can Control Generative Models at RuntimeContributor · 2024
  5. Successor Heads: Recurring, Interpretable Attention Heads In The WildContributor · 2024