← People

Mrinank Sharma

  • Language models
  • Efficient ML

Selected work

5
  1. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingAuthor · 2025
  2. Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal JailbreaksAuthor · 2026
  3. Who's in Charge? Disempowerment Patterns in Real-World LLM UsageContributor · 2026
  4. Many-shot JailbreakingContributor · 2024
  5. Towards Understanding Sycophancy in Language ModelsContributor · 2023

Source-linked involvement

1
  1. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingAuthor · 2024