← People

Devina Jain

Source-listed: Lambda

  • Language models
  • Interpretability
  • Evaluation

Selected work

5
  1. Red-teaming with Mech-InterpretabilityMaintainer · 2025
  2. From Scores to Trust: The Benchmark Confidence Rubric for LLMsContributor · 2025
  3. Red-teaming with Persuasive RephrasingMaintainer · 2025
  4. One Probe Won't Catch Them All: Towards Targeted Deception DetectionContributor · 2026
  5. Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent SecurityContributor · 2026