Devina Jain
Source-listed: Lambda
- Language models
- Interpretability
- Evaluation
Selected work
5- Red-teaming with Mech-InterpretabilityMaintainer · 2025
- From Scores to Trust: The Benchmark Confidence Rubric for LLMsContributor · 2025
- Red-teaming with Persuasive RephrasingMaintainer · 2025
- One Probe Won't Catch Them All: Towards Targeted Deception DetectionContributor · 2026
- Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent SecurityContributor · 2026