Find the people behind AI papers and projects.
/
Search public work, explore its connections, and check the sources. How coverage works
4,006 saved profiles
Matching saved profiles. Coverage and source-listed affiliations may be incomplete or historical.
Search results
1–53 of 53
- Ajay Pravin MahaleAffiliation: Not recordedWork: Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
- AnonRishAffiliation: Not recordedWork: Lumen: A Technical Alignment, Interpretability, and Governance Audit Framework
- Devina JainAffiliation: LambdaWork: Red-teaming with Mech-Interpretability
- Iuliia VitiugovaAffiliation: Not recordedWork: Mechanistic Interpretability of LLM Behaviours via Transcoders
- Justin ShawAffiliation: Not recordedWork: mechanistic-interpretability
- Neel NandaAffiliation: Google DeepMindWork: Progress measures for grokking via mechanistic interpretability
- Steven AbreuAffiliation: MakerMaker.AIWork: hybrid-interpretability
- Yatharth AnandAffiliation: ZipteamsWork: Mechanistic Interpretability of LLMs
- YuraloAffiliation: Not recordedWork: Mechanistic Interpretability
- Abhimanyu DubeyAffiliation: MetaWork: Scalable Interpretability via Polynomials
- Andy CoenenAffiliation: Google DeepMindWork: The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models
- Ben MannAffiliation: AnthropicWork: Scaling Laws and Interpretability of Learning from Repeated Data
- Carey RadebaughAffiliation: Google DeepMindWork: The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models
- Christopher OlahAffiliation: AnthropicWork: The Building Blocks of Interpretability
- Claudio BorileAffiliation: CENTAI InstituteWork: Mechanistic Interpretability
- Dawn DrainAffiliation: AnthropicWork: Scaling Laws and Interpretability of Learning from Repeated Data
- Dhruv MahajanAffiliation: Resolve AIWork: Neural Basis Models for Interpretability
- Dieuwke HupkesAffiliation: MetaWork: Interpretability of Language Models via Task Spaces
- Filip RadenovićAffiliation: MetaWork: Neural Basis Models for Interpretability
- Hao Henry ZhouAffiliation: Not recordedWork: Building Bayesian Neural Networks with Blocks: On Structure, Interpretability and Uncertainty
- Ian TenneyAffiliation: Google DeepMindWork: The Language Interpretability Tool (LIT)
- Jinwei XingAffiliation: Not recordedWork: Achieving efficient interpretability of reinforcement learning via policy distillation and selective input gradient regularization
- Kartik GargAffiliation: Birla Institute of Technology and Science, Pilani (BITS Pilani), K. K. Birla Goa CampusWork: LLM Anchoring Mechanistic Interpretability
- Lucas DixonAffiliation: Google DeepMindWork: Interpretability Illusions in the Generalization of Simplified Models
- Matthew RahtzAffiliation: Not recordedWork: Tracr: Compiled Transformers as a Laboratory for Interpretability
- Mohammadamin (Amin) BanayeeanzadeAffiliation: University of Southern CaliforniaWork: Mechanistic Interpretability of Emotion Inference in Large Language Models
- Ninad JoshiAffiliation: Tata Consultancy ServicesWork: Mechanistic-Interpretability-for-LLMs
- Santhosh JanakiramanAffiliation: Rutgers, The State University of New JerseyWork: Refusal Direction Experiments (LLM-Jailbreaks-Mechanistic-Interpretability)
- Siva ReddyAffiliation: ServiceNow AI ResearchWork: Post-hoc Interpretability for Neural NLP: A Survey
- Sunipa DevAffiliation: Google ResearchWork: The Geometry of Distributed Representations for Better Alignment, Attenuated Bias, and Improved Interpretability
- Thomas IcardAffiliation: GoodfireWork: Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
- Vladimir MikulikAffiliation: AnthropicWork: Tracr: Compiled Transformers as a Laboratory for Interpretability
- Nando de FreitasAffiliation: Microsoft AIWork: Genie: Generative Interactive Environments
- Yuzi HeAffiliation: Not recordedWork: The Llama 3 Herd of Models
- Zhufeng PanAffiliation: Not recordedWork: Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Tom HenighanAffiliation: Not recordedWork: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
- Leo GaoAffiliation: OpenAIWork: The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Hoagy CunninghamAffiliation: AnthropicWork: Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Chih-Kuan YehAffiliation: Not recordedWork: Gemini: A Family of Highly Capable Multimodal Models
- Tristan HumeAffiliation: AnthropicWork: Constitutional AI: Harmlessness from AI Feedback
- Aida AminiAffiliation: Google ResearchWork: Gemini: A Family of Highly Capable Multimodal Models
- Alvaro VidelaAffiliation: MicrosoftWork: Arithmetic Without Numbers (Draft) / Rune
- Kyle LevinAffiliation: Google DeepMindWork: Gemini: A Family of Highly Capable Multimodal Models
- Connor LeahyAffiliation: ControlAIWork: The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Ankit RamchandaniAffiliation: MetaWork: The Llama 3 Herd of Models
- Ayesha ImranAffiliation: ChipHubWork: Persona-Vector Routing: A Lightweight, Interpretable Guardrail for Mitigating LLM Hallucinations
- Roberto DessìAffiliation: Not DiamondWork: Toolformer: Language Models Can Teach Themselves to Use Tools
- Daniel P. MossingAffiliation: AnthropicWork: Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
- Linhao LuoAffiliation: Monash UniversityWork: G-reasoner: Foundation Models for Unified Reasoning over Graph-structured Knowledge
- Tom ConerlyAffiliation: AnthropicWork: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
- Yonatan BiskAffiliation: Carnegie Mellon UniversityWork: HellaSwag: Can a Machine Really Finish Your Sentence?
- Daniel D. McKinnonAffiliation: Gamow Labs, Inc.Work: Gemini: A Family of Highly Capable Multimodal Models
- Sri Priya PonnapalliAffiliation: Scale AIWork: Gemini: A Family of Highly Capable Multimodal Models