Find the people behind AI papers and projects.
/
Search public work, explore its connections, and check the sources. How coverage works
4,006 saved profiles
Matching saved profiles. Coverage and source-listed affiliations may be incomplete or historical.
Search results
1–80 of 663
- Charles FosterAffiliation: METR (Model Evaluation and Threat Research)Work: The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Elizabeth BarnesAffiliation: Model Evaluation and Threat Research (METR)Work: Evaluating Large Language Models Trained on Code
- Jeffrey WuAffiliation: AI Verification and Evaluation Research Institute (AVERI)Work: Language Models are Unsupervised Multitask Learners
- Miles BrundageAffiliation: AI Verification and Evaluation Research Institute (AVERI)Work: Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
- Alexei RobskyAffiliation: MicrosoftWork: Reliable Evaluations for LLMs and AI Agents: End-to-End Evaluation Frameworks for LLMs and Autonomous AI Agents
- Ananya KumarAffiliation: MetaWork: Holistic Evaluation of Language Models
- Axel StjerngrenAffiliation: DeepMindWork: Bifrost: End-to-End Evaluation and Optimization of Reconfigurable DNN Accelerators
- Brian IsraelAffiliation: AnthropicWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Bryan SeethorAffiliation: AnthropicWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Ce ZhangAffiliation: University of Chicago; Together AI; INSAITWork: Holistic Evaluation of Language Models (HELM)
- Chen ElkindAffiliation: Google ResearchWork: TrueTeacher: Learning Factual Consistency Evaluation with Large Language Models
- Chenel ElkindAffiliation: Google ResearchWork: TrueTeacher: Learning Factual Consistency Evaluation with Large Language Models
- Christian CosgroveAffiliation: Not recordedWork: Holistic Evaluation of Language Models
- Clément CrepyAffiliation: Not recordedWork: MIND: Monge Inception Distance for Generative Models Evaluation
- Craig PettitAffiliation: Not recordedWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Da YanAffiliation: AnthropicWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Dan MalkinAffiliation: GoogleWork: Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
- Daniela AmodeiAffiliation: AnthropicWork: Discovering Language Model Behaviors with Model-Written Evaluations
- David W. Sculley IIAffiliation: Not recordedWork: Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation
- Deepak NarayananAffiliation: NVIDIAWork: Holistic Evaluation of Language Models
- Dilara SoyluAffiliation: Stanford UniversityWork: Holistic Evaluation of Language Models
- Dimitris TsiprasAffiliation: OpenAIWork: Holistic Evaluation of Language Models
- Dominik KrzemińskiAffiliation: Arm Ltd.Work: Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
- Edwin ChenAffiliation: Surge AIWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Esin DurmusAffiliation: AnthropicWork: Holistic Evaluation of Language Models
- Faisal LadhakAffiliation: NVIDIAWork: Holistic Evaluation of Language Models (HELM)
- Felipe Maia PoloAffiliation: Not recordedWork: Rich Insights from Cheap Signals: Efficient Evaluations via Tensor Factorization
- Guro KhundadzeAffiliation: AnthropicWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Hailey SchoelkopfAffiliation: Not recordedWork: Language Model Evaluation Harness
- Hongyu RenAffiliation: Not recordedWork: Holistic Evaluation of Language Models (HELM)
- Huaxiu YaoAffiliation: University of North Carolina at Chapel HillWork: Holistic Evaluation of Language Models
- James LandisAffiliation: Not recordedWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Jeeyoon HyunAffiliation: Not recordedWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Jiaju LinAffiliation: The Pennsylvania State UniversityWork: AgentSims: An Open-Source Sandbox for Large Language Model Evaluation
- Jue WangAffiliation: Together AIWork: Holistic Evaluation of Language Models
- Kamile LukosuiteAffiliation: Not recordedWork: Saved link: Discovering Language Model Behaviors with Model-Written Evaluations
- Keshav SanthanamAffiliation: NVIDIAWork: Holistic Evaluation of Language Models
- Kiran VodrahalliAffiliation: Google DeepMindWork: Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries
- Landon GoldbergAffiliation: Not recordedWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Laurel J. OrrAffiliation: StacklokWork: Holistic Evaluation of Language Models (HELM)
- Lisa Anne HendricksAffiliation: Google DeepMindWork: CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
- Madelaine BoydAffiliation: Not recordedWork: Evaluation of Filesystem Provenance Visualization Tools
- Maribeth RauhAffiliation: Trinity College DublinWork: Gaps in the Safety Evaluation of Generative AI
- Martin LucasAffiliation: Machine Intelligence Research Institute (MIRI)Work: Discovering Language Model Behaviors with Model-Written Evaluations
- Matan EyalAffiliation: Not recordedWork: ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer
- Michihiro YasunagaAffiliation: Not recordedWork: Holistic Evaluation of Language Models
- Miranda ZhangAffiliation: AnthropicWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Mohammadamin (Amin) BanayeeanzadeAffiliation: University of Southern CaliforniaWork: Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
- Nan DingAffiliation: Not recordedWork: Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluations
- Nandan ThakurAffiliation: Microsoft Research IndiaWork: BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
- Nathan KimAffiliation: Unify (Unify GTM)Work: Holistic Evaluation of Language Models
- Nathan ScalesAffiliation: Not recordedWork: Evaluation of retrieval-based QA on QUEST-LOFT
- Nathanael SchärliAffiliation: Google DeepMindWork: Evaluation of retrieval-based QA on QUEST-LOFT
- Neel GuhaAffiliation: Columbia Law School, Columbia UniversityWork: Holistic Evaluation of Language Models
- Neerav KingslandAffiliation: AnthropicWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Niladri Shekhar ChatterjiAffiliation: OpenAIWork: Holistic Evaluation of Language Models (HELM)
- Oliver RauschAffiliation: AnthropicWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Pauline LucAffiliation: Google DeepMindWork: SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
- Percy LiangAffiliation: Stanford UniversityWork: Holistic Evaluation of Language Models
- Peter HendersonAffiliation: Princeton UniversityWork: Holistic Evaluation of Language Models
- Pidong WangAffiliation: Not recordedWork: MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
- Runpeng GengAffiliation: The Pennsylvania State University (Penn State)Work: PIArena: A Platform for Prompt Injection Evaluation
- Ryan Andrew ChiAffiliation: OpenAIWork: Holistic Evaluation of Language Models
- Scott HeinerAffiliation: Not recordedWork: Discovering Language Model Behaviors with Model-Written Evaluations
- Scott Mayer McKinneyAffiliation: OpenAIWork: International evaluation of an AI system for breast cancer screening
- Sébastien M. R. ArnoldAffiliation: Google DeepMindWork: Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluations
- Shibani SanturkarAffiliation: OpenAIWork: Holistic Evaluation of Language Models (HELM)
- Simin ChenAffiliation: Columbia UniversityWork: Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic Evaluation
- Surya GanguliAffiliation: Stanford UniversityWork: Holistic Evaluation of Language Models (HELM)
- Tatsunori HashimotoAffiliation: Stanford UniversityWork: Holistic Evaluation of Language Models
- Thomas IcardAffiliation: GoodfireWork: Holistic Evaluation of Language Models
- Tianyi ZhangAffiliation: Not recordedWork: Holistic Evaluation of Language Models
- Victoria KrakovnaAffiliation: Google DeepMindWork: Realistic honeypot evaluations for scheming propensity
- Virginie DoAffiliation: Not recordedWork: ARE: Scaling Up Agent Environments and Evaluations
- William Yang WangAffiliation: University of California, Santa BarbaraWork: Holistic Evaluation of Language Models
- Xuechen LiAffiliation: xAIWork: Holistic Evaluation of Language Models
- Yana HassonAffiliation: Google DeepMindWork: AIMIP Phase 1: systematic evaluations of AI weather and climate models
- Yian ZhangAffiliation: NVIDIAWork: Holistic Evaluation of Language Models
- Yifan MaiAffiliation: Stanford UniversityWork: Holistic Evaluation of Language Models
- Yilin ZhangAffiliation: MetaWork: GenAI Evaluation Maturity Framework (GEMF) to assess and improve GenAI Evaluations