← People

Boyi Wei

Source-listed: PhD student / PhD candidate · Princeton University

  • Language models
  • Evaluation

Selected work

5
  1. Dynamic Risk Assessments for Offensive Cybersecurity AgentsContributor · 2025
  2. Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsContributor · 2024
  3. Evaluating Copyright Takedown Methods for Language ModelsContributor · 2024
  4. On Evaluating the Durability of Safeguards for Open-Weight LLMsContributor · 2025
  5. An Adversarial Perspective on Machine Unlearning for AI SafetyContributor · 2024