Chan Jun Shern
Source-listed: Member of Technical Staff · Anthropic
- Information retrieval
- Evaluation
Selected work
5- PaperBench: Evaluating AI’s Ability to Replicate AI ResearchContributor · 2025
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringContributor · 2024
- GPT-4o System CardContributor · 2024
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI BenchmarkContributor · 2023
- Few-shot Adaptation Works with UnpredicTable DataFirst author · 2023