Yi Dong
Source-listed: Principal Research Scientist · NVIDIA
- Language models
- Reinforcement learning
- Efficient ML
Selected work
5- Nemotron-4 340B Technical ReportAuthor · 2024
- SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHFContributor · 2023
- NeMo-Aligner: Scalable Toolkit for Efficient Model AlignmentContributor · 2024
- HelpSteer2: Open-source dataset for training top-performing reward modelsContributor · 2024
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsContributor · 2025