Tianqi Liu
Source-listed: Google DeepMind
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- RRM: Robust Reward Model Training Mitigates Reward HackingSaved credit: Contributor · 2025Explore people connected to this work →
- Building Math Agents with Multi-Turn Iterative Preference LearningSaved credit: Contributor · 2025Explore people connected to this work →
- LiPO: Listwise Preference Optimization through Learning-to-RankSaved credit: Contributor · 2025Explore people connected to this work →
- Statistical Rejection Sampling Improves Preference OptimizationSaved credit: Contributor · 2024Explore people connected to this work →
- Gemma 3 Technical ReportSaved credit: Contributor · 2025Explore people connected to this work →