← People

Tianqi Liu

Source-listed: Google DeepMind

Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.

Selected work

5
  1. RRM: Robust Reward Model Training Mitigates Reward HackingSaved credit: Contributor · 2025Explore people connected to this work →
  2. Building Math Agents with Multi-Turn Iterative Preference LearningSaved credit: Contributor · 2025Explore people connected to this work →
  3. LiPO: Listwise Preference Optimization through Learning-to-RankSaved credit: Contributor · 2025Explore people connected to this work →
  4. Statistical Rejection Sampling Improves Preference OptimizationSaved credit: Contributor · 2024Explore people connected to this work →
  5. Gemma 3 Technical ReportSaved credit: Contributor · 2025Explore people connected to this work →