Maxim Krikun
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsSaved credit: Contributor · 2022Explore people connected to this work →
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingSaved credit: Contributor · 2021Explore people connected to this work →
- Beyond Distillation: Task-level Mixture-of-Experts for Efficient InferenceSaved credit: Contributor · 2021Explore people connected to this work →
- Building Machine Translation Systems for the Next Thousand LanguagesSaved credit: Contributor · 2022Explore people connected to this work →
- Massively Multilingual Neural Machine Translation in the Wild: Findings and ChallengesSaved credit: Contributor · 2019Explore people connected to this work →