Mostofa Patwary
Source-listed: Director of Large Foundational Language Model · NVIDIA Corporation
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model ParallelismSaved credit: Contributor · 2019Explore people connected to this work →
- Megatron-LMSaved credit: ContributorExplore people connected to this work →
- Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LMSaved credit: Contributor · 2021Explore people connected to this work →
- Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language ModelSaved credit: Contributor · 2022Explore people connected to this work →
- LLM Pruning and Distillation in Practice: The Minitron ApproachSaved credit: Contributor · 2024Explore people connected to this work →