Jared Casper
Source-listed: Senior Deep Learning Scientist · NVIDIA
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- Megatron-LMSaved credit: ContributorExplore people connected to this work →
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model ParallelismSaved credit: Contributor · 2019Explore people connected to this work →
- Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LMSaved credit: Contributor · 2021Explore people connected to this work →
- Reducing Activation Recomputation in Large Transformer ModelsSaved credit: Contributor · 2023Explore people connected to this work →
- An Empirical Study of Mamba-based Language ModelsSaved credit: Contributor · 2024Explore people connected to this work →