Rohan Varma
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- Training LLMs with Fault Tolerant HSDP on 100,000 GPUsSaved credit: Contributor · 2026Explore people connected to this work →
- PyTorch FSDP: Experiences on Scaling Fully Sharded Data ParallelSaved credit: Contributor · 2023Explore people connected to this work →
- PyTorch RPC: Distributed Deep Learning Built on Tensor-Optimized Remote Procedure CallsSaved credit: Contributor · 2023Explore people connected to this work →
- Introducing PyTorch Fully Sharded Data Parallel (FSDP) APISaved credit: Contributor · 2022Explore people connected to this work →
- PyTorch Distributed: Experiences on Accelerating Data Parallel TrainingSaved credit: Contributor · 2020Explore people connected to this work →