Hao Zhang
Source-listed: Assistant Professor; also listed as Staff Software Engineer at Snowflake · University of California, San Diego
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- Efficient Memory Management for Large Language Model Serving with PagedAttentionSaved credit: Contributor · 2023Explore people connected to this work →
- vLLMSaved credit: Contributor · 2023Explore people connected to this work →
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingSaved credit: Contributor · 2024Explore people connected to this work →
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningSaved credit: Contributor · 2022Explore people connected to this work →
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingSaved credit: Contributor · 2024Explore people connected to this work →