Xiangrui Yu
Source-listed: RD · PaddlePaddle, Baidu
- Language models
- Efficient ML
Selected work
5- SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUsMaintainer · 2025
- ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionContributor · 2026
- DynSpAttn: Efficient Attention via Dual-Side Dynamic Sparsity on Sparse Tensor CoresContributor · 2026
- ROME: Maximizing GPU Efficiency for All-Pairs Shortest Path via Taming Fine-Grained IrregularitiesContributor · 2026
- Balancing Computation and Communication in Distributed Sparse Matrix-Vector MultiplicationContributor · 2023