← People

Xiangrui Yu

Source-listed: RD · PaddlePaddle, Baidu

  • Language models
  • Efficient ML

Selected work

5
  1. SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUsMaintainer · 2025
  2. ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionContributor · 2026
  3. DynSpAttn: Efficient Attention via Dual-Side Dynamic Sparsity on Sparse Tensor CoresContributor · 2026
  4. ROME: Maximizing GPU Efficiency for All-Pairs Shortest Path via Taming Fine-Grained IrregularitiesContributor · 2026
  5. Balancing Computation and Communication in Distributed Sparse Matrix-Vector MultiplicationContributor · 2023