500AI
Search

Bailin Wang

  • Gated Linear Attention Transformers with Hardware-Efficient Training
  • Parallelizing Linear Transformers with the Delta Rule over Sequence Length
  • Gated Delta Networks: Improving Mamba2 with Delta Rule

All names