Bailin WangGated Linear Attention Transformers with Hardware-Efficient TrainingParallelizing Linear Transformers with the Delta Rule over Sequence LengthGated Delta Networks: Improving Mamba2 with Delta RuleAll names