Songlin Yang
Source-listed: Member of Technical Staff · Thinking Machines Lab
Record snapshot: Sep 3, 2026. Coverage and affiliations may be incomplete or historical.
Selected work
5- Gated Linear Attention Transformers with Hardware-Efficient TrainingSaved credit: Contributor · 2024Explore people connected to this work →
- Gated Delta Networks: Improving Mamba2 with Delta RuleSaved credit: Contributor · 2025Explore people connected to this work →
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSaved credit: Contributor · 2024Explore people connected to this work →
- PaTH Attention: Position Encoding via Accumulating Householder TransformationsSaved credit: Contributor · 2025Explore people connected to this work →
- Distilling to Hybrid Attention Models via KL-Guided Layer SelectionSaved credit: Contributor · 2026Explore people connected to this work →