Blog · corpus.blog/blogs/rabzelj.com/posts
rabzelj.com
rabzelj.com
2026
9 Mar 2026
11 Feb 2026
2025
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints | Paper Notesoriginal ↗
10 Sept 2025
10 Sept 2025
6 Sept 2025
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention | Paper Notesoriginal ↗
29 Aug 2025
28 Aug 2025
27 Aug 2025
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation | Paper Notesoriginal ↗
27 Aug 2025
26 Aug 2025
26 Aug 2025
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer | Paper Notesoriginal ↗
22 Aug 2025
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer | Paper Notesoriginal ↗
22 Aug 2025
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding | Paper Notesoriginal ↗
19 Aug 2025
10 Aug 2025
Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference | Paper Notesoriginal ↗
7 Aug 2025
20 Feb 2025
17 Feb 2025
2024
10 Feb 2024
2023
2021
2020
19 Jul 2020
23 Jun 2020
2018
10 Jul 2018