Blog · corpus.blog/blogs/vllm.ai/posts
vLLM Blog
vllm.ai
2026
24 Sept 2026
21 Sept 2026
vLLM x Novita AI: Chord, Faster INT4 MoE for Kimi K2.x. Up to 1.3x on H200, 2.15x on Untuned B300original ↗
15 Sept 2026
10 Sept 2026
MiniMax H3 on vLLM-Omni: From System-Wide Optimization to Real-Time Serving with FastVideo’s FastH3original ↗
1 Sept 2026
17 Aug 2026
29 Jul 2026
28 Jul 2026
23 Jul 2026
23 Jul 2026
21 Jul 2026
16 Jul 2026
EAGLE3 Speculative Decoding on AMD Instinct GPUs: Training and Serving with vLLM and AMD Quarkoriginal ↗
13 Jul 2026
6 Jul 2026
Session-Aware Agentic Routing: Continuity-Aware Model Selection for Long-Horizon LLM Agentsoriginal ↗
2 Jun 2026
28 May 2026
28 May 2026
EAGLE 3.1: Advancing Speculative Decoding Through Collaboration Between the EAGLE Team, vLLM, and TorchSpecoriginal ↗
26 May 2026
14 May 2026