Blog · corpus.blog/blogs/inferencex.semianalysis.com/posts
inferencex.semianalysis.com
inferencex.semianalysis.com
2026
MiniMax M3 on AgentX: Why B200 and B300 Beat Their Rack-Scale GB200 NVL72 & GB300 NVL72 Counterpartsoriginal ↗
25 Aug 2026
25 Aug 2026
9 Jun 2026
GB300 NVL72 vs GB200 NVL72 Inference Performance & Perf per Dollar - on DeepSeek-V4-Pro 1.6T: Up to 2.83x Throughputoriginal ↗
27 May 2026
B200 NVFP4 vs H200 FP8 on GLM-5: Up to 3.65x Better Performance per Dollar with SGLang MTPoriginal ↗
26 May 2026
B200 NVFP4 vs H100 FP8 on MiniMax-M2.5: Up to 8.2x Better Performance per Dollar with vLLMoriginal ↗
26 May 2026
26 May 2026
25 May 2026
AMD MI355X Qwen3.5 397B-A17B Inference: Up to 19x Throughput per GPU in 3 Months on SGLang FP8original ↗
25 May 2026
23 May 2026
AMD MI355X Kimi K2.5 Inference: 7.7x Throughput, Up To 15x Interactivity in 25 Days on vLLMoriginal ↗
22 Apr 2026