Record · corpus.blog/posts/01a105fb-45d4-70a5-9dc6-da4fcd0c5835
Hand-writing a Blackwell GEMM: 164 to 1400 TFLOP/
kyrieblunders.bearblog.dev · published 2 October 2026
A record is what was published and who pointed at it. The text of the post is not held here: read it at the source, or at its Wayback capture.
0blogs citing
- First cited
- —
- Most recent
- —
- Rank this month
- not ranked
- External links
- 14
- Words captured
- 4,981
Cited by
No blog we hold has cited this post yet.
Similar posts
Optimizing Matrix Multiply on CUDA
aaron-ang.github.io
27 Sept 2026
Optimizing an NVFP4 Blockscaled GEMM on RTX PRO 6000 Blackwell GPU (SM120)
research.colfax-intl.com
9 Aug 2026
Outperforming cuBLAS on NVFP4
cudaforfun.substack.com
15 Aug 2026
30 Sept 2026
10 Sept 2026
Double-checking GPU acceleration basics
uucode.com
9 Sept 2026
Double-checking GPU acceleration basics
uucode.com
9 Sept 2026
3 Oct 2026
Really Fast Bayesian Linear Regression
rukulkarni.com
11 Sept 2026
105 % of FlashAttention-4 on average (prefill, bf16, causal, hd128, GQA)
iaroslavelistratov.github.io
1 Oct 2026
Links from this post
Link
Host
github.com