corpus.blog
Most cited
Talked about
Blogs
56,966 blogs · [ { "id": "01a08785-f641-729f-b9b0-c7f9f6cc4347", "title": "Reverse engineering Apple's simdgroup async copy on M4", "url": "https://ighoshsubho.bearblog.dev/reverse-engineering-apples-simdgroup-async-copy-on-m4/", "published_at": "2026-05-21T12:12:03+00:00" }, { "id": "01a08785-f641-729f-b9b0-c7f9f7aa4dc2", "title": "My 2 cents on Fusing GEMM + Top-K + Softmax on SM100", "url": "https://ighoshsubho.bearblog.dev/my-2-cents-on-fusing-gemm-top-k-softmax-on-sm100/", "published_at": "2026-03-23T15:16:00+00:00" }, { "id": "01a08785-f641-729f-b9b0-c7f9f86b3cf0", "title": "Optimizing 3D Square Convolution for cuDNN-like Performance - A Worklog", "url": "https://ighoshsubho.bearblog.dev/optimizing-3d-square-convolution-for-cudnn-like-performance-a-worklog/", "published_at": "2025-04-18T17:07:00+00:00" }, { "id": "01a08785-f641-729f-b9b0-c7f9f949c79d", "title": "Preconditioned SGD can level up your training game", "url": "https://ighoshsubho.bearblog.dev/preconditioned-sgd-can-level-up-your-training-game/", "published_at": "2025-01-30T12:06:00+00:00" }, { "id": "01a08785-f641-729f-b9b0-c7f9f9bf4d58", "title": "Understanding Lightning Attention: A Breakthrough in Linear Attention Efficiency", "url": "https://ighoshsubho.bearblog.dev/understanding-lightning-attention-a-breakthrough-in-linear-attention-efficiency/", "published_at": "2025-01-25T07:01:00+00:00" }, { "id": "01a08785-f641-729f-b9b0-c7f9faa2ed5a", "title": "From 10 to 1000 Tokens/Second: Cursor AI's Secret Weapon Revealed", "url": "https://ighoshsubho.bearblog.dev/from-10-to-1000-tokenssecond-cursor-ais-secret-weapon-revealed/", "published_at": "2024-11-17T12:47:00+00:00" } ] posts
Claim your blog
Back to Subho's research at your service 🫡
Blog · corpus.blog/blogs/ighoshsubho.bearblog.dev/posts
Subho's research at your service 🫡
ighoshsubho.bearblog.dev
2026
Reverse engineering Apple's simdgroup async copy on M4
original ↗
21 May 2026
My 2 cents on Fusing GEMM + Top-K + Softmax on SM100
original ↗
23 Mar 2026
2025
Optimizing 3D Square Convolution for cuDNN-like Performance - A Worklog
original ↗
18 Apr 2025
Preconditioned SGD can level up your training game
original ↗
30 Jan 2025
Understanding Lightning Attention: A Breakthrough in Linear Attention Efficiency
original ↗
25 Jan 2025
2024
From 10 to 1000 Tokens/Second: Cursor AI's Secret Weapon Revealed
original ↗
17 Nov 2024