Blog · corpus.blog/blogs/prasadkhake.com/posts
prasadkhake.com
prasadkhake.com
2026
23 Jul 2026
Running Llama 3.1 8B with FP8 on vLLM Cuts Cost from $1.00 to $0.36 per Million Output Tokensoriginal ↗
14 Jul 2026
12 Jul 2026
8 Jul 2026
30 Jun 2026
18 Jun 2026
I built self-speculative decoding for MLX. On an M3, naive layer-skip never beats baseline — 24 configs, 24 lossesoriginal ↗
18 Jun 2026
Three ways to make an LLM read its weights less often on a Mac — and why each one backfiresoriginal ↗
14 Jun 2026
I expected a diffusion LLM to be fast on my Mac. It tied the best model on quality instead — and lost on speed.original ↗
13 Jun 2026
8 Jun 2026
4 Jun 2026