56,966 blogs · [ { "id": "01a087c8-ff86-7380-88e3-d2ac7ea26a61", "title": "My data-loader fix just won a $1,000 Vesuvius Challenge Progress Prize", "url": "https://prasadkhake.com/blog/vesuvius-zarr-progress-prize/", "published_at": "2026-08-05T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac7f8918eb", "title": "My readability metric found pockets in a sealed scroll. They were shredded.", "url": "https://prasadkhake.com/blog/pherc1203-torn-pockets/", "published_at": "2026-07-27T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac7fa5cba9", "title": "I mapped local readability cues inside a sealed Herculaneum scroll", "url": "https://prasadkhake.com/blog/pherc1203-readability-atlas/", "published_at": "2026-07-24T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac7fdc198f", "title": "Why a 2.4 µm Herculaneum scan still reads as noise — and what actually blocks it", "url": "https://prasadkhake.com/blog/vesuvius-readability-gate/", "published_at": "2026-07-23T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac80cbe964", "title": "sarvam-translate on three Windows laptops: 17.5 to 73.6 tok/s", "url": "https://prasadkhake.com/blog/on-device-4b-windows-old-gpu/", "published_at": "2026-07-22T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac81091504", "title": "The same Herculaneum scroll query can fetch 1.3× — or 176× too much data", "url": "https://prasadkhake.com/blog/vesuvius-chunk-size/", "published_at": "2026-07-21T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac81d5c9df", "title": "KV Cache Quantization Is 4× Slower on My Mac and 28% Faster on a Rented L4", "url": "https://prasadkhake.com/blog/kv-quantization-m3-vs-l4/", "published_at": "2026-07-19T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac82955e7c", "title": "Running Llama 3.1 8B with FP8 on vLLM Cuts Cost from $1.00 to $0.36 per Million Output Tokens", "url": "https://prasadkhake.com/blog/inference-audit-8b-rag/", "published_at": "2026-07-14T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac82a95e3b", "title": "Carmack's right about the weights. The KV cache is the part his argument skips.", "url": "https://prasadkhake.com/blog/kv-cache-cant-stream-from-flash/", "published_at": "2026-07-12T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac836d90e7", "title": "A rotating KV cache saves 36% of your memory and 100% of your recall", "url": "https://prasadkhake.com/blog/kv-cache-eviction-tax/", "published_at": "2026-07-12T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac84636367", "title": "At 32,000 tokens, the costliest thing my MacBook did was wait seven minutes to speak", "url": "https://prasadkhake.com/blog/kv-cache-tax-m3-vs-l4/", "published_at": "2026-07-08T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac84f068c3", "title": "Learning fine-tuning by building a tool-calling LoRA on an M3", "url": "https://prasadkhake.com/blog/lora-tool-calling-m3/", "published_at": "2026-07-07T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac859684f1", "title": "When to hand-write a GPU kernel on Apple Silicon (and when the compiler already won)", "url": "https://prasadkhake.com/blog/metal-kernels-from-scratch/", "published_at": "2026-06-30T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac859c3873", "title": "Atomic Chat's TurboQuant headline did not survive a chat-generation benchmark on my M3", "url": "https://prasadkhake.com/blog/atomic-turboquant-m3/", "published_at": "2026-06-18T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac85b1e64e", "title": "I built self-speculative decoding for MLX. On an M3, naive layer-skip never beats baseline — 24 configs, 24 losses", "url": "https://prasadkhake.com/blog/self-spec-mlx-m3/", "published_at": "2026-06-18T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac8690d505", "title": "Three ways to make an LLM read its weights less often on a Mac — and why each one backfires", "url": "https://prasadkhake.com/blog/three-ways-fewer-weight-reads-mac/", "published_at": "2026-06-14T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac872d2465", "title": "I expected a diffusion LLM to be fast on my Mac. It tied the best model on quality instead — and lost on speed.", "url": "https://prasadkhake.com/blog/diffusion-llm-16gb-mac/", "published_at": "2026-06-13T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac87ad49f7", "title": "Apple's on-device model ties a 4-bit Llama-3.1-8B — and won't name the M1", "url": "https://prasadkhake.com/blog/apple-on-device-model-benchmark/", "published_at": "2026-06-09T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac884ed05f", "title": "I turned on MLX's memory-saving flag and ran out of memory", "url": "https://prasadkhake.com/blog/kv-bits-memory-flag-backfires/", "published_at": "2026-06-09T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac891b82a7", "title": "Speculative decoding on a 16 GB Mac: a 20% win that becomes a 25% loss", "url": "https://prasadkhake.com/blog/speculative-decoding-16gb-mac/", "published_at": "2026-06-09T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac89222f87", "title": "Gemma 4 on a 16 GB Mac: the E4B matches the 12B at 42% less RAM and 3× the speed", "url": "https://prasadkhake.com/blog/gemma-4-qat-on-m3/", "published_at": "2026-06-08T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac8946fac3", "title": "My benchmark graded '7! = 5040' as wrong — and two other ways it lied to me", "url": "https://prasadkhake.com/blog/benchmark-bugs-that-inflated-my-scores/", "published_at": "2026-06-07T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac89760af2", "title": "One flag makes Qwen3-4B beat Llama-3.1-8B on a 16 GB Mac — at half the RAM", "url": "https://prasadkhake.com/blog/qwen3-4b-thinking-flag-16gb-mac/", "published_at": "2026-06-06T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac8a20d8bd", "title": "Gemma 4 12B on a 16 GB Mac: 11 GB RAM, 2.7 tok/s, and what my benchmark got wrong", "url": "https://prasadkhake.com/blog/gemma-4-12b-m3-benchmark/", "published_at": "2026-06-04T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac8a6cc579", "title": "Attention sinks: the four tokens that stabilize infinite context on a 16 GB Mac", "url": "https://prasadkhake.com/blog/streamingllm-attention-sinks-16gb-mac/", "published_at": "2026-06-03T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac8ae37aec", "title": "Gemma-3-12B QAT vs Qwen3-14B 3-bit: same quality on a 16 GB Mac, but the smaller model runs lighter and faster", "url": "https://prasadkhake.com/blog/gemma-3-12b-vs-qwen3-14b-16gb-mac/", "published_at": "2026-06-02T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac8b31a15c", "title": "What actually runs well on a 16 GB MacBook", "url": "https://prasadkhake.com/blog/16gb-mac-llm/", "published_at": "2026-06-01T00:00:00+00:00" }, { "id": "01a087c8-ff86-7380-88e3-d2ac8b7cd0b4", "title": "Why Mistral and Devstral models drop their spaces on Apple Silicon", "url": "https://prasadkhake.com/blog/mlx-tekken-detokenizer/", "published_at": "2026-05-30T00:00:00+00:00" } ] posts Claim your blog
Back to prasadkhake.com
Blog · corpus.blog/blogs/prasadkhake.com/posts

prasadkhake.com

prasadkhake.com

2026