corpus.blog
Most cited
Talked about
Blogs
41554 blogs · [ { "id": "01a08785-d739-70ff-a733-da2b333d9611", "title": "Fast weights and sparse attention in GLM-5.3-Flash", "url": "https://idlemachines.co.uk/essays/glm-5-3-flash", "published_at": "2026-09-04T12:00:00+00:00" }, { "id": "01a08785-d739-70ff-a733-da2b3424ca88", "title": "Tokenisation: Who decides what a token is anyway?", "url": "https://idlemachines.co.uk/essays/tokenisation", "published_at": "2026-08-09T20:00:00+00:00" }, { "id": "01a08785-d739-70ff-a733-da2b34679fe7", "title": "Adam and AdamW: adaptive optimisation and weight decay", "url": "https://idlemachines.co.uk/essays/adam-adamw", "published_at": "2026-07-31T15:00:00+00:00" }, { "id": "01a08785-d739-70ff-a733-da2b34bb1508", "title": "Thinking Machines dropped RoPE", "url": "https://idlemachines.co.uk/essays/inkling", "published_at": "2026-07-22T18:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abbd957c04", "title": "The annotated PyTorch training loop", "url": "https://idlemachines.co.uk/essays/pytorch-training-loop", "published_at": "2026-06-20T10:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abbdcd94e3", "title": "MAE vs MSE: more than just the mean vs median debate", "url": "https://idlemachines.co.uk/essays/mae-vs-mse", "published_at": "2026-06-16T00:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abbe55a57a", "title": "Heaven Knows I'm Perplexed Now", "url": "https://idlemachines.co.uk/essays/perplexed", "published_at": "2026-06-06T10:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abbf2820fc", "title": "Reading MAI's efficiency gain", "url": "https://idlemachines.co.uk/essays/efficiency-gain", "published_at": "2026-06-03T00:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abbf288fe2", "title": "Every token, everywhere, all at once", "url": "https://idlemachines.co.uk/essays/every-token", "published_at": "2026-05-16T10:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abbf6f411b", "title": "DeepSeek V4 from the inside", "url": "https://idlemachines.co.uk/essays/deepseek-v4-efficiency-release", "published_at": "2026-04-24T00:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abc0348bae", "title": "Gemma 4 is not your standard transformer", "url": "https://idlemachines.co.uk/essays/gemma4-architecture", "published_at": "2026-04-18T00:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abc03e2491", "title": "The cut in the Mixture of Experts compute graph", "url": "https://idlemachines.co.uk/essays/mixture-of-experts", "published_at": "2026-04-09T10:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abc11c41c7", "title": "Rotary Position Embeddings: a derivation", "url": "https://idlemachines.co.uk/essays/rope", "published_at": "2026-04-08T10:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abc19e39b1", "title": "Softmax, can you really derive the Jacobian? And should you care?", "url": "https://idlemachines.co.uk/essays/softmax", "published_at": "2026-03-21T10:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abc1d0e61f", "title": "Are contrastive losses just cross entropy all along?", "url": "https://idlemachines.co.uk/essays/contrastive-losses", "published_at": "2026-03-21T10:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abc25b3db9", "title": "cross-entropy: the simplest gradient in all of machine learning", "url": "https://idlemachines.co.uk/essays/cross-entropy-softmax-gradient", "published_at": "2026-03-21T00:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abc2b498e0", "title": "sigmoid: don't overflow, it's embarrassing", "url": "https://idlemachines.co.uk/essays/sigmoid-deepdive", "published_at": "2026-03-19T10:00:00+00:00" }, { "id": "01a08785-d73a-703b-8d7b-10abc2c7981b", "title": "Embeddings: why we need them, and how to build them from scratch", "url": "https://idlemachines.co.uk/essays/embeddings", "published_at": "2026-03-19T10:00:00+00:00" } ] posts
Claim your blog
Back to idlemachines.co.uk
Blog · corpus.blog/blogs/idlemachines.co.uk/posts
idlemachines.co.uk
idlemachines.co.uk
2026
Fast weights and sparse attention in GLM-5.3-Flash
original ↗
4 Sept 2026
Tokenisation: Who decides what a token is anyway?
original ↗
9 Aug 2026
Adam and AdamW: adaptive optimisation and weight decay
original ↗
31 Jul 2026
Thinking Machines dropped RoPE
original ↗
22 Jul 2026
The annotated PyTorch training loop
original ↗
20 Jun 2026
MAE vs MSE: more than just the mean vs median debate
original ↗
16 Jun 2026
Heaven Knows I'm Perplexed Now
original ↗
6 Jun 2026
Reading MAI's efficiency gain
original ↗
3 Jun 2026
Every token, everywhere, all at once
original ↗
16 May 2026
DeepSeek V4 from the inside
original ↗
24 Apr 2026
Gemma 4 is not your standard transformer
original ↗
18 Apr 2026
The cut in the Mixture of Experts compute graph
original ↗
9 Apr 2026
Rotary Position Embeddings: a derivation
original ↗
8 Apr 2026
Softmax, can you really derive the Jacobian? And should you care?
original ↗
21 Mar 2026
Are contrastive losses just cross entropy all along?
original ↗
21 Mar 2026
cross-entropy: the simplest gradient in all of machine learning
original ↗
21 Mar 2026
sigmoid: don't overflow, it's embarrassing
original ↗
19 Mar 2026
Embeddings: why we need them, and how to build them from scratch
original ↗
19 Mar 2026