#ai
- Two Cents a Task Changes What You Bother Automating
The number that stopped me wasn't 89.0% on ARC-AGI-1. It was $0.02 per task. I've spent the last year writing prompts as if every call were…
- The bottleneck is me, not the model
The uncomfortable part of Goedecke's argument is that it explains my own logs. When I'm working on the Astro site or a Cloudflare Worker…
- A moderation model I can actually host on the boat
The number that got my attention isn't the 7x claim, it's 16GB. A 3B model that fits on one 16GB GPU is a thing I can run myself, next to…
- Kimi K3 is close enough that I'll rerun my agent evals
The number that got my attention isn't the 2.8 trillion parameters. It's that Kimi K3 won three of six agent tests, and on Artificial…
- Cheap tokens only matter if the model shuts up
Luna at $0.20 in / $1.20 out per million tokens is the kind of number that changes what I'm willing to run in a loop. The jobs I've been…
- Cloudflare's quantization math is the same math I run on the boat
The part of this I keep rereading is the throughput table for the KV cache. At one concurrent request, BF16 beats FP8 — 137 tokens per…
- Documentation is state. Insight lives in moments.
Fed my AI curation panel 20 candidates harvested from my project documentation. It rejected 17 outright and drafted 3 status reports no one…
- The model I write this blog with just got pulled
Most of The Cloudy Brain is co-written with Fable 5 — it drafts, I argue with it, I sign. So this one lands close to home: as of yesterday,…
- Munich court: "the AI said it" is not a shield
If you ship anything that presents LLM answers as authoritative to end users, this is the precedent to read twice. The disclaimer-and-pray…
- "You'll never know if the AI degrades you" — except you can
I spent Fable 5's launch day building with it — a whole publishing platform, including the bumps: it over-built twice, and I had to catch…