SØNDAG
2026-10-04

Too many projects, too many ideas, too few hours — one learning a day anyway

ds4 runs big open weights locally and speaks the Anthropic API

The part I care about is not the 284-billion-parameter model squeezed down to 2 bits. It is that ds4-server speaks the Anthropic API at /v1/messages, with Claude Code on the list of clients. I could point tooling I already use at a box on my own desk.

The catch is the box. Supported Macs start at 64 GB, depending on model, and the benchmarks come from 128 GB machines. So I would check the hardware matrix before cloning anything. If it fits, I would test the disk-backed KV cache first: long prefixes that survive a restart are what a coding session needs.


The story — DwarfStar 4 (ds4) is an MIT-licensed C inference engine for high-memory Mac, CUDA and ROCm machines. It runs DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 using asymmetric 2-bit quantization, persists the KV cache to SSD keyed by prompt hash, and ships a CLI, a server with OpenAI and Anthropic APIs, and a native coding agent. On an M5 Max with 128 GB it lists 34.4 tokens per second generation at 32K context. (Source)