The interesting line in this paper isn’t a throughput number. It’s the claim that agent workloads need an elastic execution platform, not a single sandbox runtime. Anyone wiring Claude agents into their own repos meets that on a small scale: one container per task works until something untrusted wants a microVM and something trivial wants a plain function call.
DSec puts FnCall, container, microVM and full-VM backends behind one SDK, with environments built from independently versioned layers. I won’t build that on a home server. I’d steal the shape: one interface, isolation chosen per task, state kept apart from the thing doing the thinking. And note they treat reward hacking as the sandbox’s problem too.
The story — DeepSeek posted a 31-page report on DSec, the production sandbox platform behind its agentic LLM training and evaluation. It unifies four sandbox backends under one SDK, loads images on demand from the 3FS distributed filesystem, and is co-designed with the RL framework so stateful rollouts survive preemptible GPU training. One unit spans around 160 nodes and serves about 3 million sandboxes a day, with over 380,000 concurrent in production. (Source)