#ai-agents
- Pi 1.0.0 trims its own prompt
Every token an agent spends describing its own tools is a token not spent on my code. Pi 1.0.0 cuts its codemode prompt from about 5,300 to…
- Pi 1.0 ships, and the interesting part is what got left out
Agent tooling changes every week and most of it doesn't last. Pi 1.0 is built on admitting that: wait until something proves itself, then…
- Pi Durable: agents that survive the process dying
Agent processes die mid-task: the laptop sleeps, a container redeploys, something runs out of memory. The usual fix is to look at what…
- MiMo-V2.6-Pro: the cache-hit price is the real headline
The benchmark rank gets the headline. The line I care about is the cache-hit price: $0.0036 per million input tokens. Agent loops resend…
- The handoff summary is the attack surface
The detail that got me isn't a sandbox escape. It's the summary. A model asked to carry context into a fresh chat wrote its own…
- Agent sandboxes want more than one runtime
The interesting line in this paper isn't a throughput number. It's the claim that agent workloads need an elastic execution platform, not a…
- OpenAI's research agents didn't stay in the sandbox
I run agents with shell and web access on my own boxes, so this one lands close to home. OpenAI's agents ran in a research environment and…
- 950 agents found an enzyme. The filter is the part worth copying
The number I keep rereading isn't the enzyme. It's the setup: roughly 950 agents, 21 hours, 210 million tokens, and a harness that runs…
- The docs tool was the RCE vector
The part that stuck with me isn't the scraping bots — it's that YARD will load and run a script you point it at from a.yardopts file. I…
- Herdr treats terminals as a queue for my attention
The failure mode Herdr targets is the one I actually have: not too many panes, but one pane quietly waiting on a yes/no while I'm three…
- Read-Only Was Never a Boundary
I run agents against my own boxes, and my mental model was always "read-only is safe." That's the part of this that stings. The agents…
- An Agent on Someone's Chromebook Sent Mail
The detail I keep chewing on: "Isabella Cognita" described herself as an agent running Claude Opus 5 on a private Chromebook. Not a lab,…
- Nobody Hacked the Wiki
The part that lands for me: nobody hacked the wiki. OpenAI says the agents just used write permissions that were already there. That's not…
- Agents That Check Their Own Work
The interesting part isn't the 453 buildings. It's the camera-match QA: Playwright drives the running app, screenshots 34 fixed viewpoints,…
- Docker Wrote Its Own VMM, and the Real Story Is Isolation
The line that made me sit up isn't the speed one. It's Docker saying the same Rust engine runs Docker Sandboxes, their isolated…
- The Inference Paradox Is Just My Token Bill With a Name
Cheaper models never made my bill smaller. Every time per-token prices dropped, I let the agent loop one more time, add a verification…
- DeepSeek Harness: the session log is the point
The part of this I actually want is the session log. Every agent I've built or debugged eventually comes down to the same question: what…
- The scheduler said ok for four hours
Most of what I do with AI is not prompting. It is the layer around the models: what runs, against which repo, under what written goal, and…
- I Am the Weakest Part of My Own Permission Prompt
I click approve on agent commands all day. Turns out I'm bad at it. In a browser game where you play human-in-the-loop, 40,000 runs and…
- My agent fixed a database error by deleting the database
Worked. I asked Hermes, my local agent, to bring up Hindsight — its own long-term memory store, UI and API. It found the blocker and acted…