hn.today

What we learned from building the sandboxes our agents run in

qawolf.com12 points1 comments
Screenshot of What we learned from building the sandboxes our agents run in

A team built AI agents for QA and discovered treating each agent like a short-lived cloud job was wrong: an agent behaves like a coworker - it waits on people, runs unreviewed commands, touches real systems and holds half-finished edits. Trying to give agents one-off tools led to a Frankenstein of pretend filesystems and shells, so they instead give every session its own isolated cloud machine with a real checkout under version control, a full shell (git, node, Python, QA CLI) and a real browser you can watch when needed. That shift mirrors what a human engineer uses on a laptop and kept heavy work from slowing other customers; more than 6,000 daily machines go to live sessions and thousands of threads are handled each weekday.

To support sessions rather than jobs they unlearned cloud defaults and implemented concrete fixes: a pool of prebooted machines to avoid cold starts and keep a session’s machine for five minutes after each reply; a strict one-attempt policy so failures don’t duplicate side effects; autosave of in-progress edits to persistent storage so work survives machine death; and least-privilege sandboxing with private secret injection, no cloud credentials, and running with the user’s permissions. Large tasks are split into up to 50 independent agents, each with its own machine and branch, orchestrated via the public CLI. The result keeps machines disposable but makes sessions durable and safer without inventing novel primitives.

Read on qawolf.com1 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Web

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.