A 1.0 that mostly says no
Pi hit 1.0 on 1 October 2026, and the release notes are short enough to read on a phone. Seven additions: Codemode, extension support for virtual models, deferred tool loading, cache warming for Anthropic models, mid-conversation system messages, a new terminal theme, and full-screen mode turned on by default. That is the whole list for a terminal coding agent used by hundreds of thousands of people every week.
The restraint is the point. Agent tooling churns weekly, and most of that churn is dead within a quarter. Pi's maintainers hold features at arm's length until the thing has survived contact with real users, then weigh whatever it does against the complexity it drags in. Plenty of candidates were tried and dropped. We run a similar rule on our own internal tooling and it is unpopular right up until the moment a dependency gets abandoned.
Pi 1.0 landed with a deliberately short feature list and an MIT licence.
Installation is a one-liner on macOS and Linux:
curl -fsSL https://pi.dev/install.sh | sh
Everything is MIT licensed, which matters more than it sounds. If you are building client systems on top of an agent framework, a permissive licence and a codebase you can actually read are worth more than a glossy feature matrix.
The release that actually changes your architecture
Shipped the same day, as an experimental package, is Pi Durable. It is not a new coding agent. It is the plumbing underneath one: storage plus the machinery to run many model conversations in parallel, the tools those conversations call, and the environments those tools execute in.
The reason this matters is embarrassingly mundane. A terminal agent assumes one human watching one process. If it dies, you read the scrollback and tell it to carry on. Every agent product we have built for a client breaks that assumption immediately. The agent is reached from Slack, or WhatsApp, or a web dashboard. Two people poke it at once. It runs for forty minutes while the container gets redeployed under it. Nobody is watching the scrollback because there is no scrollback.
So teams write their own recovery layer. We have done it, badly, more than once. You end up with a half-finished workflow engine bolted to a chat transcript, and the bugs are the worst kind: silent, intermittent, and only visible in production.
Pi Durable is the experimental substrate for long-running agents, shipped alongside the 1.0 release.
Pi Durable's model is that every step is a task that writes a checkpoint before moving on. Process dies, a new process opens the same storage, finds the unfinished tasks, and resumes each from its last checkpoint. A model request that was cut off gets sent again, with the partial answer kept in the transcript and marked aborted. Queued messages stay queued. Submissions carry a requestId, so a client that retries after a crash gets the original submission back instead of asking the agent to do the job twice.
Storage ships in three flavours: in-memory, SQLite, and JSONL, with a conformance suite and benchmarks if you want to write your own against a key-value store or Postgres. The SQLite and JSONL backends avoid Node APIs entirely, so with a thin adapter they run on Bun or inside a Cloudflare Durable Object. On SQLite only the working set stays in memory: active transcripts, live tasks, pending submissions. Compaction keeps each active transcript bounded by the model's context window, so a conversation with tens of thousands of messages still fits comfortably in RAM.
Replay safety is where the real thinking happens
Here is the part that will bite you if you build this yourself, and the part I would steal even if you never install the package.
When a crash interrupts a tool call, the runtime has to decide whether rerunning it is safe. Reading from an issue tracker, safe. Debiting a mobile money wallet, absolutely not. Pi Durable makes that an explicit property of the tool rather than a guess at recovery time. A tool marked replay-safe simply runs again. An unmarked tool does not: the model is told the call was interrupted, handed whatever output was captured, and left to decide.
const sendPayout = defineTool({
name: "send_payout",
description: "Disburse a mobile money payment",
parameters: Type.Object({ msisdn: Type.String(), amount: Type.Number() }),
// No replay flag. A payout cut off by a crash is reported to the model,
// never silently repeated.
execute: async (args, api) => {
const ref = `payout-${api.taskId}`; // idempotency key the provider honours
const res = await momo.debit(args.msisdn, args.amount, ref);
return { content: [{ type: "text", text: res.status }] };
},
});
Note the second defence in that snippet. The task id doubles as an idempotency key at the payment provider, which is what you want anyway. Durability at the runtime layer does not excuse you from idempotency at the integration layer, and anyone selling you otherwise has not sat through a payout reconciliation at month end.
Tasks and conversations form a single ownership tree, which is the other idea worth copying. Aborting a parent aborts what it owns, bottom-up, so children clean up their own side effects before the parent reports done. Subagents are not a built-in feature at all. A subagent is just a conversation owned by the tool call that spawned it, which takes a few lines to write and inherits crash recovery, its own cost accounting, and its own place in the UI for free. Fewer concepts, more composition.
The costs people underestimate
Two numbers are worth internalising before you commit.
First, context. Compaction runs as a background task while the conversation keeps going, and it only blocks when the next request genuinely would not fit. Defaults sit around a 16K token reserve before a request waits for a summary, with background summarisation kicking off roughly 32K tokens earlier. Tune those and you are trading token spend against latency spikes in the middle of a user's turn. Get it wrong in the cheap direction and your agent stalls for fifteen seconds at the worst possible moment.
Second, the prompt cache. Pi 1.0's mid-conversation system messages exist because changing a system prompt or tool set halfway through a conversation normally invalidates the cache and you pay full freight on the next request. On providers that support it, only the delta goes over the wire and the cache survives. On providers that do not, dynamic prompts become a line item on your bill. We have watched a client's monthly inference spend roughly double from nothing more than an over-eager prompt section that rebuilt itself every turn.
Then there is the codebase itself as a cost. Pi Durable is about 15,000 lines without tests, which works out to roughly 150,000 tokens for GPT-family tokenisers and about 250,000 for Claude. That is the worst case, and it is deliberate: the whole thing is sized so your agent can read the framework it runs on. The bundled vacation planning demo is around 1,300 lines of TypeScript, most of it terminal UI.
Whether to put it in production
Not yet, not on the critical path. The API is explicitly experimental and will move. But the design is the clearest statement we have seen of what an agent runtime actually needs: checkpointed tasks, explicit replay semantics, an ownership tree for cancellation, application documents committed atomically with the transcript, and many clients attached to one conversation at once.
If you are shipping agentic features this quarter, the practical move is to spend an afternoon with the examples and audit your own tool list against the replay question. Mark every tool safe or unsafe to rerun, right now, on paper. Most teams discover two or three that will quietly double-charge somebody the first time a container restarts mid-call. That audit costs nothing and pays for itself the first time your infrastructure does what infrastructure does.





