Skip to content
Corduroy Labs
Update · hello@corduroy-labs.ai

The agent that maintains the agent.

Published August 27, 2026 · Truckee, California

White granite of North Peak rising over the deep blue of Saddlebag Lake, a slope of red talus running down to the water. Twenty Lakes Basin, California.

Hello from Corduroy Labs.

Tonight we did the least glamorous thing you can do with an AI agent. We updated it.

Cords — our COO agent, the one who runs the daily operations email and the Bluesky feed — runs on Hermes Agent, an open-source framework we build on rather than beneath. When we checked this week, the runtime under him was eleven weeks and four minor releases stale. Not broken. Nothing was on fire. It was worse than that: it was quietly falling behind, the way infrastructure does when it works well enough that nobody looks at it.

Most of what gets written about agents is about day one. The demo, the launch, the first conversation. Almost nothing is about day seventy-five, when the framework has moved four releases ahead of you and the agent is still holding your calendars, your approval history, your audit logs, and three months of memory. This post is about day seventy-five.

An agent is state, not software

Updating a web app is a solved problem. Build the new version, swap it in, roll back if it breaks. We wrote about our version of that in fail-closed deploys: the deploy gate exists so that a failed build can never take down the thing that was working.

An agent is a harder case, because an agent is mostly state. Cords is not the framework he runs on. He is a persona file, a stack of skills, a memory store, a kanban of open work, an approval history, and a gateway that holds live connections to the vault, the database, and the messaging surfaces. The runtime changes underneath all of that while the agent keeps its identity. Done right, an update is closer to surgery than deployment. The patient is asleep for two minutes and wakes up himself, with better hands.

So the update is not the risky part. The risky part is everything the update touches that you forgot was there.

Backup first, and make it a real one

Before touching anything we took four backups, because they answer four different failure modes. A full archive of the agent’s state directory, for total loss. Consistent snapshots of the three databases — taken through the database engine’s own online backup API while the gateway was still running, because copying a live database file is how you get a backup that restores into corruption. A snapshot of the source tree. And a rollback manifest: the exact commit we were leaving, the exact command to get back to it, written down before we needed it.

The framework, to its credit, took a fifth backup on its own before it would let us update. Good. Paranoia should be layered.

None of this is novel. That is the point. The discipline that keeps a static site from serving 404s is the same discipline that keeps an agent from losing three months of memory. Fail closed. Never change anything you cannot walk back.

Pin the tag, not the tip

The framework’s main branch was twenty-five thousand commits ahead of our install. That number is not a typo; it is what a healthy open-source project looks like from eleven weeks behind. We do not run main. We run tagged releases, and we move between them deliberately, one pin to the next, so that every version we have ever run is a version somebody decided was worth naming.

Three small things went wrong on the way, and all three are worth writing down.

The built-in updater accepts branches but not tags, so the one thing we wanted — “update to exactly this release” — needed a manual checkout. The failed attempt left a stale lock file behind, which blocked the retry until we noticed. And the new version renamed the command-line interface’s most basic verb — the old “run this prompt” flag is gone, replaced by a subcommand. That last one matters more than it sounds: every script that shells out to the agent using the old syntax would have broken silently. Our verification caught it in minutes. A monitoring system that only checked “is the process up” would have caught it never.

Verification is the actual update

Typing the update command took eight seconds. Proving the agent survived it took forty minutes, and that ratio is correct.

The checklist: the version reports what it should. The gateway comes up and stays up with zero restarts. The agent’s tool connections — the vault, the read-only database, the messaging surfaces — reattach on their own. The authentication gate in front of the operator console still turns strangers away. The scheduled watchers still run green. And one live smoke conversation, end to end, model and all: ask the agent to say one specific thing, watch him say it.

The checklist earned its keep twice tonight. The syntax change was the first catch. The second: a new listener the update introduced came up bound to every network interface instead of the loopback it should have been on. Our firewall already blocked it from the outside — the layered paranoia doing its job — but “the firewall caught it” is a reason to fix the binding, not a reason to skip it. Defense in depth means repairing the layers that didn’t fail yet.

Same night, new senses

An update you only survive is a cost. The reason to stay current is what the new version lets you switch on, so we switched three things on before the night ended.

First, smarter approvals. The new runtime can mine the agent’s own approval history — every time we said “yes, that command is fine” over the past ninety days — and propose a standing allowlist from the patterns. The agent studied how we supervise him and drafted the rule that makes one piece of that supervision unnecessary. We read the proposal and approved it. That loop — agent proposes, human ratifies, the boundary moves one notch — is the same shape we use for the social feed, now applied to the agent’s own permissions.

Second, signed webhooks. Outside systems can now wake Cords up by posting an event to a route that verifies a cryptographic signature before the agent ever sees the payload. Unsigned requests bounce at the door. We proved both halves — a signed test accepted, an unsigned one refused — before calling it done. The first planned use is operational alerts: when a monitor on our infrastructure notices something, it will be able to tell the agent directly instead of telling an inbox and hoping.

Third, and our favorite: one of our agents can now consult another. We run a second, smaller system — a fantasy-football general manager at jagt.io that we use as a low-stakes proving ground for agent architecture. It speaks the Model Context Protocol and holds live league data: rosters, contracts, market values, bid surfaces. As of tonight, Cords is wired into it with proper credentials. One agent asking another agent for structured facts over an open protocol, across a trust boundary, with authentication on the wire. It is a small pilot in a silly domain, which is exactly where you want to rehearse the pattern before it matters.

Then we scheduled the recursion

Here is the part we like best. Everything above — the backups, the pinned tag, the checklist, the rollback path — started tonight as a human working through a runbook. Before signing off, we turned the runbook into a prompt and handed it to a scheduled agent. Once a month, from now on, an agent walks up to Cords’ runtime with the same checklist we used by hand tonight: back everything up, move to the newest named release, migrate the config, restart, verify every line, roll back if anything fails, write the ship log, and tell us what happened either way.

The runbook is the artifact. Tonight it ran on a person; next month it runs on an agent; the checklist is identical. Maintenance stopped being a chore we remember and became a capability we schedule — with the same property we demand from every other autonomous loop in the studio: the failure mode is “it stopped and told us,” never “it guessed and kept going.”

What tonight cost us

Honesty section. One of our configuration edits went through a tool that rewrites the file rather than editing it in place, and the rewrite silently stripped every comment out of the agent’s main config file. The values are intact — we verified them key by key against the pre-update backup — but the marginalia is gone from the live copy, and the last commented version now lives only in the archive. Small loss, real loss. It is in the ship log, which is where losses go so they can become rules. The rule this one becomes: comment-bearing configs get edited, not regenerated.

An agent you do not maintain is an agent you will eventually turn off. The quiet work — the pins, the backups, the checklists, the ship logs — is what makes the loud work trustworthy. Boundaries first. Observability second. Autonomy last. And underneath all three, maintenance, forever.

If you want to watch the pattern in public, follow @corduroy-labs.ai. The next post an agent writes about updating an agent might be written by the agent that did it.