Field manual for agent operators
Make AI agents actually do the work.
Setup guides, failure post-mortems and operating notes for people who already run AI agents, written from running dozens of them side by side every day.
- Setup
- Coordination
- Verification
- Cost
- Post-mortems
Eight tracks
Where agents stop doing the work
Agents rarely fail at writing code or text. They fail around it: unclear instructions, no ownership, no check, no handoff. Six tracks cover those gaps; two more collect runnable agent builds and framework comparisons.
- SETUP01
Setup
Instruction files, permissions, tools, memory and subagents, set up so the agent can finish a job unattended.
Open track - COORD02
Multi-agent coordination
Handoffs, shared state and concurrent edits when several agents work in the same repository.
Open track - VERIFY03
Review & deploy safety
Checks that catch an agent reporting "done" when it is not, and deploys that cannot ship someone else's work.
Open track - COST04
Cost & model routing
Which model does which job, what caching and batching save, and where the bill really comes from.
Open track - POSTMORTEM05
Post-mortems
What broke, why the agent did it, the fix and the guard that stops it happening again.
Open track - CASE06
Operating cases
How real workloads are run: long tasks, background agents, queues and the daily routine around them.
Open track - LIBRARY07
Agent case library
Open-source agent builds by use case: what each one solves, how to run it, what we saw, and a link to the code.
Open track - FRAMEWORK08
Frameworks compared
Agent frameworks side by side: what each is built for, licence, language, and where it gets in the way.
Open track
Latest
New guides and post-mortems
The first guides are being written and checked against each tool's own documentation.
Agent case library
Open-source agent builds, by the job they do
Research teams, recruiting, competitor tracking, daily briefings. Each page takes one public, runnable agent project and answers four things, with a link to the original code.
- What it solvesThe job, and who would use it.
- How to run itRequirements, keys and the steps.
- What we observedDated. Read-only pages say so.
- Where it breaksLimits, cost and what to change.
The first case pages are being read, run and written up. Each one will link the original code and state the date it was tested.
Operating rules
Five rules that hold when many agents share one codebase
These come from daily operation, not from a framework's documentation. Every guide on the site applies at least one of them.
Done means checked, not reported
An agent's summary is a claim. Exit code 0, a green build and an HTTP 200 are signals. The task is done when the result has been read back from where a user would see it.
One writer per file at a time
Two agents editing the same file is a race. Give each task a clear set of paths, or give each agent its own worktree, and make shared files small, atomic edits.
Commit only your own lines
A shared working tree holds everyone's unfinished work. An agent that runs a plain commit ships all of it. Stage by path, check the diff, and never sweep the tree.
Write the handoff down
The next session knows nothing. A handoff states what is finished, what is half-done, what was ruled out and the exact next step, in a file the next agent will actually read.
Deploy from a state you can name
Before a deploy, the agent should know what is live, what is in the repository and what is only on disk. If those three differ and it cannot explain why, it stops.
Post-mortem format
Every failure written up the same way
A post-mortem is useful only if you can act on it. Each one here has four parts, and the last one is never optional.
- 01 · What broke
The visible damage
What the agent did, to which files or systems, and how it was noticed.
- 02 · Why
The cause in the setup
The missing instruction, missing check or shared resource that made the mistake likely.
- 03 · Fix
How it was repaired
The recovery steps that worked, including the ones that did not.
- 04 · Guard
What stops a repeat
The rule, hook or check added so the same failure cannot happen quietly again.
Questions
Before you read further
Who is Agent Talk for?
People who already run AI agents, such as coding agents in a terminal or agents built on a framework, and want them to finish real work without constant supervision. If you are still deciding whether to try an agent at all, most guides here will assume more than you need.
Why does my agent say a task is done when it is not?
Usually because nothing in its instructions defines what done means, so it reports the last step it ran. The fix is a completion check the agent must run and quote before it reports: a test, a build, or reading the result back from where a user would see it.
Can several agents work in the same repository at once?
Yes, with rules. Each agent needs a clear owner for the files it edits, commits that contain only its own changes, and a written handoff when work moves between sessions. Without those, agents overwrite and commit each other's unfinished work.
Do I need a multi-agent framework to run more than one agent?
No. Many working setups are several independent sessions sharing a repository, an instructions file and a task list. A framework helps when agents must call each other programmatically; it does not replace ownership rules or verification.
How do I keep agent costs under control?
Route each kind of task to the cheapest model that passes your checks, keep long stable instructions cacheable, hand noisy searches to a subagent so the main context stays small, and measure spend per task type rather than per day.
Ask
Describe what your agent did
Tell the assistant what you asked for and what happened instead. It answers now, and questions that keep coming up become full guides.