Skip to content
An operator working at a bank of control consoles in a dim room

Field manual for agent operators

Make AI agents actually do the work.

Setup guides, failure post-mortems and operating notes for people who already run AI agents, written from running dozens of them side by side every day.

  • Setup
  • Coordination
  • Verification
  • Cost
  • Post-mortems

Latest

New guides and post-mortems

All guides

    The first guides are being written and checked against each tool's own documentation.

    Agent case library

    Open-source agent builds, by the job they do

    Research teams, recruiting, competitor tracking, daily briefings. Each page takes one public, runnable agent project and answers four things, with a link to the original code.

    • What it solvesThe job, and who would use it.
    • How to run itRequirements, keys and the steps.
    • What we observedDated. Read-only pages say so.
    • Where it breaksLimits, cost and what to change.

    Open the case libraryFrameworks compared

      The first case pages are being read, run and written up. Each one will link the original code and state the date it was tested.

      Operating rules

      Five rules that hold when many agents share one codebase

      These come from daily operation, not from a framework's documentation. Every guide on the site applies at least one of them.

      Read the coordination track

      1. Done means checked, not reported

        An agent's summary is a claim. Exit code 0, a green build and an HTTP 200 are signals. The task is done when the result has been read back from where a user would see it.

      2. One writer per file at a time

        Two agents editing the same file is a race. Give each task a clear set of paths, or give each agent its own worktree, and make shared files small, atomic edits.

      3. Commit only your own lines

        A shared working tree holds everyone's unfinished work. An agent that runs a plain commit ships all of it. Stage by path, check the diff, and never sweep the tree.

      4. Write the handoff down

        The next session knows nothing. A handoff states what is finished, what is half-done, what was ruled out and the exact next step, in a file the next agent will actually read.

      5. Deploy from a state you can name

        Before a deploy, the agent should know what is live, what is in the repository and what is only on disk. If those three differ and it cannot explain why, it stops.

      Post-mortem format

      Every failure written up the same way

      A post-mortem is useful only if you can act on it. Each one here has four parts, and the last one is never optional.

      Read post-mortems
      1. 01 · What broke

        The visible damage

        What the agent did, to which files or systems, and how it was noticed.

      2. 02 · Why

        The cause in the setup

        The missing instruction, missing check or shared resource that made the mistake likely.

      3. 03 · Fix

        How it was repaired

        The recovery steps that worked, including the ones that did not.

      4. 04 · Guard

        What stops a repeat

        The rule, hook or check added so the same failure cannot happen quietly again.

      Questions

      Before you read further

      Who is Agent Talk for?

      People who already run AI agents, such as coding agents in a terminal or agents built on a framework, and want them to finish real work without constant supervision. If you are still deciding whether to try an agent at all, most guides here will assume more than you need.

      Why does my agent say a task is done when it is not?

      Usually because nothing in its instructions defines what done means, so it reports the last step it ran. The fix is a completion check the agent must run and quote before it reports: a test, a build, or reading the result back from where a user would see it.

      Can several agents work in the same repository at once?

      Yes, with rules. Each agent needs a clear owner for the files it edits, commits that contain only its own changes, and a written handoff when work moves between sessions. Without those, agents overwrite and commit each other's unfinished work.

      Do I need a multi-agent framework to run more than one agent?

      No. Many working setups are several independent sessions sharing a repository, an instructions file and a task list. A framework helps when agents must call each other programmatically; it does not replace ownership rules or verification.

      How do I keep agent costs under control?

      Route each kind of task to the cheapest model that passes your checks, keep long stable instructions cacheable, hand noisy searches to a subagent so the main context stays small, and measure spend per task type rather than per day.

      Ask

      Describe what your agent did

      Tell the assistant what you asked for and what happened instead. It answers now, and questions that keep coming up become full guides.