On October 3, at 10:50 in the morning, the plan to add basketball to GameDay Scoreboard was drafted. At 1:35 that afternoon, basketball shipped as version 0.8.0. I didn't write a line of the code, and I didn't review one either.

The model didn't make that possible on its own. The process did. Agents are fast, but speed without structure just gets you to the wrong place sooner. This is the workflow I've spent the last year working through with engineering teams: four steps, two gates, and a loop.

The short version

  1. Plan: what I want to get done, in plain English. The agents interview me and draft it; I own it.
  2. Requirements: how it should be built. The agents write it, and nothing gets built until it's right.
  3. Implementation: the agents build it to the requirements, with a test for every behavior the requirements promise.
  4. Technical docs: what was actually built and how it works today. The agents write it, and nothing merges without it.

Each step produces a deliverable, and three of the four are written down. That's the whole trick. Agents don't remember last week, and neither does a team, really. Documents do.

1. Plan: the what

The plan is mine, but I don't type it from scratch. An agent interviews me: what are we building, who is it for, what have we learned, what's out of scope. It drafts the plan from my answers, and I review it until it says what I mean. Nothing moves until I approve it.

A plan reads like a story: the context, what we want from the user's point of view, the decisions that matter and why, what "done" looks like, and what's out of scope. It has no file paths, no code and no library choices. The moment I start telling agents how, I'm doing their job and skipping mine.

My job in the plan is judgment, because agents are very good at solving the problem you give them and terrible at noticing it's the wrong problem. The basketball plan is a good example. The most important line in it isn't technical at all: points come in twos and threes every few seconds, so parents will fall behind and copy the gym scoreboard at the next timeout. That makes "catch up" a main way to keep score, not an error-correction feature. No agent was going to figure that out from the code. I figured it out from the bleachers.

The plan also carried lessons from the sports we'd already built, like settling the screens and the rules before writing requirements so nothing gets built twice. That's institutional memory, written down where the next cycle will find it.

2. Requirements: the how

This is where people get surprised: I don't write the requirements. The agents do.

They read the plan, then read the existing technical docs to see how the system works today. Then they lay out exactly how the new work fits: what changes, which patterns to follow, how the API and the web app talk to each other, every edge case, and how each piece gets tested. No lines of code, just how it should be coded.

For basketball, that came to 1,255 lines: every scoring rule and edge case (overtime, ties, catching up mid-period), the exact contract between the API and the web app, each screen and what it does in every state, and a named test for each behavior. Each agent wrote its own section of one shared document and reviewed the others' side of the contract. That volume isn't padding. It's every decision made once, in writing, instead of guessed at during the build.

GameDay's repo is private, but if you want to see what this looks like, a public project of mine, stack-benchmark, has its full requirements and technical docs in the repo.

Every requirements doc must also end with a section called Technical Documentation Follow-Through: every doc that will need to be created, updated or retired once the work is done, and why. If it's missing, the requirements go back.

Gate one: alignment

Then comes the first hard gate. A coordinator agent, whose whole job is to run the workflow, cross-checks the requirements against the plan before I ever see them:

  • Is every part of the plan covered?
  • Is anything quietly dropped?
  • Is there any ambiguity an agent would have to guess at?

Then it brings me what's left: every open doubt, each with a proposed answer, not just a list of problems. I make the calls, the requirements get updated, and I approve them. Nothing is built before that.

This is where the team, human and agent, gets aligned. The rule I use: by implementation, no known product or architecture question should be open. Building still turns up surprises. When it does, the agent doesn't decide quietly and move on. The question goes back through the requirements, gets answered, and gets written down. If agents keep hitting big decisions mid-build, the requirements were too thin, and that's my fault, not theirs.

3. Implementation: the easy part

This used to be the whole job. Now it's the shortest step.

Every piece of work gets its own isolated copy of the repo, a git worktree, so parallel work doesn't collide. The agents build to the requirements there and run their own tests. The bar doesn't move: every behavior in the requirements has a named test, and no change breaks an existing one. The goal isn't a coverage percentage. It's that if something the requirements promise ever stops working, a test fails.

Because the requirements did the thinking, implementation is mostly execution: building what's already been decided. That's exactly the part agents are best at.

Pro tip: you don't need a high-end model for this step. Thorough requirements make implementation a good fit for cheaper models, or even local ones running on your own machine. Save the expensive models for planning, requirements and review, where the judgment happens.

Gate two: the ship check

The second gate is a script, not a person, so nobody can talk their way past it. Before anything merges, it rebases the work onto the main branch and runs the full test suite. It refuses to merge if any test fails, and it refuses to merge a change that touched behavior without touching the technical docs.

That second refusal is the one that matters most.

4. Technical docs: the most important step

After the code is written, the agents document what they actually built: how it works today, what changed, and where it lives. The follow-through list from the requirements says exactly which docs to touch.

In my projects, the technical docs are the source of truth for understanding the system. The running code is what actually happens, but nobody, human or agent, starts by reading all of it. They start from the docs. The plan and the requirements are a record of what we intended, and they're expected to go stale. The technical docs are the maintained map of what's real, kept true on every change.

A map is only useful if it's right, so drift gets caught, not assumed away. The ship check refuses any change to behavior that doesn't touch the docs. Review agents compare the docs against what was built. And when the docs and the code ever disagree, that's a bug, fixed before anything else.

The loop

Here's why step 4 matters so much: it feeds step 2.

When the agents write requirements for the next feature, they don't start by reading the raw code. They start by reading the technical docs. An agent that starts from good docs understands the system in a few thousand tokens instead of a few hundred thousand, and it starts with the intended structure and fewer guesses.

So every cycle leaves the repo easier to work on than it found it. Basketball was the fourth sport. It went from plan to production in an afternoon because the first three left behind docs explaining exactly how a sport plugs in.

Most software works the other way. As it grows, every new feature leans on decisions nobody wrote down, until the system becomes a house of cards that nobody dares touch. That's what usually happens on large teams. This process flips it: because every change has to leave the docs true, the system gets easier to maintain as it grows, not harder.

That's what "scales" means here. Not that agents type faster, but that each round makes the next one cheaper.

The process is a product too

This workflow didn't arrive finished. It's built on principles software teams have scaled with for years, and it improves with every project, the same way the code does: when something goes wrong, the process gets a new rule.

Worktrees are the best example. They weren't part of the process at first, and one feature in progress could block everything behind it. On one bad day, half-finished design work stopped a release, a change committed straight to the main branch broke a test for every other session, and two sessions editing the same backlog created duplicate IDs.

So the process changed. Now every change gets its own worktree, even a one-line fix, and a git hook refuses commits that try to skip it. Features move in parallel, and nothing reaches the main branch except through the ship check. That one change did more for throughput than any model upgrade.

When to skip it

Not everything needs all four steps. A quick fix that follows existing patterns, touches a handful of files and adds no new behavior skips the plan and the requirements. It goes straight to an agent.

It never skips the docs. If behavior changed, the docs change. That's the one rule I don't bend, because it protects whoever has to understand the system next, person or agent.

What's actually new here

None of this is a new idea. Good teams have worked this way for decades, and the ones who skipped steps usually paid for it later. What's new is the economics.

Writing 1,255 lines of requirements used to cost a week, so teams skipped it. Keeping docs current used to cost a person, so nobody did. Tests got cut the same way: clients didn't want to pay for them, and teams never had the time. Now agents write all three in a fraction of the time.

That matters most for tests. Here, a test for every promised behavior is a hard requirement, and it's a big part of what keeps the system stable as it grows. Because agents write them, they're cheap enough to actually get written, on every change.

The human time goes where it should: the plan at the start and the judgment at the gates.

The agents do the labor. You own the what, the gates, and the outcome. The next part is about that last one: how to own code you didn't write.