Most engineers I've worked with started the same way. They tried vibe coding. They described an app in a sentence, watched an agent build it in minutes, and felt like magic was happening. Then the app grew, the agent started breaking things it had built the day before, and the magic turned into a mess. Most of them decided agents don't work for real software.
They were half right. Vibe coding doesn't scale. But the answer isn't a better prompt or a smarter model. It's a different way of working, and you can learn it in a week.
This is a field guide for that week. Each day has one job, and each job has a clear "done when."
Why vibe coding hits a wall
Vibe coding works on day one because the whole project fits in the agent's context window. Then the project grows, old decisions fall out of the window, and nobody wrote down why things were built the way they were. The agent starts guessing, and it guesses wrong. You can't follow it either, because you didn't write the code and nothing explains it.
Pointing an agent at a big existing codebase fails for the same reason: no real codebase fits in a context window. It doesn't fit in your head either. Great engineers don't know every line. They know where to look. That's what this week teaches your agents to do, using documents.
Before Monday: the setup
You need less than you think:
- One agent tool. Claude Code, Codex, or something like them: a tool that can read your repository, run commands and make changes, not just autocomplete.
- One subscription to the model behind it.
- A fresh repository with git and a test runner.
- A short instructions file at the root of the repo. Most tools read one automatically, like
CLAUDE.mdorAGENTS.md. Five lines is enough to start:
# Project rules
- Follow the workflow: plan → requirements → implementation → docs.
- Plans live in docs/plans/, requirements in docs/requirements/, technical docs in docs/technical/.
- Never start implementation until the requirements are approved.
- Run the tests before calling anything done. Never break an existing test.
- When behavior changes, update docs/technical/ in the same change.
That file is how you'll keep the agent on the process when it, or you, wants to skip a step. It's a starting point for the workflow, not a security policy.
Pick the right first project
The first project matters more than the tools. Pick something that:
- Solves a problem you actually have. You'll know what "done" means, and you'll keep going when it gets hard.
- Starts from scratch. Don't learn the workflow and a ten-year-old legacy codebase at the same time.
- You can finish in a week. One or two screens, or one small service.
- Is low stakes. No payments, no sensitive data, nothing other people depend on yet.
- You'll use. Using it is half of how you'll know it works.
Monday: the framework
Spend day one understanding the four steps and why they exist: plan, requirements, implementation, technical docs, with gates in between.
It clicks faster than people expect, because none of it is new. Write down what you want. Agree on how before you build. Test what matters. Document what you built. If you've worked on a team that shipped well, you already know this. The only difference is who does each step.
Tuesday: the plan
The deliverable: docs/plans/<project>.md
Have an agent interview you about what you're building and why, then draft the plan from your answers. Edit it until it says what you mean. A good plan covers:
- Context: the problem, and who has it
- What we want: what the user can do when it's done
- Decisions: the choices that matter, and why
- Done when: how you'll know it works
- Out of scope: what this version won't do
- Open questions
No file names, no code, no library choices. That's tomorrow's job.
Done when: a stranger could read it and know what you're building and why.
Wednesday: the requirements
The deliverable: docs/requirements/<project>.md
Have the agent write the requirements from the plan: how it fits together, what each piece does, the edge cases, and how each behavior will be tested. Then check it yourself:
- Is every part of the plan covered?
- Does every behavior have a test described?
- Are the edge cases named, like empty states, bad input and what happens offline?
- Is every open question answered, with nothing left for the agent to guess?
- Does it list which technical docs will need writing?
Done when: you could hand it to someone else to build, and they wouldn't have to ask you anything.
Thursday: the build
The deliverable: working code, with a test for every behavior.
Let the agent implement the requirements. Your job is not to watch it type. Your job is to check the result: do the tests pass, and does the thing actually do what the plan said?
If the agent hits a question the requirements didn't answer, don't let it decide quietly. Answer it, add the answer to the requirements, and keep going.
Done when: all the tests pass, and you've used it yourself.
Friday: the docs, and the acceptance check
The deliverable: docs/technical/
Have the agent document what it actually built: how it works, where things live, how to run it and test it.
Then run the acceptance check:
- The tests pass from a fresh checkout.
- You've used it for a real task, not a demo.
- The docs explain how it works without reading the code.
- A brand-new agent session can answer "how does X work?" from the docs alone.
That last one is the real test. If a fresh agent can understand your project from the docs, so can you in six months, and so can anyone you hand it to.
A failure, and the fix
My own first failure is one most people hit. I'd ask an agent to build something on day one, and it would. On day two, in a fresh session, I'd ask for the same kind of task and get a completely different result: different structure, different patterns, different code, all reaching the same goal another way.
At first I thought the model was hallucinating. It wasn't. It just didn't know what it had decided the day before. Every new context window starts from zero, and with nothing written down, the agent found a new path each time. Two days in, my project already had two ways of doing everything.
So I looked at what works in software teams, and what breaks them. A context window turned out to have the same problem as expecting one engineer to remember everything and code anything. Teams solved that a long time ago: you write it down. The decisions go in the plan, the how goes in the requirements, and what was built goes in the docs. Once the agent started every session by reading those, day two looked like day one.
That's why Wednesday and Friday matter so much. The requirements are where a decision gets made once. The docs are how tomorrow's session remembers it.
Week two: add a feature
Now come back and add one feature, starting from the docs, not the code. Write a short plan, let the agent write requirements from the plan and the docs, build, and update the docs.
This is usually the moment it clicks. The agent picks up a project it has never seen and gets it right, because the docs told it where to look.
Once that works, and you're running more than one agent at a time, consider a coordinator: one agent whose job is to take your requests, follow the workflow rules, and hand the actual work to a few specialists split along real boundaries, like your API and your UI. Each specialist gets only the context for its task. Don't start there, though. One agent, one repo, one small project comes first.
The mistakes most people make
You'll make some of these.
- Treating it like autocomplete. Using agents for snippets instead of whole, well-defined tasks.
- Vague asks with no plan. "Add user accounts" in one sentence, then blaming the agent for guessing wrong.
- Micromanaging the how. Dictating files and functions instead of describing the outcome.
- Trusting without checking. Accepting output with no tests and no verification.
- Vibe coding past its limit. Fine for a prototype. Past a certain size, nobody knows how the thing works.
- Over-engineering the agents. An elaborate multi-agent setup before you've shipped anything.
- One giant context window. Running everything through one endless conversation instead of focused tasks.
If you're skeptical
The most common objection hides behind all the others: I tried it, and it fell apart. Almost always, what fell apart was vibe coding on a growing project, or asking an agent to understand a big codebase in one go. That fails for a real reason. You were asking the agent to hold an entire system in its head, which no engineer does either.
Give it what good engineers rely on: a plan, a spec, and docs that say where to look. Then try again.
By Friday
If you follow this, by the end of the week you'll have done the whole job once: planned something, specified it, had it built, checked it, and documented it, without writing the code yourself.
It won't feel like magic, and that's the point. It'll feel like engineering, with a lot more leverage.
The next part goes deep on the workflow itself: The Workflow That Scales.