Here's a confession that makes some engineers wince: I don't read the code my agents write. On my scoreboard app, I didn't write a line of it, and I haven't reviewed one.
That sounds reckless. The strongest critic of agentic coding I've read, Alex Ewerlöf, puts the objection plainly in Coding is NOT solved: "You cannot be accountable for what you don't understand." He's right. So this post is about how I understand, and own, software whose code I've never read, and where that approach stops being enough.
The short answer: ownership never came from reading the code. It came from the system around it. How much of that system you need depends on what's at stake. My scoreboard app has no accounts, no payments and no sensitive data, and that's a big part of why gates alone are enough there.
You've done this before
For years I was a CTO running engineering teams of more than 30 people across offshore and nearshore time zones. I didn't read every line they wrote. No CTO does. Nobody could.
I was still accountable for all of it. When something broke, it was my problem. So how did I own code I never read? The same way every engineering leader does: clear specs, agreed designs, tests, reviews, demos, documentation, and using the product myself. I owned the system that produced the code, and that system is what made the code trustworthy.
You've done a version of this too. You've never read the source of your database or most of the packages your app depends on. But be honest about why you trust them: mature dependencies have outside maintainers, years of public scrutiny, release histories and security reporting. Code an agent wrote this morning has none of that. That's exactly why the system around it has to carry more weight, not less.
Understanding a system isn't reading its code
Ewerlöf's other line is the one I keep coming back to: "AI can explain it to you but it cannot understand it for you."
I agree, and I'd add that reading the code was never how anyone understood a large system. Even great engineers don't know every line of the codebase they work in. What they know is how it fits together and where to look.
That's what the technical docs are for in my workflow. They're the maintained map of what was built and how it works today, and they're the same docs the agents read first every time I start a new plan. That's what keeps them honest: if they drifted, the next round of work would start from a wrong map. When I need to understand part of the system, I start with the docs too. They tell me how it works and where in the code it lives. When the docs and the system's behavior disagree, the agents go to the code, and whichever one is wrong gets fixed.
So I do understand my systems. I understand them from the docs and from using them, the same way I understood the systems my teams built.
The layers that replace code review
Code review is one person reading a diff, usually in a hurry, and hoping to spot a problem. What I have instead is a stack of checks, each catching something the others miss:
- The plan. I own what it should do. An agent interviews me and drafts it, and nothing moves until I approve it.
- The requirements gate. Before anything is built, the requirements are checked against the plan and every open question is answered. Reading the requirements is reading the design, and it's where most mistakes are cheap to fix.
- Tests. Every behavior in the requirements has a named test, and no change may break an existing one.
- Review agents. Separate agents review the requirements and the work against them before I ever see it.
- Browser automation. Agents drive the real app in a real browser, end to end, the way a user would.
- The ship check. A script refuses to merge anything with a failing test, or a change to behavior without a change to the docs.
- Using it for real. I test the product myself, in the real world. For my scoreboard app, that means keeping score at an actual game from the bleachers.
Keeping the checks independent
There's a real weakness here, and it's worth naming. If one agent writes the code, the tests and the docs, a single misunderstanding can pass every check. The code is wrong, the test expects the wrong thing, and the docs describe it confidently.
So the layers can't all come from the same place. The tests come from the requirements, which were written and approved before any code existed, not from the code itself. Review agents work in separate sessions and check the work against the requirements, not against the implementation. Browser automation tests behavior from the outside, the way a user would. And the check furthest from any agent is me, using the product for real. It won't catch everything, like a security hole or a slow query, and it can share my own assumptions, but it sees the product the way users do.
That doesn't eliminate correlated mistakes. It makes them much less likely to survive.
When something slips through
It happens. A bug gets past the tests, or a flow works in the browser automation and still feels wrong at a real game.
I don't think of those as getting burned. When something slips, I fix the bug, and then I fix the process that let it through: a missing test, a gap in the requirements checklist, a rule the agents didn't have. That rule gets written into the instructions every agent reads, so the same mistake is much less likely to happen the same way twice. We train the agents, not the models.
Every slip makes the system stronger. That's the difference between owning code and owning a process. A bug in code gets fixed once. A gap in the process, once closed, protects every feature after it.
When it breaks in production
Owning code you didn't write also means owning it at 2 a.m. That's where the docs pay off a second time.
When something goes wrong in production, the agents that built the system aren't the ones I send to diagnose it. Separate operations agents, ones that never touched the code, start from the same technical docs. That's how they learn how the system is supposed to behave. Then they do something agents are very good at: read through far more logs and metrics than a person could, and spot what's out of place fast.
The same goes for people. An SRE team that reads the technical docs can understand the system as well as the team that built it, without starting from the code. And because the operations agents aren't the build agents, they're one more independent check. They look at what the system actually does, not what its authors believed it would do.
Starting from the docs doesn't mean never opening the code. Some incidents, like a race condition, a resource leak or a dependency behaving badly, only make sense at the implementation level. When that happens, the operations agents follow the evidence from the docs and the logs into the code.
Spend your attention by risk
How much human review you need depends on two things: what's at stake, and how much you've learned to trust your own system. Here's where I'd start as the stakes go up:
| What you're building | What earns your trust |
|---|---|
| Personal tools and prototypes | The gates, the tests, and using it yourself |
| Ordinary product features | All of the above, plus independent review and production monitoring |
| Authentication, payments, permissions, data migrations | Review agents first, then human review by someone who knows the domain. The human review stays. |
| Regulated or safety-critical systems | Formal controls and domain-specific assurance. This post isn't your playbook. |
I run even that third tier through agents, because I trust the system I've built around them. I don't recommend that as anyone's starting point. Here's how I'd set it up: review agents go first, flag the risky parts of a change and explain them, so the human reviewer knows exactly where to look. A human who knows the domain stays in the loop for authentication, permissions, payments, destructive migrations and secrets. With good agents, that review gets much faster, but it doesn't go away because recent reviews found nothing. A quiet stretch doesn't prove the problems are gone, and the models, tools and prompts underneath keep changing.
Every time the human catches something the agents missed, it becomes a rule the review agents follow from then on. And whenever the models or tools change, recheck the review process itself.
The argument isn't that nobody should read code anymore. It's that reading every line was never the best use of anyone's attention, and the time it saves should go where the risk is.
What you're actually accountable for
When you don't type the code, it's tempting to feel like you're not responsible for it. You are, completely. What changes is what you're responsible for:
- That it does the right thing. That's the plan.
- That it was built the way you agreed. That's the requirements and the gates.
- That it keeps working. That's the tests and the ship check.
- That someone can understand it next year. That's the docs.
- That it actually works for real people. That's you, using it.
That's a heavier load than "I read the diff and it looked fine." It's also the job engineering leaders have long had.
This only works if the layers are real. Tests that check nothing, docs nobody keeps current, a gate everyone waves through: then you really are shipping code nobody understands, and Ewerlöf's warning applies in full.
You can't delegate accountability. You can build a system that earns it.
Next: how to bring a whole team through this shift, in Leading the Switch.