Getting one engineer to work with agents is easy. Getting a team there is the hardest thing I've done in this whole shift, and I'm still learning how.
The tools aren't the problem. The workflow isn't the problem. The human transition is the hard part, and not because people are slow or difficult. It's because this asks a lot of them, and some of what it asks is personal.
The aha moment is the whole game
Everything depends on one moment: when someone watches something that failed for them suddenly work.
Usually they've already tried agents, vibe coded something that fell apart, and filed agents under "hype." Then they see the same kind of project built with a real workflow, and it keeps working as it grows.
The second half of the aha matters just as much: it isn't some strange new way of building software. It's what has worked on good teams for decades. The only thing that changed is who does the typing.
That moment can come top-down or bottom-up. A leader can bring it to the team, or one engineer can bring it to their peers. What I've seen work best is two steps:
- Show it to a group. Get people in a room and walk through the process on something real. The questions are the point. Every doubt someone says out loud is one you can answer in front of everyone.
- Then have each person build something. Watching isn't enough. The aha really lands when someone sits down with a problem of their own, digests it, and builds the solution through the process, start to finish. That's the same advice I give individuals in your first week, and it holds for teams too.
Start with a pilot
Don't roll this out to everyone at once. Start with a pilot: a team lead and a couple of volunteer engineers. Volunteers matter, because people who want to learn this will push through a rough first week. But the real requirement is that they can judge the quality of what comes out, both the code and the docs.
Make the pilot something of real value to the organization, something already sitting in the backlog. A pilot that proves the process on work the business actually wants is far more convincing than a demo. And pick work with as few dependencies as you can. A greenfield project makes the team's life much easier, because there's no history to untangle.
If the pilot has to be an existing project, start backwards. The workflow runs on technical docs, and most real codebases don't have them. So before any new work, backfill the docs. There's no shortcut: agents analyze the code a piece at a time, and the team lead guides them with what the team already knows, like why that service exists, what that weird module protects against, and which parts nobody dares touch. The team reviews every doc against the actual code. Once the agents have an accurate picture of the system, start the workflow: add a feature, fix a bug.
Either way, review the output together as a team at every stage, from the plan to the requirements to the code to the docs. Spend the most care on the docs. Tests catch the behaviors they exercise, so broken code tends to show itself. A doc that misrepresents the code is much harder to catch, and a wrong doc misleads every agent and every person who reads it afterward. When the team's reviews consistently stop finding problems, step back to spot checks and let the gates carry more of the load.
Guardrails that grow with trust
Most guardrails you add as you learn how much to trust the output. One has to be settled before the pilot starts.
Access, decided up front. Coding rules aren't access control. Before agents touch company code, decide:
- Which models and tools are approved, and what happens to your source code and customer data when they're used (retention and training policies)
- Where secrets and production credentials live, and which ones agents never get
- What agents can reach on the network and the filesystem
- How agent work is logged and attributed, so you can tell who, or what, changed something
- How new dependencies and their licenses get reviewed
- Who is allowed to deploy to production
None of that needs to be elaborate for a pilot, but every item needs an answer.
Separate workspaces. In my setup, every piece of work, from the plan to the finish, happens in its own git worktree: its own branch and working copy, like a single engineer with their own checkout. Testing and validating an agent's work doesn't interrupt anyone else's, and rebasing between agents and sessions becomes routine. But a worktree isn't a security boundary. An agent working in one can still reach the rest of the machine, the network and any credentials it's given. If agents can run commands or touch sensitive systems, give them least-privilege credentials and run them in a real sandbox, like a container or a virtual machine, with the worktree inside it.
Specialist agents that carry your team's standards. If you want control over how the code gets written, this is where you take it. Create stack-specific agents for implementation, like a Python agent or a TypeScript agent, and put your team's existing standards in their instructions: the patterns, the rules, the things your senior engineers enforce in code review today. A Python agent's instructions might include:
- Read before writing. Explore the existing code and match its patterns instead of introducing new ones.
- Make the smallest correct change. No drive-by refactors, no unrequested features.
- Verify, don't assume. Run the tests, the linter and the type checker before reporting done.
- No fake passes. Never weaken an assertion, skip a test, or special-case test inputs to get a green run.
- No hallucinated APIs. Confirm a function exists in the installed package before using it.
- When a requirement conflicts with the code, stop and ask instead of guessing.
- Parameterized SQL only. Never hardcode secrets.
The coordinator delegates implementation to those agents, so every session starts from the same rules. That's how a team's discipline survives the switch: not in people's heads or in review comments, but in instructions every agent starts from, every session. Agents don't follow instructions perfectly, which is why the gates still check the result, but written rules make the same drift far less likely. And writing those rules is a natural job for your most senior engineers. Their standards become everyone's.
Measure the process, not the typing
Once the pilot is running, time it. How long does a feature take to go from a plan to a change that's tested, documented and ready to review? How much of that is human time? Those numbers, not lines of code or the number of pull requests, become your team's new performance baseline.
Speed alone isn't the goal, though. Track quality right next to it:
- Rework needed before a change is accepted
- Defects that reach users, and rollbacks
- Requirements reopened during the build
- Doc errors caught by a fresh agent session or by the people running the system
- Human review time, broken down by risk
If speed goes up and these get worse, the process isn't working yet.
One note on the evidence. The rollout in this article is the approach I recommend from working with teams. The numbers that follow come from my own projects, not from a measured team rollout.
Mine still surprise me. I can start with a single sentence, like "add a news feed section to the app." After about ten minutes of planning, mostly the agents interviewing me, the work runs on its own. What comes back is a working change, tested end to end and documented, and it takes me 30 to 45 minutes to review and approve.
That speed came from refining the workflow, and from time spent with the agents making sure they meet my expectations. It also carries over. When I copied the workflow into a new project, my scoreboard app, I had a working product in days, not weeks. When a friend said a feed would be cool, the feed was live the same day, with under an hour of my own time.
The same idea applies to the software itself. Give your agents a benchmark that defines what fast means, and they can measure, analyze and optimize your code against it. I go deeper on that in part 6, The Skills That Matter Now.
What gets in the way
The obstacles are predictable:
- Bad first experiences. Someone tried it a year ago, or tried it the wrong way, and the doubt is already baked in before you start.
- Mixed buy-in. Half the team is all in and half isn't, and a process breaks at the people who don't follow it.
- Legacy code with no docs. Agents need to know where to look, and an old codebase with no documentation gives them nothing.
- Identity. This is the big one.
For a seasoned engineer, the idea that a computer can write better code than they can isn't just wrong. It can feel insulting. Many of the best engineers I know built their identity around their craft. They're proud of their work, and they should be. Asking them to stop typing can sound like asking them to stop being who they are.
I don't think you get past that by arguing. You get past it by respecting it, and by showing them where their skill goes next. The judgment that made them great, knowing what good looks like, spotting the wrong problem, understanding why the system is the way it is, is exactly what this new way of working needs most.
On an existing codebase, that's literal. The senior engineer who feared being replaced turns out to be the person the docs backfill depends on. Their knowledge becomes the foundation every agent builds on. That's not a consolation prize. It's the most valuable work on the team.
What changes for leaders
If you lead an engineering team, two things change.
You own the process. Your job stops being assigning tickets and reviewing pull requests. It becomes the workflow itself: the gates, the rules the agents follow, the standard for plans and requirements, and fixing the process every time something slips through.
Teams get smaller. A few people with a good process can deliver what used to take a department.
The part I struggle with
That second point is the one I haven't made peace with.
The honest truth is that you don't need large teams, departments and layers of job titles to build software anymore. Agents are very, very good at writing specs, gathering requirements, reviewing work, and running the process. That's the work many people spent years building their careers around.
It's personal for me, too. Engineers I used to work with on my offshore teams recently reached out looking for work. I had to be honest with them: I don't need outside engineering help anymore. Agents do that work now, through my process. That was a hard conversation, and I suspect a lot of leaders are about to have it.
The same question hangs over the next generation. If teams hire fewer junior engineers, where do tomorrow's senior engineers come from? I honestly don't know. Judgment has always come from years of doing the work, and a lot of that work is going away.
I don't have a tidy answer for that, and I don't trust anyone who claims to. What I do believe is that pretending it isn't happening helps no one. The kindest thing a leader can do is be honest about it early, and help people move toward the work that's left: deciding what to build, owning the outcome, and the judgment no agent has.
That's what leading the switch really means. Not rolling out a tool. Helping people find where they fit on the other side of it.
The next part is about exactly that: The Skills That Matter Now.