From Squad to Pod: Who Does What in AI-DLC
AI-DLC changes roles more than it changes tools. Humans decide, the AI operates, squads shrink toward pods, the manager becomes the keeper of pace, and the engineer still owns the code in production. What each role does now, and the competency-matrix problem nobody has solved.
The agent writes the code. You still get paged when it breaks.
Most of what gets written about AI-DLC is about the machinery: phases, stages, agents, gates. The machinery is the easy part. What actually changed for the teams I watched adopt it was who does what, and whether the organization noticed in time.
AI-DLC is the AI-Driven Development Life Cycle, a method published by AWS in which an AI agent plans the work and writes the artifacts while humans approve or reject them at gates. I followed its rollout inside a large engineering organization, from pilot squads to the people trained to spread it. The tooling questions got answered in weeks. The role questions are still open, and they are the ones that decide whether the method survives.
Humans decide, the AI operates
The cleanest description of the split came from one of the people who brought AI-DLC into that organization, in the first session of the training: the AI handles the operational work, and the humans keep the definition of the problem.
The split of the work
| The AI does | The human does |
|---|---|
| Refines a vague intent into structured questions | Defines the problem and says what is out of scope |
| Writes requirements, stories, designs, and plans | Judges each artifact at the gate: Approve or Request Changes |
| Decomposes the work into Units and tasks | Decides architecture and product trade-offs the AI can only propose |
| Generates code and tests | Builds the guardrails that tell the AI when it is wrong |
| Keeps the state and the audit trail | Corrects the course before an error becomes an incident |
Read the right column carefully, because it is not less work. It is different work. Judging a requirements document you did not write takes more domain knowledge than writing one, because you have to spot what is missing, and missing things do not show up on the page. The job moved from producing to deciding, and deciding is the part that never got cheaper.
Squads shrink toward pods
The second change is the size of the team around a piece of work.
Agile squad
- 01Four to six people per squad
- 02Humans start the work, AI assists
- 03Two-week sprints
- 04Daily, planning, review, retro
AI-DLC pod
- 01One engineer and a PM, plus data or design when the work has that surface
- 02The AI starts and drives, humans validate at gates
- 03Bolts of hours or days
- 04Mob Elaboration, Mob Construction, approval gates
A pod works because the agent covers the operational breadth that used to need more hands. It does not work by cutting people and hoping. The organizations that treated pods as a headcount move got the worst of both: fewer people, the same process, and an agent producing more work than anyone had time to review.
The daily changes too. When the state of the work lives in a file the agent updates after every step, the standup stops being a status report. The team guide in that rollout set it at five minutes. Planning and Mob Elaboration also stopped being the same meeting: planning decides what to work on, the mob decides what that work actually is.
Who sits in the mob
Mob Elaboration is the AI-DLC ritual where the team and the agent turn an intent into requirements, stories, and Units in one session. The rule that made it work in practice was strict: only people with decision power over some dimension of the work join. Observers inflate the cost without adding context.
Mob composition by kind of work
| Kind of work | Who joins |
|---|---|
| User-facing feature | Engineering, product, design |
| Data or machine learning | Engineering, data, product |
| Pure technical change | Engineering and an architect |
| Algorithmic optimization | Engineering and a domain specialist |
That caption is a lesson from a pilot, not from the method. After the agent reverse-engineers the code, it asks a long list of questions, and most of them are technical. The pattern that stuck was simple: engineers answer the technical questions on their own and pull product in for the business questions, in short blocks. The full ritual is in the Mob Elaboration guide.
What each discipline does now
Before and after, by discipline
| Discipline | Before | With AI-DLC |
|---|---|---|
| Engineering | Writes the code | Validates and decides at gates, owns what ships |
| Product | Writes the PRD | Brings a well-shaped intent, answers the business questions |
| Design | Produces wireframes | Validates the UX the agent proposes, inside the mob |
| Data | Analyzes after deploy | Defines the success metrics before Inception starts |
The product row is the one teams underestimate. AI-DLC assumes the intent arrives in good shape, and the method says almost nothing about how. When product brings a vague request, the agent fills the gaps with generic assumptions and the whole downstream inherits them. That gap gets its own article.
Backend engineers ship front-end now
One principle in the original AWS whitepaper is to streamline responsibilities: with the AI doing decomposition and generation, developers can work across the old silos of front-end, back-end, infrastructure, and security. The pilots I followed reported exactly that: backend engineers delivering front-end work, and new hires shipping faster with less onboarding, because the agent carried the codebase context they did not have yet.
It is a real gain and it has a cost. An engineer can now ship a screen in a stack they could not have written by hand last month. That is fine until the screen breaks in production and the person on call cannot read the code. Convergence widens what an engineer can deliver. It does not widen what an engineer understands, unless you make room for that.
The engineer still owns the code
Here is the sentence I heard in the very first session of the training, and it is the one I would put on the wall: if this goes to production and causes an incident, the one who gets called is not the AI. It is us.
AI-DLC does not move accountability. It moves authorship. The agent wrote the code, and a named human approved it at a gate, so a named human owns it. The risk the pilots flagged early is that people start focusing on how to use the AI instead of understanding the code, and ownership quietly erodes.
The research is starting to describe the same thing from the outside. A survival analysis of over 200,000 code units in 201 open-source projects, “Will It Survive?”, found that agent-written code is modified less often than human code, and concluded that the bottleneck for agent-generated code may not be generation quality but “the organizational practices that govern its long-term evolution.” A second study of AI-generated files across 100 repositories found they receive less frequent maintenance, and that human developers perform the large majority of the maintenance they do get. Code that nobody feels they wrote is code that nobody wants to touch.
The manager keeps the pace
The role that changed most was the one nobody planned for: the engineering manager.
AI-DLC produces artifacts in seconds and asks a human to judge each one. Without someone setting the pace, the pressure to keep up with the machine wins. In one of the training sessions an engineer described reaching the last documents of an Inception and realizing they had no attention left to review them, and that if they kept going they would start approving anything. The fix the pilot teams converged on was organizational, not technical: sessions of about an hour, then a pause, and the state file means the pause costs nothing.
Someone has to enforce that, and a tech lead often cannot, because stopping looks like slowing down. So the manager joins the first two or three mobs of every new team. Not to watch. To decide when to stop, to notice when approvals come without questions, and to tell the J-curve, the expected dip in cycle time while a team learns, apart from a real structural problem. Presence without intervention is worse than absence, because it creates the illusion that someone is supervising. After a few cycles the team should regulate its own pace, and if the manager is still needed in every mob, something else is wrong. The curve itself is in why AI-DLC makes you slower first.
Seniors slow down, juniors speed up, and that is the problem
The pattern that worried me most showed up in the upstream work, where the agent drafts the documents that define a feature. Senior people with deep domain knowledge interrogated every claim the agent made and took longer. People with less domain knowledge accepted the first output and moved fast. The person who knows the least trusts the most.
That inverts the usual assumption that AI helps juniors most. Speed at a gate is not a sign of skill. Often it is the opposite. The answer the pilots reached was not to keep juniors away from AI-DLC, but to stop putting them at gates before they had the context to judge: study time and onboarding first, then a gate, and for a while, a senior reviewer next to them.
Your competency matrix rewards the wrong thing
This is the open problem, and I have not seen anyone solve it.
A manager raised it at the end of the training. Competency matrices reward technical depth in code: writing it, debugging it, designing it. In AI-DLC the agent writes the code. The engineer’s value moved to specifying, reviewing, deciding, and catching code that works and is still wrong. If the matrix keeps rewarding volume of code, engineers will optimize for volume, and in an agentic workflow that means approving generated code without reading it.
The example from the pilot sticks with me. The agent generated a process that fired hundreds of database queries for one operation. It worked. The tests passed. It would not have survived production load. It was caught in code review, by someone reading the generated code line by line, the slow way. That kind of catch is among the most valuable things an engineer does now, and almost no matrix I know would score it.
Behaviors worth rewarding in an AI-DLC team
- Required:Specs that an agent builds correctly on the first pass.Rigor of specification is now the main input to code quality.
- Required:Rejections at the gate with a clear reason.A Request Changes that stops a wrong design early is worth more than ten approvals.
- Required:Catching code that works and is wrong.Performance traps, security holes, logic that passes the tests and misses the requirement.
- Required:Fixing the cause, not the symptom.When the agent errs, going back to the spec or the rules instead of patching the code by hand.
- Required:Owning what was approved.Being able to explain, debug, and run in production the code you signed off on.
- Anti-pattern:Removing technical depth from the matrix.The wrong fix. You still need to understand code to review it. The emphasis shifts. The depth stays.
The way to get there is not a new framework from HR. It is managers who have sat in the mobs writing down which behaviors they saw that the current matrix ignores, and adjusting it in small steps. The measurement side, what to track instead of story points, is in after story points.
The agents are a role map
AI-DLC 2 makes the human side visible in a way that is easy to miss. Its 14 agents are named after roles: product, design, delivery, architect, AWS platform, compliance, DevSecOps, developer, quality, pipeline and deploy, operations, plus two reviewers and a composer. The documentation calls the philosophy “Small Mob, Broad Agents” and explains it as mirroring how effective human teams work: a mob of three to five people covering a whole feature, each with broad skills instead of one narrow specialty.
That is a useful mirror for a leader. For every agent on that list, ask which human in your pod judges its output. If nobody on the team can judge what the compliance agent or the DevSecOps agent produces, the agent is not adding coverage. It is adding unreviewed output. AI-DLC 2 also supports team mode, where people claim individual Units and build them in parallel in their own checkouts. More parallel work makes the question sharper, not softer.
FAQ
Does AI-DLC eliminate roles?
It changes them more than it removes them. The operational work of writing documents and code moves to the agent. Defining the problem, judging the artifacts, deciding trade-offs, and owning production stay with people. Teams get smaller around a piece of work, but the judgment they need does not shrink.
Who approves the gates in AI-DLC?
The person with decision power over that dimension of the work: an engineer or architect for design and code, product for requirements and stories. AI-DLC 2 records each approval in an audit log. Make sure every approval has a named owner and enough time to actually read what they approve.
Do we still need QA engineers?
Yes, with a different job. AI-DLC 2 has a quality agent that writes test strategy and tests, which removes the typing, not the judgment. Someone still has to decide what good enough means, find the cases the generated tests miss, and own the release.
How should juniors work in AI-DLC?
Give them study and onboarding time before putting them at a gate, and pair them with a senior reviewer at first. The pilots showed that people with less domain knowledge tend to accept the agent's first output. Speed at a gate is not a sign of skill.
How do we evaluate engineers who no longer write most of the code?
Nobody has a finished answer. Start by rewarding what the work now needs: specs the agent builds correctly, well-reasoned rejections at gates, catching code that works and is wrong, and ownership of what you approved. Let the managers who attend the mobs adjust the matrix from what they observe.
Where to go next
AI-DLC moved the code to the agent and left the accountability with people. The organizations that do well with it are the ones that redesign the human side on purpose: smaller pods, a manager who protects the pace, gates with named owners, and a matrix that rewards judgment instead of keystrokes. The ones that skip that part get faster code and nobody who can answer for it.
The newsletter
Don’t Code, Specify. A weekly dispatch from where AI agents meet real production. No hype, just what shipped and what broke.
Subscribe on Substack (opens in a new tab)