AI-DLC 2: What Changed, and What It Made Obsolete
AWS rewrote AI-DLC. Version 2 ships as a deterministic engine that takes routing, grading, and state away from the model: 5 phases, 33 stages, 14 agents, 11 workflow profiles, a learning loop, verified construction checkpoints. What changed from version 1, which homegrown layers it made obsolete, and what survived.
AWS did not bump AI-DLC to version 2. It rewrote it, and in one release most of the plumbing teams had built on top of version 1 stopped being worth maintaining.
AI-DLC 2 is the second generation of AWS’s AI-Driven Development Life Cycle. The first stable 2.x release, 2.7.0, landed on September 1, 2026. Three weeks later 2.10.0 was out. Version 1 was a set of rules files that told an agent how to walk you through a project. Version 2 is an engine: a native aidlc command that installs one deterministic workflow into seven coding agents, with 5 phases, 33 stages, 14 agents, 11 workflow profiles, and a loop that turns your corrections into rules.
If you are new to the method itself, start with what AI-DLC is. This article is for people who already know version 1, or who are deciding whether to start on version 2.
I followed AI-DLC version 1 inside a large engineering organization, through pilot squads and a training program. The teams there did what every serious adopter does with a young framework: they built layers on top of it to cover what it lacked. Then AI-DLC 2 shipped most of those layers as native features. Watching that happen taught me more about adopting AI frameworks than the framework did, so the second half of this article is about that.
From rules files to an engine in fourteen months
How AI-DLC got to version 2
- Jul 2025The whitepaper
Raja SP publishes the AI-DLC method definition and AWS announces it on the DevOps blog. Ten principles, three phases, no code.
- Nov 2025Open source as rules files
AWS releases the workflows as Amazon Q Developer rules and Kiro steering files. The agent reads the rules and follows the lifecycle.
- Jan to Apr 2026The 0.1 series
Nine releases harden the rules-driven workflow.
- Jun 19, 2026Version 1.0.0
The first stable release. 1.0.1 follows on June 30 and becomes the last of the line.
- Sep 1, 2026Version 2.7.0
The first stable 2.x release. The 2.x line had been built separately; no 2.0 to 2.6 release was ever published as stable.
- Sep 24, 2026Version 2.10.0
Stronger construction checkpoints, coexisting harnesses, worktree fixes. Previews of 2.10.1 ship almost daily.
The version number tells you something on its own. AWS went straight from 1.0.1 to 2.7.0, which means the 2.x line spent months maturing in parallel before anyone was asked to depend on it. When it arrived on the main branch it was already on its seventh minor version.
AI-DLC 2.10.0 in numbers
- 33
- stagesacross 5 phases
- 14
- agents11 experts, 2 reviewers, 1 composer
- 11
- workflow profilesplus a composer for custom routes
- 108
- audit event typesin an append-only log
The engine decides, the model executes
Of everything in version 2, this is the change I would put on the box. Version 1 was rules files: prose the agent read and was trusted to follow. The agent decided which stage came next, whether a step could be skipped, whether its own output was good enough, and what the state file said. When it drifted, nothing stopped it except a tired human at a gate. That is where slop comes from: a model that routes itself, grades itself, and edits its own records.
Version 2 splits the work in two. A deterministic engine, plain code with a handful of subcommands, owns the routing: which stage runs next, which profile applies, when to stop, what a gate is waiting for. It emits a typed instruction, and the conductor (the /aidlc session, which is the LLM) carries it out and reports back. The model still does the thinking inside a stage. It no longer decides the shape of the process. AWS’s own docs put the split in one sentence: the engine owns the routing, the conductor owns execution quality.
Powers the model had in version 1, and who holds them in version 2
| In version 1 the model could | In version 2 |
|---|---|
| Decide which stage comes next, or skip one | The engine routes from a stage graph compiled when the workflow starts. A skipped stage is recorded as skipped. |
| Say the tests passed | A Unit is verified only by the check command a human approved, and only through a receipt the tool writes. A handwritten proof file cannot verify a Unit. |
| Grade its own artifact | Separate reviewer agents judge it, bound to a fingerprint of the artifact, so changed work cannot inherit an earlier approval. Deterministic sensors (linters, type checks) fire on the output. |
| Edit the workflow state | A guard blocks the model's tools from writing session and plan-approval records. State moves through the engine. |
| Approve its own gate | A gate put to you needs a real message from you. The human-presence guard cannot be lowered per workflow. |
| Change the rules halfway through | Rules, stages, and sensors are compiled once per workflow. A new rule applies next time. |
| Lose the links between artifacts | Automated verification gates check traceability between phases before the next one builds on them. |
None of this makes the model smarter. It makes the process less dependent on the model behaving, which is a different and better goal. It is the same lesson I keep landing on in harness engineering: a deterministic check that fails the build beats a paragraph of instructions the agent can ignore. AWS just applied it to the whole lifecycle. Even the packaging follows it: every harness gets the same engine, and the build that produces the seven distributions runs twice and compares the output byte for byte.
The lifecycle grew two phases
Version 1 had the three phases of the whitepaper. Inception turned an intent into requirements, stories, and Units of work. Construction turned each Unit into design, code, and tests. Operations deployed and watched it run. A human approved every step.
AI-DLC version 1
- 01
Inception
Mob Elaboration: the AI proposes, the team refines.
- Workspace detection
- Reverse engineering (brownfield)
- Requirements analysis
- User stories
- Workflow planning
- Application design
- Units generation
Gate: Human approval per stage
- 02
Construction
Mob Construction, per Unit.
- Functional design
- NFR requirements
- NFR design
- Infrastructure design
- Code generation
- Build and test
Gate: Human approval per stage
- 03
Operations
The whitepaper describes deploy, monitoring, and approved runbook actions. The version 1 rules shipped it as a placeholder.
Version 2 keeps those three and wraps them. Initialization runs first, with no human, and sets up the workspace in under a second. Ideation runs before Inception and asks whether the thing is worth building at all: intent capture, market research, feasibility, scope, team formation, rough mockups, and an approval to proceed. Between phases, a verification gate now checks traceability automatically, so an orphaned requirement or a story that maps to nothing gets caught before the next phase builds on it.
AI-DLC version 2
- 0.1 to 0.3
Initialization
Workspace, detection, state. Automatic.
- 3 stages, no gate
- 1.1 to 1.7
Ideation
Is this worth building?
- Intent Capture
- Feasibility
- Scope Definition
- plus 4 more
Gate: Verification gate 1
- 2.1 to 2.9
Inception
Requirements, design, Units.
- Reverse Engineering
- Practices Discovery
- Requirements Analysis
- Units Generation
- plus 5 more
Gate: Verification gate 2
- 3.1 to 3.7
Construction
Design, code, tests per Unit.
- Functional Design
- NFR Design
- Code Generation
- Build and Test
- plus 3 more
Gate: Verification gate 3
- 4.1 to 4.7
Operation
Deploy, observe, feed back.
- Deployment Pipeline
- Observability Setup
- Incident Response
- plus 4 more
The new Inception stage worth noticing is Practices Discovery. Before writing requirements, the agent looks at how your team already works (testing posture, deployment habits, code style), drafts it, has three other agents inspect the draft independently, interviews you about the gaps, and writes the result into the team’s memory. Version 1 learned your conventions by being corrected. Version 2 asks first.
Fourteen agents instead of one voice
Version 1 was one agent following steering rules and playing roles as the rules told it to. Version 2 ships 14 named agents: 11 domain experts (product, design, delivery, architect, AWS platform, compliance, DevSecOps, developer, quality, pipeline and deploy, operations), two reviewers, and a composer.
AWS calls the philosophy “Small Mob, Broad Agents”. The alternative was obvious and wrong: thirty narrow specialists, each owning one artifact, passing work down a chain. That is waterfall with extra steps, and every handoff loses context. Eleven broad agents that each work across several stages carry more of the picture with them.
Most stages still run inline, in your conversation. Four dispatch work: Reverse Engineering as a two-step pipeline (a developer agent scans, an architect agent writes the synthesis), Practices Discovery and Code Generation as subagents, and User Stories as a mob where design, developer, and quality agents write in parallel. The count in 2.10.0 is 29 inline, 2 subagent, 1 pipeline, 1 mob.
The two reviewers matter more than the headcount. After a stage produces its artifact, the product-lead reviewer judges requirements and stories, and the architecture reviewer judges technical design. An adversarial review can send the work back up to two times. And then it stops and hands you the findings, because the reviewer never blocks. The human decides. That rule is the whole method in one line, and I am glad AWS wrote it into the engine instead of leaving it to good intentions.
Profiles replace the one-size lifecycle
The whitepaper’s tenth principle said no workflow should be hard-wired: the AI proposes the depth, the human adjusts it. Version 1 implemented that as a judgment call inside the rules. Version 2 turned it into 11 named workflow profiles (the engine calls them scopes), each a fixed route through the 33 stages with a default depth and test strategy.
The profiles you will actually use first
| Profile | Stages | What it is for |
|---|---|---|
| Classic | 18 / 33 | The version 1 ceremony: Inception and Construction, one approval per stage. The engine's default when you name nothing. |
| Express | 10 / 33 | Requirements already understood. Shortest path to code and tests. |
| Feature | 33 / 33 | A production feature through the whole lifecycle at standard depth. |
| Bugfix | 9 / 33 | A known defect, a focused repair, a regression test, and the deploy. |
Classic is the migration path, and AWS made it the default on purpose. A team that knows version 1 can start on version 2 without learning a new ceremony. Depth and test strategy are now separate dials (Minimal, Standard, Comprehensive), and you can turn either one without switching profiles.
Corrections become rules: the learning loop
This is the feature I would have paid for. Version 1 forgot. You corrected the agent on Monday, and on Thursday, in a new workflow, it made the same mistake, because nothing carried the correction forward except your own memory.
Version 2 keeps a diary for every stage, a file called memory.md, with four headings: Interpretations, Deviations, Tradeoffs, and Open questions. When the agent makes a call the stage instructions did not cover, it writes it down. At the approval gate, the engine shows you those lines verbatim and asks if you want to keep any, plus a free-text “Anything to add for next time?”.
How a correction becomes a rule
Input
You correct the agent during a stage
- DIARYThe agent records the call it made
The stage's memory.md gets an entry under Interpretations, Deviations, Tradeoffs, or Open questions.
- GATEThe gate shows you the candidates
Lines are surfaced as written, no paraphrase. You tick the ones worth keeping and can add your own.
- CHECKA conflict check against org rules
If the new line contradicts an organization rule, you revise it, skip it, or escalate it.
- WRITEThe kept line lands in project memory
It goes to project.md, one click promotes it to team.md. Open questions never become rules.
Output
The rule applies from the first stage of the next workflow. Never in the middle of the current one.
The docs walk through a real example. On a banking project, a stakeholder note said “the transaction shouldn’t duplicate on retry”. The product agent read “transaction” as a database transaction and wrote an ACID requirement. The person meant a payment. They corrected it, the diary caught the interpretation, and the kept rule fixed the term for every workflow after that.
The last detail is the one that shows someone thought about it. A learning never changes the rules mid-run. The engine compiles stages, rules, and checks once, when a workflow starts, and keeps that compiled view until the end. The gates you approved earlier attested to a stable rule set, and the ground does not move under a workflow in progress.
Rules and sensors: AWS now talks about harness engineering
Version 2 splits steering into two halves. Rules are prose instructions loaded before work (feedforward). Sensors are deterministic checks fired on the output, like a linter or a type check (feedback). AWS calls the pair the control loop.
Rules resolve through five layers, and the model is strictly additive: every applicable layer is in context at once, and nothing silently overrides anything.
The five layers of rules
- ORG
Organization
Framework and company defaults: trunk-based development, testing posture, walking skeleton policy.
- TEAM
Team
Practices your team affirmed, including the ones Practices Discovery found.
- PROJ
Project
This project's specialization. Where the learning loop writes by default.
- PHASE
Phase
Rules for every stage in one phase, like documenting two alternatives for each architecture decision in Inception.
- STAGE
Stage
Reserved for a future release.
The repository even ships a Harness Engineer Guide for reshaping stages, agents, rules, sensors, and knowledge without touching the engine code. I have been arguing that the harness is where your edge lives for months. It was good to see AWS put the word in its own documentation.
Construction got verified checkpoints
Version 1 asked you to approve every Construction stage for every Unit. On a feature with six Units and five stages each, that is thirty approvals, most of them on design documents nobody would read twice. Approval fatigue was one of the first things to break in the pilots I followed.
Version 2 changed the default for new solo work with Units. It builds one Unit at a time, in dependency order, and asks you to approve each completed Unit at a verified checkpoint instead of approving every intermediate document. Verified means something specific: during Delivery Planning, the agent proposes a real project check (bun test, pytest, make check), you approve that exact command, and the engine reuses it at every checkpoint. Changing it later requires your approval again. A handwritten proof file cannot verify a Unit.
What Construction asks you in version 2
- Required:Approve the plan for each Unit before code is generated.Plan Approval stays human in every mode.
- Required:Approve the verification command, once.Reused at every checkpoint. Changing it needs a new human approval.
- Required:Approve the walking skeleton, when it is on.The first Unit is the smallest working end-to-end slice, verified by the command before the rest begin.
- Optional:Choose: continue automatically, or review each checkpoint.Automatic skips routine completion questions. It never skips plan approval, the command, or failures.
- Required:Decide on every failure.A failed Unit halts Construction and offers retry, skip, or abort, even in automatic mode.
There is also a team mode. With Unit ownership set to team, people claim Units and build them in their own checkouts at the same time, with their own gate rhythm. Version 1 had nothing like it. Teams that wanted parallel construction had to invent it.
Guards, and who is allowed to lower them
Version 2 adds a Guard Policy per piece of work, with three values: strict, relaxed, and off. It controls five fences: plan approval, review freeze, state transition, reviewer read scope, and human presence. Enterprise defaults to strict. The other profiles default to off, which lets four of the fences stand aside for undirected work, and every time a fence stands aside the audit log records it.
The fifth fence is the interesting one. Human presence cannot be lowered for a single workflow. A gate put to you needs a real message from you. The only way around it is a machine-wide environment variable, which is the right level of friction for “let the agent approve its own work”.
Multi-repo is native
Version 1 was single-repository. Real features rarely are: a back end, a front end, maybe a mobile app, each in its own repo. Teams I watched worked around it by running one session per stack and pointing the agent at the other repo’s directory as reference material, with the contract between them agreed by hand beforehand.
Version 2 supports multi-repo intents. You declare the repository set when the intent is created (or let the engine discover sibling repositories), Construction anchors each git operation to a specific repo, and Reverse Engineering has to complete one full scan chain per repository before the stage can be approved. The workaround became a flag.
The record is now something you can query
Every intent gets its own record folder under aidlc/spaces/<space>/intents/. Inside: the state file, with a six-state checkbox per stage (not started, in progress, awaiting approval, revising, completed, skipped), every artifact, every questions file, and an append-only audit log with 108 event types. Version 1 had a state file and an audit file too. Version 2 made them the engine’s source of truth, and added a runtime graph compiled from the audit log after every transition.
That makes the workflow measurable without new tooling. /aidlc-session-cost prints duration, stage outcomes, and learnings. /aidlc-replay narrates the session for people who were not there. aidlc attest tells you whether a commit’s content still matches what a reviewer approved. I wrote about using that data in after story points.
What AI-DLC 2 made obsolete
Now the part the release notes do not say.
The team I followed most closely maintained an overlay on top of version 1: a set of extensions, commands, and patches that made the method fit their organization. It was good work, built by people who ran the method every day and fixed what hurt. Then AI-DLC 2 arrived and shipped most of it natively.
What teams built on version 1, and what version 2 ships
| Homegrown layer on version 1 | Native in version 2 |
|---|---|
| A shared state file every custom skill read and updated, so a new session could resume after clearing context | A state machine with six-state checkboxes, a recovery breadcrumb before compaction, and /aidlc --resume |
| Patches injected into the core rules without forking them | A plugin system that only adds, never edits the core, with validate, build, and sync commands |
| An automatic retrospective after Build and Test that scanned for where the AI failed | The learning loop: per-stage diaries, gate candidates, rules that apply next workflow |
| A plan to parallelize construction with git worktrees | Swarm execution with one worktree per Unit, and team mode with Unit claims |
| Running one session per stack and pointing at the other repo by hand | Multi-repo intents with per-repo reverse engineering |
| Archiving each finished task's artifacts to start the next one clean | Per-intent record folders, spaces, and intent archive commands |
This is what happens to every layer of plumbing you build on a fast-moving framework. If the gap is generic, the vendor sees the same gap in every customer and ships it. Your clever state handoff becomes their feature, and your code becomes migration work.
What survived the rewrite
Here is the list that did not get absorbed, and it is the list I would protect.
What version 2 did not replace
- Required:The rules against vibe coding.Never edit generated code by hand. Ask for impact before a big change. Get the critique in a fresh context. They came from watching the agent fail, and no engine encodes them for you.
- Required:The ritual before the code.How an intent gets shaped before AI-DLC sees it: the entry document, the curated sources, the complexity triage. Ideation helps. It still assumes someone knows the problem.
- Required:Who sits in the mob.Only people with decision power on some dimension of the work. Version 2 has a Team Formation stage; it cannot tell you who in your company decides.
- Required:Pace.One-hour sessions, managers in the first mobs, pauses between gates. Version 2 made the gates better. Tired people still approve anything.
- Required:The organization's own knowledge.Internal libraries to inject before code generation, the decision records, where the trusted documentation lives.
- Required:Knowing when not to use it.A one-line change does not need AI-DLC, in any version.
So the lesson for anyone adopting an AI framework in 2026: build judgment, rent plumbing. Spend your engineering time on the parts that encode how your organization decides, and keep the generic machinery thin, because the vendor is about to ship it. That is the argument of don’t adopt AI-DLC, steal from it, and version 2 is its strongest evidence.
Should you move from version 1
If you have a workflow in flight on version 1, finish it, and start the next intent on version 2. Version 2 itself refuses to refresh a project while a workflow is active, which tells you how AWS thinks about changing the ground under work in progress.
Pick Classic for that first run. It reproduces the version 1 ceremony in 18 of the 33 stages, so your team learns the new engine without also learning a new process. Once Classic feels routine, try Feature on a real production change and turn on the walking skeleton.
And before you port your homegrown layers, sort them into the two lists above. The plumbing probably has a native equivalent now; delete it. The judgment does not; move it into rules, team memory, or a plugin that only adds. The setup itself is in AI-DLC with Claude Code, and if your codebase is old and large, read AI-DLC in brownfield first.
FAQ
When was AI-DLC 2 released?
The first stable 2.x release, AI-DLC 2.7.0, was published on September 1, 2026. 2.8.0 followed on September 8, 2.9.0 on September 15, and 2.10.0 on September 24. Versions 2.0 to 2.6 were never published as stable releases.
Is AI-DLC 2 compatible with version 1 workflows?
Not as a drop-in. Version 2 is a different installation, a native aidlc command plus harness runtimes, and it keeps state in per-intent record folders.
Finish in-flight version 1 workflows where they are and start new intents on version 2. The Classic profile reproduces the version 1 ceremony so the process feels familiar.
What is the difference between the Classic and Feature profiles?
Classic runs 18 of the 33 stages: Inception and Construction with one approval per stage, skipping Ideation and Operation, like version 1. Feature runs all 33 stages at standard depth, from intent capture through deployment and feedback.
Does AI-DLC 2 still require Kiro or Amazon Q?
No. Version 2 runs one engine on seven harnesses: Claude Code, Kiro CLI, Kiro IDE, Codex CLI, Cursor, opencode, and GitHub Copilot. The README recommends Claude Opus 4.8 as the model.
What is the AI-DLC learning loop?
Each stage keeps a diary of the calls the agent made. At the approval gate you choose which to keep, and the kept ones become rules in project or team memory. They apply from the start of the next workflow, never in the middle of the current one.
Is AI-DLC 2 open source?
Yes. The workflows are published at github.com/awslabs/aidlc-workflows under the MIT-0 license. You pay only for the model your coding agent uses.
Where to go next
AI-DLC 2 is a real engine now, and it closed most of the gaps that made version 1 hard to run at scale. Use it. But notice what it absorbed and what it left alone, because that line tells you where your own work should go.
The newsletter
Don’t Code, Specify. A weekly dispatch from where AI agents meet real production. No hype, just what shipped and what broke.
Subscribe on Substack (opens in a new tab)