Skip to content
← articles
updated AI-DLCAI AgentsAWSHarness EngineeringSoftware Engineering

AI-DLC 2: What Changed, and What It Made Obsolete

AWS rewrote AI-DLC. Version 2 ships as a deterministic engine that takes routing, grading, and state away from the model: 5 phases, 33 stages, 14 agents, 11 workflow profiles, a learning loop, verified construction checkpoints. What changed from version 1, which homegrown layers it made obsolete, and what survived.

AWS did not bump AI-DLC to version 2. It rewrote it, and in one release most of the plumbing teams had built on top of version 1 stopped being worth maintaining.

AI-DLC 2 is the second generation of AWS’s AI-Driven Development Life Cycle. The first stable 2.x release, 2.7.0, landed on September 1, 2026. Three weeks later 2.10.0 was out. Version 1 was a set of rules files that told an agent how to walk you through a project. Version 2 is an engine: a native aidlc command that installs one deterministic workflow into seven coding agents, with 5 phases, 33 stages, 14 agents, 11 workflow profiles, and a loop that turns your corrections into rules.

If you are new to the method itself, start with what AI-DLC is. This article is for people who already know version 1, or who are deciding whether to start on version 2.

I followed AI-DLC version 1 inside a large engineering organization, through pilot squads and a training program. The teams there did what every serious adopter does with a young framework: they built layers on top of it to cover what it lacked. Then AI-DLC 2 shipped most of those layers as native features. Watching that happen taught me more about adopting AI frameworks than the framework did, so the second half of this article is about that.

From rules files to an engine in fourteen months

How AI-DLC got to version 2

  1. Jul 2025The whitepaper

    Raja SP publishes the AI-DLC method definition and AWS announces it on the DevOps blog. Ten principles, three phases, no code.

  2. Nov 2025Open source as rules files

    AWS releases the workflows as Amazon Q Developer rules and Kiro steering files. The agent reads the rules and follows the lifecycle.

  3. Jan to Apr 2026The 0.1 series

    Nine releases harden the rules-driven workflow.

  4. Jun 19, 2026Version 1.0.0

    The first stable release. 1.0.1 follows on June 30 and becomes the last of the line.

  5. Sep 1, 2026Version 2.7.0

    The first stable 2.x release. The 2.x line had been built separately; no 2.0 to 2.6 release was ever published as stable.

  6. Sep 24, 2026Version 2.10.0

    Stronger construction checkpoints, coexisting harnesses, worktree fixes. Previews of 2.10.1 ship almost daily.

Stable releases from github.com/awslabs/aidlc-workflows. Preview builds omitted.

The version number tells you something on its own. AWS went straight from 1.0.1 to 2.7.0, which means the 2.x line spent months maturing in parallel before anyone was asked to depend on it. When it arrived on the main branch it was already on its seventh minor version.

AI-DLC 2.10.0 in numbers

33
stagesacross 5 phases
14
agents11 experts, 2 reviewers, 1 composer
11
workflow profilesplus a composer for custom routes
108
audit event typesin an append-only log
From the repository README and user guide at version 2.10.0.

The engine decides, the model executes

Of everything in version 2, this is the change I would put on the box. Version 1 was rules files: prose the agent read and was trusted to follow. The agent decided which stage came next, whether a step could be skipped, whether its own output was good enough, and what the state file said. When it drifted, nothing stopped it except a tired human at a gate. That is where slop comes from: a model that routes itself, grades itself, and edits its own records.

Version 2 splits the work in two. A deterministic engine, plain code with a handful of subcommands, owns the routing: which stage runs next, which profile applies, when to stop, what a gate is waiting for. It emits a typed instruction, and the conductor (the /aidlc session, which is the LLM) carries it out and reports back. The model still does the thinking inside a stage. It no longer decides the shape of the process. AWS’s own docs put the split in one sentence: the engine owns the routing, the conductor owns execution quality.

Powers the model had in version 1, and who holds them in version 2

From the 2.10.0 user guide and glossary. Every refusal or stand-aside is written to the audit log.
In version 1 the model couldIn version 2
Decide which stage comes next, or skip oneThe engine routes from a stage graph compiled when the workflow starts. A skipped stage is recorded as skipped.
Say the tests passedA Unit is verified only by the check command a human approved, and only through a receipt the tool writes. A handwritten proof file cannot verify a Unit.
Grade its own artifactSeparate reviewer agents judge it, bound to a fingerprint of the artifact, so changed work cannot inherit an earlier approval. Deterministic sensors (linters, type checks) fire on the output.
Edit the workflow stateA guard blocks the model's tools from writing session and plan-approval records. State moves through the engine.
Approve its own gateA gate put to you needs a real message from you. The human-presence guard cannot be lowered per workflow.
Change the rules halfway throughRules, stages, and sensors are compiled once per workflow. A new rule applies next time.
Lose the links between artifactsAutomated verification gates check traceability between phases before the next one builds on them.
From the 2.10.0 user guide and glossary. Every refusal or stand-aside is written to the audit log.

None of this makes the model smarter. It makes the process less dependent on the model behaving, which is a different and better goal. It is the same lesson I keep landing on in harness engineering: a deterministic check that fails the build beats a paragraph of instructions the agent can ignore. AWS just applied it to the whole lifecycle. Even the packaging follows it: every harness gets the same engine, and the build that produces the seven distributions runs twice and compares the output byte for byte.

The lifecycle grew two phases

Version 1 had the three phases of the whitepaper. Inception turned an intent into requirements, stories, and Units of work. Construction turned each Unit into design, code, and tests. Operations deployed and watched it run. A human approved every step.

AI-DLC version 1

  1. 01

    Inception

    Mob Elaboration: the AI proposes, the team refines.

    • Workspace detection
    • Reverse engineering (brownfield)
    • Requirements analysis
    • User stories
    • Workflow planning
    • Application design
    • Units generation

    Gate: Human approval per stage

  2. 02

    Construction

    Mob Construction, per Unit.

    • Functional design
    • NFR requirements
    • NFR design
    • Infrastructure design
    • Code generation
    • Build and test

    Gate: Human approval per stage

  3. 03

    Operations

    The whitepaper describes deploy, monitoring, and approved runbook actions. The version 1 rules shipped it as a placeholder.

Version 1 as its rules files define it (release 1.0.1). The Operations phase was a placeholder reserved for future expansion.

Version 2 keeps those three and wraps them. Initialization runs first, with no human, and sets up the workspace in under a second. Ideation runs before Inception and asks whether the thing is worth building at all: intent capture, market research, feasibility, scope, team formation, rough mockups, and an approval to proceed. Between phases, a verification gate now checks traceability automatically, so an orphaned requirement or a story that maps to nothing gets caught before the next phase builds on it.

AI-DLC version 2

  1. 0.1 to 0.3

    Initialization

    Workspace, detection, state. Automatic.

    • 3 stages, no gate
  2. 1.1 to 1.7

    Ideation

    Is this worth building?

    • Intent Capture
    • Feasibility
    • Scope Definition
    • plus 4 more

    Gate: Verification gate 1

  3. 2.1 to 2.9

    Inception

    Requirements, design, Units.

    • Reverse Engineering
    • Practices Discovery
    • Requirements Analysis
    • Units Generation
    • plus 5 more

    Gate: Verification gate 2

  4. 3.1 to 3.7

    Construction

    Design, code, tests per Unit.

    • Functional Design
    • NFR Design
    • Code Generation
    • Build and Test
    • plus 3 more

    Gate: Verification gate 3

  5. 4.1 to 4.7

    Operation

    Deploy, observe, feed back.

    • Deployment Pipeline
    • Observability Setup
    • Incident Response
    • plus 4 more
Every stage outside Initialization still ends at a human approval gate. The full stage list is in the pillar guide.

The new Inception stage worth noticing is Practices Discovery. Before writing requirements, the agent looks at how your team already works (testing posture, deployment habits, code style), drafts it, has three other agents inspect the draft independently, interviews you about the gaps, and writes the result into the team’s memory. Version 1 learned your conventions by being corrected. Version 2 asks first.

Fourteen agents instead of one voice

Version 1 was one agent following steering rules and playing roles as the rules told it to. Version 2 ships 14 named agents: 11 domain experts (product, design, delivery, architect, AWS platform, compliance, DevSecOps, developer, quality, pipeline and deploy, operations), two reviewers, and a composer.

AWS calls the philosophy “Small Mob, Broad Agents”. The alternative was obvious and wrong: thirty narrow specialists, each owning one artifact, passing work down a chain. That is waterfall with extra steps, and every handoff loses context. Eleven broad agents that each work across several stages carry more of the picture with them.

Most stages still run inline, in your conversation. Four dispatch work: Reverse Engineering as a two-step pipeline (a developer agent scans, an architect agent writes the synthesis), Practices Discovery and Code Generation as subagents, and User Stories as a mob where design, developer, and quality agents write in parallel. The count in 2.10.0 is 29 inline, 2 subagent, 1 pipeline, 1 mob.

The two reviewers matter more than the headcount. After a stage produces its artifact, the product-lead reviewer judges requirements and stories, and the architecture reviewer judges technical design. An adversarial review can send the work back up to two times. And then it stops and hands you the findings, because the reviewer never blocks. The human decides. That rule is the whole method in one line, and I am glad AWS wrote it into the engine instead of leaving it to good intentions.

Profiles replace the one-size lifecycle

The whitepaper’s tenth principle said no workflow should be hard-wired: the AI proposes the depth, the human adjusts it. Version 1 implemented that as a judgment call inside the rules. Version 2 turned it into 11 named workflow profiles (the engine calls them scopes), each a fixed route through the 33 stages with a default depth and test strategy.

The profiles you will actually use first

The other seven are Enterprise, MVP, Proof of concept, Refactor, Infrastructure, Security patch, and Workshop. /aidlc compose builds a custom route and waits for your approval.
ProfileStagesWhat it is for
Classic18 / 33The version 1 ceremony: Inception and Construction, one approval per stage. The engine's default when you name nothing.
Express10 / 33Requirements already understood. Shortest path to code and tests.
Feature33 / 33A production feature through the whole lifecycle at standard depth.
Bugfix9 / 33A known defect, a focused repair, a regression test, and the deploy.
The other seven are Enterprise, MVP, Proof of concept, Refactor, Infrastructure, Security patch, and Workshop. /aidlc compose builds a custom route and waits for your approval.

Classic is the migration path, and AWS made it the default on purpose. A team that knows version 1 can start on version 2 without learning a new ceremony. Depth and test strategy are now separate dials (Minimal, Standard, Comprehensive), and you can turn either one without switching profiles.

Corrections become rules: the learning loop

This is the feature I would have paid for. Version 1 forgot. You corrected the agent on Monday, and on Thursday, in a new workflow, it made the same mistake, because nothing carried the correction forward except your own memory.

Version 2 keeps a diary for every stage, a file called memory.md, with four headings: Interpretations, Deviations, Tradeoffs, and Open questions. When the agent makes a call the stage instructions did not cover, it writes it down. At the approval gate, the engine shows you those lines verbatim and asks if you want to keep any, plus a free-text “Anything to add for next time?”.

How a correction becomes a rule

Input

You correct the agent during a stage

  1. DIARYThe agent records the call it made

    The stage's memory.md gets an entry under Interpretations, Deviations, Tradeoffs, or Open questions.

  2. GATEThe gate shows you the candidates

    Lines are surfaced as written, no paraphrase. You tick the ones worth keeping and can add your own.

  3. CHECKA conflict check against org rules

    If the new line contradicts an organization rule, you revise it, skip it, or escalate it.

  4. WRITEThe kept line lands in project memory

    It goes to project.md, one click promotes it to team.md. Open questions never become rules.

Output

The rule applies from the first stage of the next workflow. Never in the middle of the current one.

How a correction becomes a rule: flow of 4 steps from “You correct the agent during a stage” resulting in “The rule applies from the first stage of the next workflow. Never in the middle of the current one.”.

The docs walk through a real example. On a banking project, a stakeholder note said “the transaction shouldn’t duplicate on retry”. The product agent read “transaction” as a database transaction and wrote an ACID requirement. The person meant a payment. They corrected it, the diary caught the interpretation, and the kept rule fixed the term for every workflow after that.

The last detail is the one that shows someone thought about it. A learning never changes the rules mid-run. The engine compiles stages, rules, and checks once, when a workflow starts, and keeps that compiled view until the end. The gates you approved earlier attested to a stable rule set, and the ground does not move under a workflow in progress.

Rules and sensors: AWS now talks about harness engineering

Version 2 splits steering into two halves. Rules are prose instructions loaded before work (feedforward). Sensors are deterministic checks fired on the output, like a linter or a type check (feedback). AWS calls the pair the control loop.

Rules resolve through five layers, and the model is strictly additive: every applicable layer is in context at once, and nothing silently overrides anything.

The five layers of rules

  1. ORG

    Organization

    Framework and company defaults: trunk-based development, testing posture, walking skeleton policy.

  2. TEAM

    Team

    Practices your team affirmed, including the ones Practices Discovery found.

  3. PROJ

    Project

    This project's specialization. Where the learning loop writes by default.

  4. PHASE

    Phase

    Rules for every stage in one phase, like documenting two alternatives for each architecture decision in Inception.

  5. STAGE

    Stage

    Reserved for a future release.

All of it is plain Markdown under aidlc/spaces/<space>/memory/, readable and editable by hand.

The repository even ships a Harness Engineer Guide for reshaping stages, agents, rules, sensors, and knowledge without touching the engine code. I have been arguing that the harness is where your edge lives for months. It was good to see AWS put the word in its own documentation.

Construction got verified checkpoints

Version 1 asked you to approve every Construction stage for every Unit. On a feature with six Units and five stages each, that is thirty approvals, most of them on design documents nobody would read twice. Approval fatigue was one of the first things to break in the pilots I followed.

Version 2 changed the default for new solo work with Units. It builds one Unit at a time, in dependency order, and asks you to approve each completed Unit at a verified checkpoint instead of approving every intermediate document. Verified means something specific: during Delivery Planning, the agent proposes a real project check (bun test, pytest, make check), you approve that exact command, and the engine reuses it at every checkpoint. Changing it later requires your approval again. A handwritten proof file cannot verify a Unit.

What Construction asks you in version 2

  • Required:
    Approve the plan for each Unit before code is generated.Plan Approval stays human in every mode.
  • Required:
    Approve the verification command, once.Reused at every checkpoint. Changing it needs a new human approval.
  • Required:
    Approve the walking skeleton, when it is on.The first Unit is the smallest working end-to-end slice, verified by the command before the rest begin.
  • Optional:
    Choose: continue automatically, or review each checkpoint.Automatic skips routine completion questions. It never skips plan approval, the command, or failures.
  • Required:
    Decide on every failure.A failed Unit halts Construction and offers retry, skip, or abort, even in automatic mode.
Parallel builds exist too. You have to ask for them: stage-major order plus swarm execution, each Unit in its own git worktree.

There is also a team mode. With Unit ownership set to team, people claim Units and build them in their own checkouts at the same time, with their own gate rhythm. Version 1 had nothing like it. Teams that wanted parallel construction had to invent it.

Guards, and who is allowed to lower them

Version 2 adds a Guard Policy per piece of work, with three values: strict, relaxed, and off. It controls five fences: plan approval, review freeze, state transition, reviewer read scope, and human presence. Enterprise defaults to strict. The other profiles default to off, which lets four of the fences stand aside for undirected work, and every time a fence stands aside the audit log records it.

The fifth fence is the interesting one. Human presence cannot be lowered for a single workflow. A gate put to you needs a real message from you. The only way around it is a machine-wide environment variable, which is the right level of friction for “let the agent approve its own work”.

Multi-repo is native

Version 1 was single-repository. Real features rarely are: a back end, a front end, maybe a mobile app, each in its own repo. Teams I watched worked around it by running one session per stack and pointing the agent at the other repo’s directory as reference material, with the contract between them agreed by hand beforehand.

Version 2 supports multi-repo intents. You declare the repository set when the intent is created (or let the engine discover sibling repositories), Construction anchors each git operation to a specific repo, and Reverse Engineering has to complete one full scan chain per repository before the stage can be approved. The workaround became a flag.

The record is now something you can query

Every intent gets its own record folder under aidlc/spaces/<space>/intents/. Inside: the state file, with a six-state checkbox per stage (not started, in progress, awaiting approval, revising, completed, skipped), every artifact, every questions file, and an append-only audit log with 108 event types. Version 1 had a state file and an audit file too. Version 2 made them the engine’s source of truth, and added a runtime graph compiled from the audit log after every transition.

That makes the workflow measurable without new tooling. /aidlc-session-cost prints duration, stage outcomes, and learnings. /aidlc-replay narrates the session for people who were not there. aidlc attest tells you whether a commit’s content still matches what a reviewer approved. I wrote about using that data in after story points.

What AI-DLC 2 made obsolete

Now the part the release notes do not say.

The team I followed most closely maintained an overlay on top of version 1: a set of extensions, commands, and patches that made the method fit their organization. It was good work, built by people who ran the method every day and fixed what hurt. Then AI-DLC 2 arrived and shipped most of it natively.

What teams built on version 1, and what version 2 ships

Not a criticism of the overlay. It covered real gaps. That is exactly why the vendor closed them.
Homegrown layer on version 1Native in version 2
A shared state file every custom skill read and updated, so a new session could resume after clearing contextA state machine with six-state checkboxes, a recovery breadcrumb before compaction, and /aidlc --resume
Patches injected into the core rules without forking themA plugin system that only adds, never edits the core, with validate, build, and sync commands
An automatic retrospective after Build and Test that scanned for where the AI failedThe learning loop: per-stage diaries, gate candidates, rules that apply next workflow
A plan to parallelize construction with git worktreesSwarm execution with one worktree per Unit, and team mode with Unit claims
Running one session per stack and pointing at the other repo by handMulti-repo intents with per-repo reverse engineering
Archiving each finished task's artifacts to start the next one cleanPer-intent record folders, spaces, and intent archive commands
Not a criticism of the overlay. It covered real gaps. That is exactly why the vendor closed them.

This is what happens to every layer of plumbing you build on a fast-moving framework. If the gap is generic, the vendor sees the same gap in every customer and ships it. Your clever state handoff becomes their feature, and your code becomes migration work.

What survived the rewrite

Here is the list that did not get absorbed, and it is the list I would protect.

What version 2 did not replace

  • Required:
    The rules against vibe coding.Never edit generated code by hand. Ask for impact before a big change. Get the critique in a fresh context. They came from watching the agent fail, and no engine encodes them for you.
  • Required:
    The ritual before the code.How an intent gets shaped before AI-DLC sees it: the entry document, the curated sources, the complexity triage. Ideation helps. It still assumes someone knows the problem.
  • Required:
    Who sits in the mob.Only people with decision power on some dimension of the work. Version 2 has a Team Formation stage; it cannot tell you who in your company decides.
  • Required:
    Pace.One-hour sessions, managers in the first mobs, pauses between gates. Version 2 made the gates better. Tired people still approve anything.
  • Required:
    The organization's own knowledge.Internal libraries to inject before code generation, the decision records, where the trusted documentation lives.
  • Required:
    Knowing when not to use it.A one-line change does not need AI-DLC, in any version.
All of it is judgment. None of it is plumbing.

So the lesson for anyone adopting an AI framework in 2026: build judgment, rent plumbing. Spend your engineering time on the parts that encode how your organization decides, and keep the generic machinery thin, because the vendor is about to ship it. That is the argument of don’t adopt AI-DLC, steal from it, and version 2 is its strongest evidence.

Should you move from version 1

If you have a workflow in flight on version 1, finish it, and start the next intent on version 2. Version 2 itself refuses to refresh a project while a workflow is active, which tells you how AWS thinks about changing the ground under work in progress.

Pick Classic for that first run. It reproduces the version 1 ceremony in 18 of the 33 stages, so your team learns the new engine without also learning a new process. Once Classic feels routine, try Feature on a real production change and turn on the walking skeleton.

And before you port your homegrown layers, sort them into the two lists above. The plumbing probably has a native equivalent now; delete it. The judgment does not; move it into rules, team memory, or a plugin that only adds. The setup itself is in AI-DLC with Claude Code, and if your codebase is old and large, read AI-DLC in brownfield first.

FAQ

When was AI-DLC 2 released?

The first stable 2.x release, AI-DLC 2.7.0, was published on September 1, 2026. 2.8.0 followed on September 8, 2.9.0 on September 15, and 2.10.0 on September 24. Versions 2.0 to 2.6 were never published as stable releases.

Is AI-DLC 2 compatible with version 1 workflows?

Not as a drop-in. Version 2 is a different installation, a native aidlc command plus harness runtimes, and it keeps state in per-intent record folders.

Finish in-flight version 1 workflows where they are and start new intents on version 2. The Classic profile reproduces the version 1 ceremony so the process feels familiar.

What is the difference between the Classic and Feature profiles?

Classic runs 18 of the 33 stages: Inception and Construction with one approval per stage, skipping Ideation and Operation, like version 1. Feature runs all 33 stages at standard depth, from intent capture through deployment and feedback.

Does AI-DLC 2 still require Kiro or Amazon Q?

No. Version 2 runs one engine on seven harnesses: Claude Code, Kiro CLI, Kiro IDE, Codex CLI, Cursor, opencode, and GitHub Copilot. The README recommends Claude Opus 4.8 as the model.

What is the AI-DLC learning loop?

Each stage keeps a diary of the calls the agent made. At the approval gate you choose which to keep, and the kept ones become rules in project or team memory. They apply from the start of the next workflow, never in the middle of the current one.

Is AI-DLC 2 open source?

Yes. The workflows are published at github.com/awslabs/aidlc-workflows under the MIT-0 license. You pay only for the model your coding agent uses.

Where to go next

AI-DLC 2 is a real engine now, and it closed most of the gaps that made version 1 hard to run at scale. Use it. But notice what it absorbed and what it left alone, because that line tells you where your own work should go.