Skip to content
← articles
Agentic SDLCAI-DLCAI AgentsEngineering LeadershipSoftware Engineering

The Agentic SDLC Playbook: The Build Shrank. Now Fix Both Sides.

For decades the software lifecycle was built around one slow, expensive step. Agents made it cheap, and everything that leaned on it fell over. A field playbook for the agentic SDLC, written in the order you adopt it: the repository, the checks, the context, the spec, the review, the flow, and the loop.

For decades the software lifecycle was built around one slow, expensive step. Agents made that step cheap. Everything that leaned on it fell over.

The build shrank. What used to take a team weeks of writing code now takes an agent hours. I spend my working days with engineering teams going through this, and the first thing everyone expects turns out to be wrong. Nothing ships ten times faster. The build got cheap, and both sides of it broke.

This playbook is about the two sides. It is written in the order you adopt it, so you can read it top to bottom and get the sequence for free: what to fix first, what each step looks like on disk, how to tell where you are, and how high to climb. It works with any harness.

The build shrank. The calendar barely noticed

Keep one question in your pocket for the whole article: if the agent writes the code ten times faster, why doesn’t the feature ship ten times faster?

Here is the measurement that made me take the question seriously. One feature from a team I work with: ten tasks, thirteen merge requests, six repositories, five languages, with tests and documentation. Before agents, the team estimated it at 10 to 15 working days for five developers full time. With agents it took three working days and six and a half hours of human work, summed across everyone, plus $67 in tokens.

One feature, before and after agents

Human effort÷ 60 or more
Without agents (team estimate): 5 developers full time for 10 to 15 working days
With agents (measured): 6.5 hours, summed across everyone
Calendar time÷ 4
Without agents (team estimate): 10 to 15 working days
With agents (measured): 3 working days
One feature, measured once and after the fact, on work chosen because it suited the agent. Bars use the low end of the team's estimate.

Human effort fell by a factor of about 60. The calendar fell by about 4. Same feature, same team. It is one measurement, taken retroactively, on a feature picked because it fit the agent, so read it as an order of magnitude and nothing finer. But the shape is the whole story. The work stopped being what took the time.

Picture a factory line built around one slow press. Every station is tuned to its pace. The people feeding it have time to prepare the next batch. The people after it inspect a part every few minutes. The operator stands at the press while it runs and thinks about the next job. Now swap in a press ten times faster. The feeders cannot keep up, the inspectors drown, and the operator, who used to rest while the machine worked, now stands at it all day. Output barely moves. The line was never limited by the press alone. It was tuned around it.

That is a software team after agents.

The slow build was doing three jobs nobody wrote down

Before agents, the build was the expensive step, so everything around it existed to protect it. Requirements, refinement, estimates, sprints, review: all of it was arranged so that the weeks of people writing code would go as smoothly as possible. The bottleneck was the build, and the build was paying for three things nobody ever put in a process document.

It gave the upstream time to prepare. While engineers spent two weeks building, product had two weeks to work out the next thing. A vague requirement got clarified in a hallway conversation halfway through the sprint, before anyone needed it.

It gave the downstream a trickle it could absorb. Review, QA and deploy were sized for the amount of code people produce. A reviewer could read every line because a person had typed every line.

It gave the developer time to think. Typing was thinking time. While you wrote the code and sat in the meetings, you were assembling the context in your head. Experienced developers in one rollout told me they now end the day more tired than before they used AI. The context arrives ready-made, it is a lot to filter, and the slow part where they used to think is gone.

When the build flattened, all three jobs disappeared at once. Both sides broke.

What the slow build gave you for free

The rest of this playbook is how to build the right-hand column, in order.
The slow build gave youWhen it shrankWhat you now build on purpose
Time for the upstream to prepareSpecs arrive vague or late, and the agent fills every gap with a confident guessWritten context before the ticket: the problem, the scope, the acceptance criteria
A trickle the downstream could absorbReview queues, merge requests too big to read, approvals that mean nothingSmall units, deterministic checks, and review of intent and risk
Thinking time for the developerContext lands ready-made, people end the day drained, and what they learned dies with the chat windowSpec time before execution, sessions with a limit, and work logs that keep what was learned
The rest of this playbook is how to build the right-hand column, in order.

So the bottleneck moved. It used to be the build. Now it sits on both sides: context on the left, verification on the right. On the left, the question is whether the agent knows what to build and under which rules. On the right, the question is whether anyone can tell that what came back is correct. Every step below attacks one of the two.

What an agentic SDLC actually is

An agentic SDLC is a software lifecycle where each stage leaves a versioned artifact the next stage can read, whether the next reader is a person or an agent. The ticket feeds the plan. The plan feeds the code. The checks judge the code. What the cycle taught you goes back into the repository, so the next cycle starts smarter.

Three things are different from the lifecycle you know. Humans stop reviewing lines and start reviewing intent and evidence. Rules that a machine can check are checked by a machine, every time. And the repository becomes the place where everybody meets: product, design and engineering each write into it in their own language, and the agent reads all of it.

The agentic SDLC as a chain of artifacts

  1. 01

    Upstream

    What to build, and why.

    • A written problem behind the epic
    • A ticket with acceptance criteria
    • Scope seen by engineering

    Gate: A person accepts the spec

  2. 02

    Build

    The agent writes; the repository constrains.

    • AGENTS.md at the root
    • A plan with units sized for review
    • Code and tests

    Gate: The one command passes

  3. 03

    Verify

    Machines check what nobody reads anymore.

    • Lint and type rules
    • Tests that block the merge
    • Review of intent and risk

    Gate: A named owner approves

  4. 04

    Deliver

    Ship small, watch, undo fast.

    • Canary with an abort
    • A rollback that has been run
    • A trace of what the agent did

    Gate: Release authorized

  5. 05

    Learn

    What the cycle taught goes back in.

    • Work logs
    • Repeated review findings become rules
    • Incidents become eval cases

    Gate: Someone curates before it becomes truth

The last phase is the one teams skip, and it is the one that makes the next cycle cheaper.

The seven steps that follow build this chain, in the order that holds up in practice.

1. Make the repository able to receive work

Start with the build, not with the spec. This is the decision that matters most, and the natural instinct runs the other way.

Some groups next to mine started upstream. One of them automated the path from product requirements to technical spec first and planned to deal with the rest later. I do not believe in that order. A perfect spec handed to a repository with no context file, no single command that proves the work is done, and no tests that block a merge still produces bad code. The agent guesses. The developer babysits it, prompting and pasting. The team concludes the tool does not work, and the beautiful spec gets the blame.

The way I put it to one engineering leader: if the build is not ready when the upstream speeds up, you get an avalanche of stressed, annoyed developers who no longer write code. They babysit it.

There is a second reason, and it is about trust. When the build is hardened, people start trusting what the agent produces, because they know how hard a merge request got hammered before it came out. You cannot get that trust from a better prompt.

The repository needs two things before anything else. A context file the agent reads every time, and one command that tells it the work is done.

# AGENTS.md

## Done means

Run `make check` and paste the output. It runs lint, types, unit and
integration tests, and exits non-zero on any failure. Fix the code, not
the test.

## Conventions the linter does not cover

- Money is integer cents, never floats.
- Every new endpoint gets an integration test in tests/integration/.

## Where the why lives

- Specs and decisions: docs/specs/ and docs/adr/
- Work logs: docs/worklogs/ (unofficial notes; verify before trusting)

## Mistakes we have seen twice

- The billing client retries on 409. Stub it in tests.

Keep AGENTS.md as the single source and point every tool-specific file at it. For Claude Code, a CLAUDE.md with the single line @AGENTS.md imports it. Teams switch harnesses. Your conventions should not have to move with them.

2. Turn every rule you can check into a failing check

A line in AGENTS.md is a request. The model reads it, and depending on how much else is piled into its context, it may skip it. A model is a statistical machine. Deterministic is everything that is not statistical, and that is where your rules belong whenever they can live there.

So when a review finding shows up twice, it should not become one more line of Markdown. If a machine can check it, it becomes a lint rule, a type, or a test, and the error message points to where the fix is written.

// eslint.config.js
export default [
  {
    rules: {
      "no-restricted-imports": [
        "error",
        {
          paths: [
            {
              name: "moment",
              message: "Use date-fns. Why and how: docs/conventions.md#dates",
            },
          ],
        },
      ],
    },
  },
];

That rule used to be a sentence the agent followed most of the time. Now it holds on every run, and the agent only needs to know one thing: run the check, read the error, and the error tells it where the answer is. You take the job of judging away from the model, and the model gets better at the job you left it.

A healthy context file shrinks over time. Every rule that can be checked moves down a layer, into the linter, the types or the tests, and what stays in Markdown is judgment. The pipeline enforces the rest: red tests and exposed secrets do not merge, no matter who or what opened the merge request.

3. Stop throwing context away

Every time someone closes a chat window, the organization loses what that session learned. The rule the developer had to explain by hand, the dead end they found, the reason they split the migration: all of it was in the session, and the session is gone. That was the token, the time, and the thinking of the person who did it.

Context is the new bottleneck, and it is expensive in three ways. It has to be captured, because most of it lives in people’s heads. It has to be curated, because what gets written down is often wrong or contradictory. And it has to be connected, because a great knowledge base that one team keeps to itself does not compound.

Capture starts with a work log. At the end of a unit of work, before you clear the session, ask the agent to extract what it needed and did not find, and write it to a Markdown file.

# Work log: invoice retry, 2026-09-18

## Context I had to add by hand

- Invoices to public-sector customers must not retry on weekends.
  The rule is not written anywhere; finance confirmed it in chat.
- The email provider returns 202 before validating the address.

## Decisions taken in the session

- Moved the migration to its own merge request to keep review small.

## Promote?

- [ ] Weekend rule to docs/specs/invoicing.md (ask finance to sign off)
- [ ] Provider behavior to AGENTS.md, under mistakes seen twice

A work log is not documentation. It is raw material. Curation decides what gets promoted to the spec, the context file or a check, and it is patient, unglamorous work: label what you know as green, yellow or red by how much you trust it, and never let a red note become the truth an agent builds on.

The stakes are higher than they look. In one rollout, half of a business rule lived in a product manager’s head and the other half in an engineering manager’s head, and it was written nowhere. On older systems I keep hearing the same complaint: nobody dares change a rule, because one day someone had a reason for it, and nobody knows where that reason is today. Technical debt is often exactly this, a decision whose rationale was never written down. The wiki, when it exists, tells maybe a tenth of how the story actually happened.

Each cycle that feeds the repository makes the next one cheaper. That is the compounding part of an agentic SDLC, and it is the part a slow build never forced anyone to build.

4. Move the spec before the execution

I sat with a developer who was doing what most developers do in the first months with agents. They read the product requirement, read the ticket, and started prompting, filling in the details that were missing as the agent asked or failed. That is specifying at execution time. It feels like productivity. It is the clearest sign of a team stuck between level 2 and level 3.

The agent changed the economics of a vague ticket. A developer with a vague task gets stuck, walks over and asks, and that question was free spec review that nobody ever counted. The agent never gets stuck. It fills the gap with an assumption and delivers so fast that nobody sees the assumption until it sits inside a diff too big to read. The model will satisfy any spec, good or bad.

This does not require a new format. A ticket with a problem, a scope and acceptance criteria is already the artifact the agent needs to read. What changes is the reader: you now write it for the machine too.

# Show why an invoice failed to send

## Problem

Customers open support tickets to ask why an invoice never arrived.
Written problem: docs/problems/invoice-visibility.md

## Scope

- Billing page, read-only. Out of scope: resending.

## Acceptance criteria

- An invoice that failed shows the reason and the last attempt time.
- An invoice that was delivered shows no badge.
- `make check` passes, with one integration test per criterion.

## Context

- Failure reasons: docs/specs/invoicing.md#failures
- Design: link to the approved mock

Two things have to be true before a ticket like this exists. The epic points to a written problem: what it is, who it affects, and how you will know it worked. And engineering saw the scope before the ticket was written, because that is where the cheap questions get asked.

The flow also runs backward. If you do not take what the build learned and feed it to the next requirement, you push the problem to the left, and the next product requirement comes out broken. The work logs from step 3 are how the build talks back to the upstream.

5. Change what arrives at review

Back to the question from the top. The time the agent saved did not vanish. Most of it moved into review. Developers on teams that adopt agents complain about it first: the number of merge requests they have to read keeps growing, and the queue is now the slowest step.

Reviewing faster does not fix it. A mandatory approval per merge request, promised within a day, was sized for the amount of code people produce. Multiply the merge requests and either the promise breaks or the approval stops meaning anything. Reading every line made sense when a person wrote every line. It does not keep up with an agent.

Fix what arrives instead. Cap the size of a unit of work when the plan is approved, before the code exists, so a reviewer can read the result in one sitting: split by endpoint, give the database migration its own merge request. Then change what the reviewer looks at. Step 2 already blocks the mechanical problems, so review can be about intent and risk: does this do what the ticket meant, and what breaks if it is wrong. Prefix every comment with its weight, blocking, major or minor, so the author knows which ones stop the merge. I wrote up how one team climbed out of the review jam in the AI-DLC J-curve.

The reviewer is the scarcest part of the system, and the most fragile. In one rollout I followed, people approved documents after long sessions with no capacity left to read them, and about an hour per session turned out to be the practical limit. The opposite failure is quieter. The tenth artifact the agent produced gets approved with less care than the first, because the first nine looked fine, and the reviewer just presses the button. A gate passed by a tired or trusting person is a rubber stamp. Approval gates covers the four ways gates fail.

6. Put the human gate where the work is

There are two kinds of human gate. A review gate comes after the work: it is done, a person validates it before the merge. A start gate comes before: nothing begins because the one person who could review it is busy. The second kind is circular. Nothing starts because there is no reviewer, and there is nothing to review because nothing started.

I sat in a meeting about a team that had the product requirement written, a self-contained piece of work, and a repository with its contracts mapped, and nothing was moving, because the engineer who would review it was allocated elsewhere. Someone in the room said they had built the car and were still walking.

The agent works asynchronously. Let the human gate travel with the work instead of standing in front of it. One team I work with drew its board like this.

A board where the human shows up twice

Input

A written problem with scope

  1. humanSpec

    A developer validates the ticket and its acceptance criteria before any agent runs.

  2. queueReady for agent

    Anyone can drag it here. The agent reads the column and the ticket.

  3. agentIn progress

    The agent plans, implements and runs the one command until it passes.

  4. humanReview

    A named owner reviews intent and risk. Comments on the ticket go back to the agent.

  5. learnDone

    Merged, and the work log is filed for curation.

Output

Agent work and human work run in parallel instead of in sequence

A board where the human shows up twice: flow of 5 steps from “A written problem with scope” resulting in “Agent work and human work run in parallel instead of in sequence”.

The sprint carries the same hidden assumption as the start gate. Two weeks made sense when delivery was limited by how fast people type: you batch the work into a coherent package. When code takes hours, waiting two weeks to start something new costs more than doing it. Shortening the sprint does not fix that, because its length was never the problem. Dropping the agent into the ceremonies you already have does not fix it either. That automates the factory instead of redesigning it, and automating a flow that does not work yet only speeds up the bottleneck.

7. Earn the loop

The last stretch is delivery and the closed loop, and it comes last for a reason: every step before it is what makes it safe.

Delivery has to be boring before an agent goes near it. Anyone on the team brings up the same environment with one command. Every change goes out behind a canary with an abort. And rollback has actually been run, recently, by someone on the team. Without those, every autonomous change is all or nothing in production.

Then the agent gets graded like an employee who never stops learning: by a suite. Collect twenty or more real tasks from your own history, each with the result you accepted, and run them in CI whenever a prompt, a skill or the model changes. Every incident becomes a new case. Without evals, “autonomous” means “unverified”, and the project stays at level 3 no matter how good the model is.

Run the agent in a sandbox with scoped credentials, keep a trace of the commands it ran, and give each team a visible token budget. Without a sandbox, a mistake becomes an incident. Without a trace and a budget, failures are silent and the bill is a surprise.

Only then close the loop. A deterministic script watches production. When a metric drifts outside its normal range, it calls the agent, the agent diagnoses and writes the next ticket, and a person triages it. The detection stays deterministic and the model does the part that needs a model, which is the same split that runs through the whole playbook. It is also the same split that makes AI-DLC 2 work: the engine routes, the model executes.

The ladder has no shortcuts

Those seven steps map onto a ladder of autonomy. It defines each level by what the agent takes on, what stays with the human, and who does the verification. Mine is adapted from an internal framework, and the levels will look familiar if you have seen any of the public ones.

The agentic ladder

  1. L1

    Assisted

    "The agent drafts and completes. The developer writes and decides everything."

    Verified by human eyes

  2. L2

    Supervised

    "The agent changes files on request. A person reviews every diff."

    Verified by human eyes, backed by tests

  3. L3

    Pull request

    "The agent takes a whole task, edits many files, opens the pull request. A person reviews the pull request."

    Verified by layered tests and CI

  4. L4

    Autonomous

    "The agent plans, implements, grades itself and delivers behind an approval gate. A person approves at the gate."

    Verified by an eval gate

  5. L5

    Full agentic

    "From spec to production, re-entering the loop on its own. The human defines intent and is on call."

    Verified by autonomous evals

The agentic ladder: 5 progressive levels, from the most basic (Assisted) to the most advanced (Full agentic).

The jump that matters is L2 to L4. That is where the work moves from operating to judging, from moving hands to stating intent. At L2 the agent is a better autocomplete. At L3 you hand it a task the way you would hand one to a colleague. A better model does not produce that jump.

You cannot skip rungs. A team at L2 that forces its way to L5 without layered tests, without context in the repository and without writing specs has close to a hundred percent chance of something going badly wrong, because what it is really doing is vibe coding at scale. A level only holds when automation is reliable enough to take over the supervision the human stopped doing. Climbing without that is the absence of control.

Watch for the plateau, too. Teams that are still prompting are between L2 and L3, and the risk is that they get used to it. Prompting feels like progress. It is the step before the step.

Two more things about the ladder. A score describes readiness, not practice: a repository at L3 is ready to receive agentic work, which does not mean the team is doing it. And the ladder runs on three tracks. Repositories are infrastructure, development teams are the downstream, business teams are the upstream, and each of them has its own L1 to L5.

Your ceiling is your weakest line

To find where a project is, score it. These are the eleven lines I ask teams to score, each one tied to the level it unlocks. Every line gets 0 to 3: 0 absent, 1 ad hoc, 2 established, 3 automated and required. Do it with the whole team, in a retro, and repeat it every cycle. Most of it can be read straight from the repository; I hand teams a skill that reads the repo and proposes the score, and the team corrects it.

Eleven lines to score, 0 to 3

Score the project, not the team: a service on the critical path and a throwaway experiment are two projects with two scores.
Line (level it unlocks)What a 2 looks like
AGENTS.md at the repo root (L2)In the main repos, with commands and conventions, touched this cycle, pointing to the spec when it lives elsewhere
One command that runs everything (L2)One place the agent reads declares everything that must pass before merge, and every check exits non-zero on failure
Acceptance criteria on the ticket (L3)The ticket carries a verifiable criterion before anyone starts the agent
Automated tests (L3)Unit, integration and end-to-end, with the pipeline blocking merge on a red test or an exposed secret
Versioned agent setup (L3)Skills, hooks and the point where the agent stops for approval live in the repo, and run the same in any harness
Deploy coupling (L3)You can merge and deploy one part without shipping another, and the contract between them is versioned
Review discipline (L3)Reviews use severity prefixes and look at intent and risk, not formatting
Environment and delivery (L4)Anyone brings up the same environment, the canary has an abort, and rollback has actually been run
Eval harness (L4)A suite of 20+ real tasks runs in CI with versioned expected results
Sandbox and trace (L4)The agent runs sandboxed, with secret scanning, cost visible per team, and a record of the commands it ran
Conventions in the linter (L4)The conventions that show up most in review are lint rules, and the error message says where the fix is written
Score the project, not the team: a service on the critical path and a throwaway experiment are two projects with two scores.

Then apply the one rule that changes the conversation.

A barrel made of wooden staves holds water up to its shortest stave. Make the other staves taller and you have built a taller barrel that holds exactly the same water.

On an illustrative project, three services and seven people, shaped like the ones I see: tests score 3, with unit, integration and end-to-end blocking the merge. Delivery scores 3, with a canary by default and a rollback run in the last incident. The team sees itself at L3. Its ceiling is L1, decided by one line: AGENTS.md exists in one of the three repositories, last touched five months ago. Without context in the repository the agent works by assumption, and good coverage does not compensate.

Perception follows the best tool. The ceiling follows the one that is stuck. The flashiest numbers on that scorecard are the zeros in the L4 gate, three levels away. The effective task is the cheapest one: a context file in three repositories, which on its own moves the project from L1 to L2.

Climb only as high as your cost of error

The same scorecard applies to everyone. What changes is how high it is worth climbing, and the cost of an error decides that. Every project is a balance between context on one side and speed, safety and reliability on the other.

Target level by cost of error

What keeps an ocean liner safe would sink a jet ski.
ProfileWhat it isAn error costsTarget
Core platformCritical path, deep, many consumersA production incidentL5
Base productCustomer- and partner-facing, solid, less deep than a platformA recoverable regressionL3
ExperimentOne or two people, high uncertainty, a disposable hypothesisThrowing it away and rebuildingL2
What keeps an ocean liner safe would sink a jet ski.

An experiment’s value is speed, so it can afford more mistakes than a platform can. Size does not make a project an experiment, though. Blast radius does. Two people working on money, personal data or the checkout path are not an experiment, however small the team.

Cross the ceiling with the target. Below target, the gap is your roadmap. At target, stop spending on harness there and put the effort into delivery. Above target, you built what you did not need, which usually means the profile was wrong. Invest in the weakest line below the target, and only in that one. An experiment stuck at L1 does not need an eval harness. It needs a context file.

The repository becomes the meeting room

The biggest change I have watched is not in the code. When generating code costs almost nothing, the discipline you came from stops being what organizes the work. The common delivery is always a repository, and everyone now writes into it in their own language: product writes the problem, design writes the mock and the rules behind it, engineering writes the architecture and the checks. The agent reads all of it. Organizations built as boxes of specialists will have to take on multidisciplinary work, whether they plan for it or not.

The engineer’s job moves up a level of abstraction, the same way it did when we stopped managing registers by hand. Less typing, more providing context, designing architecture, integrating systems and deciding what is worth building. Half jokingly, I tell developers the future of programming is writing Markdown files. It is only half a joke.

That move has a cost, and someone has to manage it. The thinking time the slow build used to give developers has to be scheduled on purpose now, as spec time before execution. Sessions need a limit. What worked in the rollout I followed was a manager who sat in on the first sessions, calibrated the pace and stopped approvals from becoming automatic. Without someone close guarding it, every gate turns into a formality.

And not every task should go to the agent. When the bottleneck is knowledge already in the head of the person doing the work, there is nothing for the agent to unlock, and describing the solution costs more than writing it. A randomized trial by METR found experienced developers working on code they knew well were slower with AI, while believing they were faster. Context and verification are the investment. Expecting the next model to fix it is not.

Measure what shrank and what didn’t

The cost of AI measures itself. Your token bill arrives every month without anyone asking. The return does not. When one side of the ledger fills itself and the other needs someone to fill it, every budget conversation happens with half the numbers.

Measure the cost of delivering a feature against what the team would have spent without the agent. That is the measurement at the top of this playbook, and here is where each number comes from.

Reproducing the feature measurement

The numberWho answersWhere it comes from
Size of the feature, on the team's own scaleTech lead, manager or the team in a retro. Never the agentYour tracker
People and time the team would need without the agentDerived from history; the team confirms or correctsEpics closed before the agent arrived
Calendar time to mergeNobodyTracker and git host
Human hours, per personThe people who did the workOnly by asking
Token costNobodyYour provider's usage export

Never let the agent size its own work. On one feature, the agent estimated more than twice the points the project owner believed it was worth. Take the counterfactual from history instead: how many points the team spends on each kind of work is already in the closed epics. Use a window from before the agent, or the baseline comes in deflated and the gain looks smaller than it is.

One line resists automation. Human hours cannot be deduced from the calendar. In the measurement above, deducing them gave two developers part time for three days, against six and a half hours actually worked: a factor of eight, straight into the denominator. Ask the people, on the day the epic closes, while they still remember.

Lines of code, commits and merge requests were fair proxies while writing code dominated the cost of the cycle. With an agent they grow without any matching change in the result. The AI-DLC metrics guide has the four I would put on the dashboard instead.

Every process you add needs a removal date

Score before you prescribe. A board, a role or a new stage is a prescription, and prescriptions come after diagnosis. A process created to find where work stalls is an instrument, and an instrument has a removal date. A process created to stay, in a project whose ceiling is still L1 or L2, becomes the ceiling: whatever cannot be automated ends up setting how high the team climbs.

Write the removal date next to the process when you create it. “We review every spec in a mob until the acceptance criteria stop coming back wrong” is an instrument. “We review every spec in a mob” is a ceremony, and ceremonies outlive the problem they were built for.

The playbook on one page

The agentic SDLC, in the order you adopt it

  • Required:
    Give every repository an AGENTS.md and one command that proves done.Tool-specific files point to it. This alone moves most projects a full level.
  • Required:
    Move every rule a machine can check into a failing check.Lint, types, tests. Red tests and exposed secrets never merge. The context file shrinks.
  • Required:
    Keep a work log for every unit of work, and curate it.Capture what the session learned before you clear it. Promote what you trust.
  • Required:
    Write the spec before the agent runs, not during.A written problem behind the epic, acceptance criteria on the ticket, engineering in the scope.
  • Required:
    Size units for review at the plan, and review intent and risk.Cap sessions, prefix comments by weight, watch for the tenth easy approval.
  • Required:
    Let the human gate travel with the work.People at the spec and at review. The agent works in between, in parallel.
  • Required:
    Earn the loop: boring delivery, evals, sandbox, trace.Then let deterministic detection call the agent, and a person triage.
  • Required:
    Score the eleven lines every cycle and attack the weakest below target.The weakest gate is the ceiling. The target comes from the cost of an error.
  • Anti-pattern:
    Push the spec to the agent and skip the scorecard.That is how a team ends up prompting at execution time and blaming the tool.

The build shrank. That part is done, and it was the easy part. The agent made code cheap. It did not make judgment cheap, and it did not make context cheap. Build those on purpose, in this order, and the calendar starts moving too.

FAQ

What is an agentic SDLC?

It is the software development lifecycle rebuilt for a world where AI agents write most of the code. Each stage leaves a versioned artifact the next stage can read, humans review intent and evidence instead of every line, machine-checkable rules are enforced by machines, and what each cycle learns goes back into the repository. You will also see it called the AI-native SDLC or the AI SDLC.

If our agents write code faster, why didn't our lead time drop?

Because the build was only one part of your lead time, and the rest of the lifecycle was tuned around it being slow. When the build shrank, the work moved to both sides: context on the left, review and verification on the right. Until you rebuild those, the calendar barely moves.

Where should a team start?

In the repository, not in the product requirements. An AGENTS.md at the root and one command that runs everything that must pass before merge. Better specs pay off only after the repository can receive work; before that they produce confident wrong code and frustrated developers.

How is the agentic SDLC related to AI-DLC and Spec-Driven Development?

AI-DLC is AWS's method for the same shift, with its own vocabulary and an engine that runs it. Spec-Driven Development is the practice of writing the spec before the agent runs, which is step 4 here. This playbook is the order and the ceiling that apply to any of them.

Does every team need to reach full autonomy?

No. Set the target by the cost of an error. A disposable experiment stops at L2, a customer-facing product at L3, a core platform on the critical path aims for L5. Spending on harness above your target buys nothing.

How do I measure the ROI of AI coding agents?

Measure the cost of delivering a feature against what the team would have spent without the agent. Take the counterfactual from epics closed before the agent, ask people for their real hours when the epic closes, and add the token bill. The answer is an order of magnitude, which is what a budget decision needs.