After Story Points: What to Measure in AI-DLC
Story points and velocity stop meaning anything when an agent writes the code. What to measure instead in AI-DLC: lead time from intent to tested code, rework at the gates, discovery time upstream, throughput read next to cycle time, and the data AI-DLC 2 already records for you.
A story point is a guess about how hard a task is for a human. When an agent writes the code, nobody is that human anymore.
In AI-DLC, measure four things instead of points and velocity: how long it takes from an approved intent to tested, deployable code, how often humans send the agent’s work back at the gates, how long it takes to get from a problem to an approved intent, and what actually reached users. And never read throughput alone. Read it next to cycle time, every time.
That is the answer. The rest of this guide is why the old metrics break, which pairs keep the new ones honest, and where AI-DLC 2 already writes the data down for you. If you are new to the method, start with what AI-DLC is.
I learned most of this by watching pilot squads inside a large engineering organization adopt AI-DLC. The single most useful thing I saw was a pair of charts that disagreed. Throughput went up from the first week. Cycle time got worse. A team reporting only the first chart would have declared victory, and a team reporting only the second would have killed the pilot. Both would have been wrong.
Story points measure the wrong thing now
The AI-DLC whitepaper asks the question out loud: would effort estimation, like story points, still be as critical if AI blurs the line between simple, medium, and hard tasks? And would velocity still be relevant, or should teams start replacing it with business value?
The honest answer is no, and the reason is mechanical. A story point is a proxy for human effort: how long a person needs to understand, write, and test a change. Velocity adds those proxies up per sprint. When an agent generates the code in minutes, the expensive parts move. Writing gets cheap. Deciding, reviewing, and integrating do not. A task the team would have called an 8 might take the agent ten minutes and the reviewer two days, and points cannot tell you which half you are paying for.
AI-DLC also kills the container points lived in. Sprints become bolts, delivery slices measured in hours or days. Estimating a two-week sprint in points made sense. Estimating a three-hour bolt in points is theater.
Four metrics to replace points and velocity
The whitepaper only proposes business value. Augment Code’s guide makes the same point from the adoption side: individual productivity rises while delivery barely moves. The teams I followed converged on a fuller set, and each one replaces a specific old number for a specific reason.
What replaces what
| Old metric | AI-DLC metric | What it actually measures |
|---|---|---|
| Story points | Construction Lead Time | Elapsed time from an approved intent to a tested, deployable Unit. The agent's speed and the humans' review time, together. |
| Velocity | Business value delivered | What reached users and moved the metric the intent named. Points burned never did that. |
| Bug count | Rework rate | How often humans answer Request Changes at a gate, by stage. Where the agent keeps getting it wrong. |
| Time in sprint | Discovery Lead Time | Elapsed time from a problem being identified to an approved intent. Everything upstream of the agent. |
The last row is the one teams forget. When construction compresses from weeks to days, the slowest part of delivery is often everything before it: the meetings, the documents, the approvals that turn “customers complain about X” into a clear intent. In the rollout I followed, discovery took far longer than construction, and nobody had been measuring it because product and design work had never been tracked the way tickets were. If you only measure the agent’s half, you optimize the half that is already fast. I wrote about that upstream gap in AI-DLC’s blind spot.
Throughput lies when you read it alone
Here is the pattern I saw in every pilot squad. In week one, the agent produces code immediately, so merged changes per developer jump. Nobody redesigned code review for the new volume, so changes pile up waiting for a human, and the time from first commit to deploy goes up before it comes down.
Throughput says the pilot is a success. Cycle time says it is failing. Neither is the truth on its own. Microsoft’s study of its own rollout of Claude Code and Copilot CLI found adopters merged roughly 24% more pull requests, and the authors were careful to say that a merged pull request “is not the same as the value it delivers” (arXiv 2607.01418). That caveat is the whole point of this section.
Read throughput and cycle time together
Stuck
Little ships and what ships takes long. The method is not working yet, or the repository is not readable by an agent.
Review jam
The agent produces plenty, humans cannot absorb it. The J-curve dip. Shrink the Units and redesign review, do not kill the pilot.
Careful or idle
Fast when it ships, but little ships. Either the team is cautious on purpose or the agent is barely used.
Where you want to be
More ships and it ships faster. Usually reached after the dip, never on day one.
A team moving from bottom left to top right is not failing. It is in the dip on the way to bottom right.
The fix for the review jam is not a faster reviewer. It is smaller pieces at the source. Ask the agent to estimate the size of each Unit before Construction starts, split anything that would produce a huge merge request (by endpoint, migrations on their own), and agree with the team on a size reviewers can actually read. The whole J-curve, and when to call it a structural problem instead of a learning curve, is in why AI-DLC makes you slower first.
AI-DLC 2 already records most of this
The good news about AI-DLC 2 is that you do not need a new analytics tool to start. Every piece of work gets a record folder with an append-only audit log that knows 108 event types, and the engine compiles a runtime graph from it after every stage transition.
Questions you can answer from the record
| The question | Where AI-DLC 2 writes it down |
|---|---|
| Where does the agent keep getting it wrong? | GATE_APPROVED and GATE_REJECTED per stage, plus the revision count in the state file. Rejections over total decisions is your rework rate by stage. |
| How long do humans keep the work waiting? | STAGE_AWAITING_APPROVAL to GATE_APPROVED. /aidlc --status shows how long the open gate has been waiting; doctor flags gates waiting over 24 hours. |
| Are the deterministic checks catching anything? | SENSOR_FIRED, SENSOR_PASSED, SENSOR_FAILED, tallied in the runtime summary. |
| Is the team teaching the agent? | Learnings captured at the gates, counted in the runtime summary, and the rules they added to project and team memory. |
| Does the merged code match what was reviewed? | aidlc attest classifies each changed path as verified, drifted, unattested, or unverifiable against the review receipts. |
The fastest way in is /aidlc-session-cost, which prints duration, stage outcomes, sensor results, and learnings for the current workflow, or the command it reads from, which also outputs JSON:
aidlc engine runtime summary --json
One signal I would add on top: the time between a gate opening and the approval. A requirements document approved forty seconds after it appeared was not read. It is a rough measure, but a team where that number keeps shrinking through the afternoon is a team getting tired, and tired gates are where bad decisions get through. More on that in gates are a loss function.
Harness metrics: is the agent getting better?
The four metrics above tell you whether delivery improved. A second set tells you whether your setup around the agent is improving: the context, the rules, the checks. That setup is the harness, and it is the part you control.
Four harness metrics
- 01
Escalation rate
The share of tasks that end up needing a human to step in beyond the planned gates. It should fall as the harness matures. If it does not, the rules and context are not learning.
- 02
Agent spiral rate
The share of tasks where the agent loops without progress: the same fix attempted three times, the same question asked again. Read it from session logs. A rising spiral rate usually means context rot or a vague spec.
- 03
Code retention rate
The share of agent-written code still present N days later, straight from git history. No instrumentation needed.
- 04
Attention allocation
The share of engineers' time spent on decisions (specifying, reviewing, choosing) versus operating. Self-reported, so treat it as a conversation, not a number.
Retention is the one that looks simplest and is the easiest to misread. A large study of 201 open-source projects found agent-written code survives longer than human code, with a 16% lower hazard of being modified (Will It Survive?). That could mean the code is good. It could also mean nobody feels they own it, so nobody touches it. Another study found AI-generated files get maintained less often, and when they are, humans do the large majority of the work (arXiv 2605.06464). The first paper’s own conclusion is the right one: the bottleneck may not be generation quality but the organizational practices around the code. Read retention next to who modifies the code and why. High retention plus nobody touching a module that should be changing is not quality. It is orphaned code.
Three ways to ruin your metrics
Every metric becomes a target the moment you put it in a performance review. These are the three I watched happen, or nearly happen.
Measurement mistakes in AI-DLC rollouts
- Anti-pattern:Turning a readiness score into a KPI.Tools that score how readable a repository is for an agent are useful as diagnosis. Make the score a target and teams create an empty CLAUDE.md to raise it.
- Anti-pattern:Trusting how fast people feel.In METR's controlled study, experienced developers were 19% slower with AI tools while believing they were 20% faster. Self-reported speed is not a metric.
- Anti-pattern:Counting merged changes as output.Throughput rises in week one no matter what. Without cycle time and rework next to it, you are measuring how much the agent typed.
- Required:Measure the pairs, per stage, from the record.Lead time with throughput, rework by stage, discovery time upstream. All of it already sits in the audit log.
The METR result deserves one more line, because it is the cleanest evidence that perception fails here (METR). The developers were not careless. They were experienced people working on their own repositories, and their sense of speed was wrong in the opposite direction from the truth. If perception fails for them, it fails for your team. Measure.
A dashboard for the first quarter
What to put on the wall
- Required:Cycle time (median and 75th percentile) next to throughput, weekly.Never one without the other.
- Required:Rework rate by stage.Requirements, design, and code plans separately. The stage with the highest rate is where your context or your intent is weakest.
- Required:Discovery Lead Time.From problem identified to approved intent. Start measuring it even if it embarrasses you.
- Required:Gate wait time and time-to-approve.Long waits mean a review bottleneck. Instant approvals mean fatigue.
- Optional:Code retention at 30 and 90 days, with who modified it.From git history. Read with ownership, not alone.
- Required:A four-week decision point.If cycle time has not started to come down after about four weeks, find out whether it is the J-curve or a structural problem before deciding anything.
None of this needs a data team. The audit log, git history, and a spreadsheet will carry you through the first quarter. What it needs is the discipline to report the pair that disagrees, especially when one half looks great. And it changes what you reward, which is its own problem: when the job becomes specifying and reviewing, the competency matrix has to follow. That is in from squad to pod.
FAQ
Do story points still make sense with AI agents?
Not much. Points estimate human effort, and when an agent writes the code the effort moves from typing to deciding and reviewing. A task can take the agent minutes and the reviewer days.
The AI-DLC whitepaper itself asks whether estimation and velocity still matter, and suggests business value instead.
What replaces velocity in AI-DLC?
Business value delivered, read together with Construction Lead Time (approved intent to tested code), rework rate at the gates, and Discovery Lead Time upstream. Throughput stays useful only when read next to cycle time.
How do I measure rework in AI-DLC?
Count Request Changes against total gate decisions, per stage. AI-DLC 2 records both as GATE_REJECTED and GATE_APPROVED in the audit log, plus a revision count per stage in the state file.
Are DORA metrics still useful with AI agents?
Yes, especially lead time for changes, which is closest to what agents affect. What changes is that you need to add the upstream (discovery time) and the gates (rework, wait time), because that is where AI-DLC moves the bottleneck.
Why did our cycle time get worse after adopting AI-DLC?
Because the agent produces more changes than your review process was built to absorb. It is the most predictable pattern in AI-DLC rollouts. Shrink the Units, redesign review, and give it about four weeks before deciding.
Where does AI-DLC 2 store workflow data?
In a record folder per piece of work under aidlc/spaces/<space>/intents/, with a state file, every artifact, and an append-only audit log. aidlc engine runtime summary --json aggregates it.
Where to go next
Measure the human part. The agent’s speed will take care of itself in the first week, and it will look wonderful on a chart. What decides whether AI-DLC works is how long people take to decide, how often they have to send work back, and whether what ships is what someone meant.
The newsletter
Don’t Code, Specify. A weekly dispatch from where AI agents meet real production. No hype, just what shipped and what broke.
Subscribe on Substack (opens in a new tab)