Skip to content
← articles
updated AI-DLCAI AdoptionEngineering MetricsCode ReviewAI Agents

Why AI-DLC Makes You Slower First

Adopt AI-DLC and throughput jumps on day one while cycle time gets worse. Why the J-curve happens, why code review becomes the bottleneck, and how to tell a learning curve from a real problem before you kill the pilot.

The agent removes the bottleneck you could see. The one you could not see was always review, and now it is the only one left.

Every pilot squad in the rollout I followed drew the same curve. The number of merged changes per developer went up from the first week. The time from first commit to deploy went up too, which is the wrong direction, and stayed there for a while before it came down.

That is a J-curve: things get worse before they get better. Change-management people have drawn it for decades, for every serious process change. What makes the AI-DLC version worth an article is that the two curves split. One metric looks great from day one. The other looks terrible. If you only watch the first, you declare victory too early. If you only watch the second, you kill the pilot right at the bottom of the dip, which is the most expensive moment to quit.

This is what causes the dip, why it lands on code review, and how to tell a team that is learning from a team that is stuck.

Two metrics, two curves

Start with definitions, because the whole argument depends on them.

Throughput is how much gets merged: merge requests per developer per month, for example. It measures output. Cycle time is how long one change takes from the first commit to the start of the deploy. It measures flow, how fast a single piece of work moves through the system.

In a normal team, those two move together. Write more code, ship more code. In an AI-DLC rollout they come apart, because the method makes one part of the work almost free and leaves the other part exactly as expensive as it was.

The shape of an AI-DLC adoption

  1. First weeksThroughput jumps, cycle time climbs

    The agent generates code immediately, so more changes get opened and merged. The new process is slow: long Mob Elaboration sessions, agents guessing in repositories nobody prepared for them, Inceptions restarted from scratch.

  2. The dipA queue forms in front of review

    More and bigger merge requests arrive at the same humans who reviewed them before. Changes wait. Cycle time hits its worst point while throughput still looks like a success.

  3. RecoverySmaller Units, fluent gates, prepared repositories

    The team caps the size of each Unit, the repo gets an agent-readable map, people learn which questions to answer in the mob. Cycle time starts to fall and keeps falling.

  4. AfterBoth curves point the right way

    Output stays high and each change moves faster than before the method. This is the state the pilot was supposed to reach.

The dip is in cycle time, not in throughput. That asymmetry is the whole story.

Why the curve dips

Three things slow a team down at the start, and none of them is the agent being slow.

What slows an AI-DLC team down first

The first two fade with practice. The third one is structural, and it is the one that kills pilots.
CauseWhat it looks likeWhen it goes away
Process shockSynchronous Mob Elaboration replaces async refinement. Everyone in one session for hours, more discussion than before.When the team learns which questions need everyone and which do not.
Unprepared repositoriesThe agent asks trivial questions, guesses wrong patterns, and the team restarts the Inception.When the repo has a short agent-readable map of its stack and patterns.
The review queueMore code arrives, in bigger merge requests, at the same reviewers.Only when someone changes the size of the work. It does not fix itself.
The first two fade with practice. The third one is structural, and it is the one that kills pilots.

The first two are a learning curve in the literal sense. People and repositories get better at the method, and the cost drops. You can wait those out.

The third one you cannot wait out. It is the reason the curve has its shape.

Review is the new bottleneck

Here is the mechanism, in the words of an engineer who lived it in one of the pilots. At first cycle time went up from the adoption itself. Then throughput shot up, but cycle time would not come down, because the team had fallen into a simple trap: the agent writes code very fast, people open a lot of merge requests, and everything stalls in review.

It got worse when the team parallelized. Several engineers ran their own Units at the same time, each producing large merge requests, and the repository turned chaotic. One or two Units that were supposed to be small ended up as merge requests nobody could review in one sitting.

Think of a highway that gets two extra lanes right up to a toll booth that still has one open gate. More cars arrive faster. The line at the booth gets longer. The trip takes more time, not less.

That is what happens when generation gets cheap and review does not. The agent did not make your team faster at the slowest step. It made every step before the slowest step faster, which only makes the line in front of it longer. I have written about the same shift for teams using Spec-Driven Development: once everyone has an agent, producing code stops being the scarce thing and integrating it becomes the job.

Cap the size of the Unit, not the speed of the agent

You cannot fix a review bottleneck by reviewing faster. Reviewers who rush are how bad code reaches production. You fix it by changing what arrives at review.

AI-DLC gives you the lever at the right moment. Before Construction starts, the agent proposes how to split the work into Units, and you approve that plan. That is the point to impose a size limit, not after the merge requests are open. The teams that recovered did two things: they asked the agent to estimate the size of each Unit before approving the plan, and they agreed as a team on the largest merge request a reviewer can read carefully in one sitting.

Keeping Units reviewable

  1. 01

    Ask for the size before you approve the plan.

    The plan is the cheapest place to split work. Once the code exists, splitting it means redoing it.

    Type this

    Before I approve these Units, estimate the files and lines each one will change, and flag any Unit that touches more than one layer.
  2. 02

    Split anything above the team's limit.

    A merge request with hundreds of changes gets skimmed, not reviewed. Even a task that sounds small can touch migration, entity, API, and tests at once.

    Type this

    Unit 3 is too big to review in one sitting. Split it by endpoint, and put the database migration in its own Unit.
  3. 03

    Make the limit a rule, not a memory.

    A limit that lives in one reviewer's head gets forgotten in the next session. In AI-DLC 2 a project rule is loaded at the start of every workflow.

    Type this

    Add to the project rules: no Unit may produce a merge request a reviewer cannot read carefully in one sitting. Estimate size before proposing Units.
Keeping Units reviewable: 3 rules, each with the words to type.

The exact limit is a team decision. What matters is that it exists, that it is enforced at the plan, and that the agent knows about it before it proposes anything.

The public evidence says the same thing

None of this is unique to the organization I watched. The public research points the same way, from different angles.

The METR randomized trial is the sharpest one. Experienced open-source developers, working on code they knew well, took about 19% longer with AI tools, while believing they had been about 20% faster. The lesson for an AI-DLC rollout is not that AI makes people slow. It is that feeling faster and being faster are different measurements, and the feeling is the one everybody reports in the retro. Measure cycle time. Do not ask people how it feels.

DORA’s research has made the same case for years from the other side: lead time, deployment frequency, change failure, and recovery time are what predict delivery performance. Volume of output is not on that list. Augment Code’s analysis of AIDLC summarizes the current state bluntly: AI adoption among developers is close to universal, and delivery metrics barely moved, because tools that attach to one engineer optimize keystrokes and nothing coordinates the rest.

Microsoft’s study of its own early-2026 rollout of Claude Code and Copilot CLI to tens of thousands of engineers (arXiv 2607.01418) found that adopters merged roughly 24% more pull requests than they would have otherwise. The authors are careful to say what that does not mean: a merged pull request is not the same as the value it delivers. That is the throughput curve, measured well and labeled honestly. It still tells you nothing about the queue.

A learning curve, or a real problem?

The dip is expected. That does not mean every dip is the J-curve. Sometimes the method genuinely does not fit the work, and waiting longer only makes the bill bigger.

Set a decision point at about four weeks, and say it out loud before the pilot starts. If cycle time has not started to turn by then, stop asking “is this the curve?” and start asking “what is structurally wrong?”

Still the J-curve

  1. 01Cycle time is high but has started to turn
  2. 02Mob sessions are getting shorter each week
  3. 03The agent asks fewer trivial questions about the repo
  4. 04Merge requests are getting smaller since the size rule
  5. 05Reviewers are behind, but the queue is shrinking

A structural problem

  1. 01Cycle time is flat or still climbing after a month
  2. 02The work is mostly one-line fixes the method was never meant for
  3. 03The repo still has no agent-readable map of itself
  4. 04Nobody in the mob knows the domain well enough to judge the artifacts
  5. 05Gates are being approved without questions, and rework shows up in review
Wait out the left column. Fix the right one, or stop.

The structural problems each have a fix, and most of them live in other parts of this series. An unprepared repository is a brownfield problem. Rubber-stamped gates are a gate design problem. Using a 33-stage lifecycle on a one-line change is a profile problem: AI-DLC 2 ships a Bugfix profile with 9 stages for a reason, and sometimes the right profile is no profile at all.

The manager’s job during the dip

The J-curve is a management problem before it is a technical one, because the person who decides whether the pilot survives is usually the one looking at the worst number.

The recommendation that came out of the rollout was boring and specific. Before the pilot starts, the manager tells the team that the first weeks will be slower and that this is expected, not a failure. Then the manager sits in the first two or three mobs, not as an observer, but as the person with the authority to say “we stop here and continue tomorrow.” A tech lead has the technical authority to question an artifact. The manager has the organizational authority to slow the whole session down without anyone reading it as an individual failure.

Read throughput only next to cycle time

The practical rule that comes out of all this fits in one sentence: never look at throughput alone.

Put three numbers on the same screen. Throughput, so you know output is going up. Cycle time, so you know whether a single change moves faster or just waits longer. And rework, the share of artifacts and merge requests sent back with requested changes, so you know whether the speed is real or borrowed from the next sprint. The metrics guide goes through what replaces story points and velocity once AI-DLC is running.

FAQ

Is the J-curve normal when adopting AI coding agents?

Yes. Any serious process change produces an initial dip before the gains, and AI-DLC is a process change, not a tool install.

What is specific to AI-DLC is that the dip shows up in cycle time while throughput rises from day one. Generation gets cheap immediately. Review and validation do not.

How long does the AI-DLC J-curve last?

Long enough to scare a manager and short enough to be worth it, if the causes are the learning kind. Set a decision point at about four weeks: if cycle time has not started to turn by then, look for a structural cause instead of waiting longer.

Why does cycle time get worse with AI agents?

Because the agent produces more and bigger changes than the team's review capacity can absorb. The changes wait in a queue in front of the same human reviewers, so each one takes longer from first commit to deploy, even though the code was written faster.

Does a faster or smarter model fix the dip?

No. A faster model writes code faster, which makes the review queue longer. The fix is on the human side: smaller Units, a size limit enforced at the plan, prepared repositories, and a pace that keeps gates meaningful.

Should we stop the pilot if cycle time gets worse?

Not in the first weeks. Expect it, and tell the team to expect it. Set a decision point around four weeks, and use it to separate a learning curve from a structural problem such as unprepared repositories, work that is too small for the method, or gates nobody can judge.

What should we measure instead of throughput alone?

Throughput next to cycle time, plus rework: how often artifacts and merge requests come back with requested changes. Throughput alone shows a queue growing and calls it success.

Where to go next

The J-curve is the most predictable thing in an AI-DLC rollout, and the most common reason pilots die early. Expect it, explain it before it starts, cap the size of the work at the plan, and put cycle time next to every throughput chart you show. Then decide with data at four weeks instead of with nerves at week two.