Skip to content
← articles
updated AI-DLCContext EngineeringAI AgentsBrownfieldSoftware Engineering

AI-DLC in Brownfield: Reverse Engineering, 1M Context and the 75% Rule

Running AI-DLC on code that already exists: scope reverse engineering to the feature and run it before the mob, budget the context window, clear it at gates, and treat context compaction as the governance failure it is.

A greenfield project has nothing for the agent to misunderstand. A brownfield project has years of decisions the agent never saw. Context management is how you keep those decisions inside the window long enough to matter.

Most AI-DLC demos start from an empty folder. The official AWS tutorial builds a river crossing puzzle, and the first thing the workflow does is detect that the project is greenfield and skip reverse engineering, because there is nothing to reverse. That makes a clean demo. It is not the job most of us have.

The job most of us have is a service that has been in production for four years, with decisions nobody wrote down, a test suite that covers some of it, and a neighbor service it talks to through a contract that lives in someone’s head. That is brownfield, and it is where AI-DLC either earns its ceremony or wastes it.

In the pilots I followed inside a large engineering organization, two things decided whether brownfield worked. How the team ran reverse engineering, and how it treated the context window. Neither is a tooling problem. Both are habits, and both have rules you can write down.

Reverse engineering is where the agent meets your code

When you start an AI-DLC 2 workflow, the Initialization phase detects the workspace. If it finds existing code, Inception opens with stage 2.1, Reverse Engineering, before the agent is allowed to ask you a single requirements question.

The reason is simple. A model that has not read your code will still propose changes to it. It will invent a pattern you do not use, duplicate a helper you already have, and break an implicit contract between two modules because nothing told it the contract existed. The original whitepaper put the fix in one line: in brownfield, first elevate the code into a semantic model, a static one (the components and their relationships) and a dynamic one (how they interact in the use cases that matter), so the context the AI works from is concise and accurate.

Think of the first week of a senior hire. You do not hand them a ticket on day one. You let them read the code, draw the boxes, and ask why the payments module talks to the ledger through a queue. Reverse engineering is that week, compressed into a stage, with the drawing written to a file.

AI-DLC 2 runs this stage as a two-link pipeline. A developer agent scans the code. An architect agent reads the scan and writes the synthesis. Then you approve it, like any other stage.

How AI-DLC 2 opens a brownfield workflow

Input

You run /aidlc in a repository that already has code

  1. 0.2Workspace Detection flags it as brownfield

    A rule-based scan of file extensions, manifests, and known config files. No model call, no gate.

  2. 2.1aThe developer agent scans the code

    Modules, entry points, how the pieces talk to each other, the stack and its versions.

  3. 2.1bThe architect agent writes the synthesis

    The reverse engineering artifacts land in the intent's record folder, where every later stage can read them.

  4. GATEYou approve the picture of your own system

    If the agent misread the architecture here, every requirement, design, and Unit after it inherits the mistake.

Output

Requirements questions that start from how your code actually works, not from how an average codebase would work.

How AI-DLC 2 opens a brownfield workflow: flow of 4 steps from “You run /aidlc in a repository that already has code” resulting in “Requirements questions that start from how your code actually works, not from how an average codebase would work.”.

That gate is the most underrated one in the lifecycle. People skim it because it describes code they already know, and that is exactly why it deserves a real read: it is the one place to check the agent’s model of your system against yours before anything is built on top of it.

Scope it to the feature, not the repository

The naive move is to point reverse engineering at the whole repository. More context sounds safer. It is not.

Every token of the reverse engineering output competes for the same window as the requirements, the design, and later the code. A full map of a large service burns a big share of that window on modules the feature will never touch. And the more irrelevant material sits in the context, the more chances the model has to pull the wrong piece into its answer. One of the engineers maintaining the setup in the rollout put it bluntly: more information means more chances to hallucinate.

What worked was focused reverse engineering. The agent maps the modules the feature will use, the boundaries it will cross, and the contracts it will depend on. It leaves the rest alone. You lose a complete architecture document. You gain a window with room left for the work.

Run it before the mob, not during it

On a small repository, reverse engineering took about twenty minutes. On a large one, longer. Twenty minutes is nothing for one engineer. It is a lot for six people sitting in a Mob Elaboration session watching a progress indicator.

The pattern the teams settled on: whoever drives the session installs the workflow, starts it, and lets reverse engineering finish before the meeting. The team joins when the synthesis is approved and the requirements questions are ready to answer. The synchronous time goes to decisions, which is the only thing synchronous time is good for.

One feature, several repositories

Real features rarely live in one repository. A change touches the API, the front-end, maybe a mobile client, and a service owned by another team.

The first version of AI-DLC was single-repository. Teams worked around it by splitting the work by stack: one session for the back-end, one for the front-end, with the contract between them agreed before either started. When the back-end needed to understand the front-end’s expectations, they added the front-end folder as an extra read-only directory in the agent’s context, so reverse engineering could see the mock or the client code without owning it.

AI-DLC 2 made multi-repository work a first-class idea. An intent can record the set of repositories it spans, either explicitly or by discovering sibling repositories in the workspace. Reverse engineering then runs one complete scan and synthesis chain per repository before you approve, and Construction anchors each git operation to the right repository. There is also a Contract Design stage in Inception for the interfaces between Units.

None of that removes the hard part. Two repositories owned by two teams still need a contract both sides approve. The engine can map both codebases. It cannot make the other team agree with you.

Budget the window like memory, because it is

The context window is the agent’s working memory, and an AI-DLC Inception is a heavy user of it. The agent reads every requirement you give it, the reverse engineering output, its own questions and your answers, and the documents it generates along the way.

In the rollout, a full Inception on a real service landed somewhere around 500,000 to 600,000 tokens. That fits in one window of a one-million-token model and does not fit in a smaller one without the harness compacting the conversation along the way. Construction is the opposite: each Unit has its own spec, so the context needed per Unit is small.

Where the window goes

The big window is for the phase that reads the most. Paying for it during Construction buys you nothing.
MomentWhat fills the contextWhat worked
InceptionReverse engineering, requirements, questions and answers, design documents.Run it in one window on a 1M-token model. Check usage before you start.
Construction, per UnitThe Unit's spec, the code it touches, the tests.Clear between Units. A smaller, cheaper model is often enough.
A long discovery conversationBack and forth, corrections, abandoned directions.Write the conclusions to a file and start clean. Do not keep talking.
The big window is for the phase that reads the most. Paying for it during Construction buys you nothing.

In Claude Code you can see how full the window is with /context. Look before you start an Inception. Finding out at 70% that you are on the small window is a bad afternoon.

The 75% rule

Here is the heuristic the teams converged on: once the window passes roughly three quarters full, quality drops. The agent starts repeating your own answers back to you, loses track of decisions made an hour earlier, and invents details with the same confident tone it used when it was right.

This is not a law and the exact number moves with the model. The mechanism behind it is well documented. Models attend best to the beginning and the end of their context and skim the middle, so the decision you made in the middle of a long session is the one most likely to get lost. Practitioners call the slow decline context rot. An engineer from one of the pilot squads gave the practical version: at around half a million tokens, clean up and resume from the state file, because that beats watching the agent start to hallucinate.

Clear at gates, never in the middle of a phase

Clearing the context is not giving up on the task. It is throwing away the noise of the session: the wrong turns, the corrections, the intermediate drafts nobody needs anymore. The work itself lives in files. The session was always disposable.

The rule is about when. Clear at a checkpoint: after reverse engineering is approved, after the requirements are approved, between Construction Units. Never in the middle of a phase that is still producing its artifact, because then the half-finished reasoning exists nowhere.

AI-DLC keeps a state file per intent, aidlc-state.md, with the current phase, the stage, and the status of every stage. In AI-DLC 2 it lives in the intent’s record folder next to an append-only audit log, and a hook writes a recovery breadcrumb before the harness compacts. That is what makes clearing safe. A new session reads the state and picks up where the last one stopped.

How to clear without losing work

  1. 01

    Commit and push before you clear.

    The artifacts are the memory. If they are not committed, clearing deletes the only copy of the reasoning that produced them.

    Type this

    Commit the approved artifacts for this stage with a message that names the stage, then push.
  2. 02

    Clear only at a gate.

    A gate means the artifact is finished and approved. Anything in flight should finish first.

    Type this

    /clear
  3. 03

    Resume from the state file, not from your memory of the session.

    The state file knows the phase, the stage, and what is done. You remember what felt important.

    Type this

    /aidlc --status
  4. 04

    Before clearing a stuck session, make the agent write down what it learned.

    A debugging session that went nowhere still ruled things out. Save that, drop the noise, and the new session starts with the same depth and none of the confusion.

    Type this

    Do not change any code. Write a summary of this issue to notes/issue-summary.md: what we tried, what we ruled out, and what we still suspect.
How to clear without losing work: 4 rules, each with the words to type.

Compaction quietly breaks your rules

When the window fills and you do not clear it, the harness does something on your behalf. It compacts: it replaces the earlier conversation with a summary so the session can continue. That sounds harmless. It is not.

A summary optimizes for continuing the task. It keeps what looks central and drops what looks peripheral. And a constraint you stated once, early in the session (“do not touch the billing module”, “never delete a migration”), looks peripheral right up to the moment the agent violates it.

This is no longer a hunch. Two 2026 papers measured it.

What compaction does to constraints

0% to 30%
constraint violationsbefore vs after compaction, across seven model families
59%
worst caseviolation rate for some models
17%
constraints kepton average, by current compactors
Governance Decay (arXiv 2606.22528) and Lost in Compaction (arXiv 2608.11242).

The Governance Decay study ran 1,323 agent episodes across seven model families. With the policy in full context, agents violated it 0% of the time. After compaction, 30%, and 59% for some models. When the constraint survived the summary, violations stayed at zero. When the summary dropped it, they jumped. The authors propose pinning governance constraints outside the lossy summary, which brought violations back to zero in their benchmark.

The Lost in Compaction study looked at instructions users give for the rest of a session, like “do not delete any emails until I confirm”. Current compactors kept 17% of them on average, and most did worse than not compacting at all. An extractor that pulls those constraints out before compaction kept over 90%.

The practical lesson for brownfield work is short. A rule you typed into the chat lives inside the conversation that gets summarized. A rule in a file the harness loads on its own, your AGENTS.md or CLAUDE.md, or the rule files AI-DLC 2 keeps in its space memory, lives outside it. Put every constraint that matters in a file. The AGENTS.md guide covers what belongs there.

A bigger window treats the symptom

When an agent starts repeating itself in a long session, the reflex is to switch to a model with a bigger window. A product manager in the rollout did exactly that during a long discovery: the agent broke, started echoing their answers back, and a larger model fixed it.

It fixed the symptom. The cause was a session that kept accumulating context that should have been written down and dropped. A bigger window lets you accumulate longer before the same failure. It does not stop the failure, and it costs more per turn while you wait for it.

The cure is the one the method already has. Every AI-DLC stage writes its artifact to a file, and the next stage reads the file instead of the conversation. Inside a stage, do the same thing by hand: when a thread of discussion reaches a conclusion, write the conclusion down and start clean. This is context engineering applied to a lifecycle. Chat is scratch paper. Files are memory. If something the next session needs only exists in the conversation, it is one compaction away from being gone.

Is your brownfield repo ready for AI-DLC?

Brownfield readiness

  • Required:
    The repo has a short AGENTS.md or CLAUDE.md with the stack, the patterns, the forbidden libraries and why, and one example of each pattern.Every gap here becomes a question the agent asks during Inception, or worse, a guess it makes without asking.
  • Required:
    Reverse engineering is scoped to the modules the feature touches.A full map of the repo spends the window on code the feature will never call.
  • Required:
    Reverse engineering runs before the mob, not during it.The team joins when the synthesis is approved and the questions are ready.
  • Required:
    Cross-repository contracts are agreed before the sessions start.AI-DLC 2 can scan several repositories in one intent. It cannot negotiate with the team that owns the other one.
  • Required:
    Inception runs on a 1M-token model, and someone checks the window before starting.Find out you are on the small window at the start, not at 70%.
  • Required:
    The team clears at gates, around three quarters full, after commit and push.Never in the middle of a phase. Resume from the state file.
  • Required:
    Every constraint that matters lives in a file, not only in the chat.Compaction keeps a fraction of what you said. It does not touch what is on disk.
Two or more unchecked and the agent will spend your Inception rediscovering what your team already knows.

FAQ

What is context rot?

The gradual drop in an AI agent's quality as its context window fills up. The model attends best to the start and the end of its context, so details from the middle of a long session get lost, and the agent starts repeating itself or contradicting earlier decisions.

In practice, it becomes visible somewhere past three quarters of the window. The fix is to externalize decisions into files and clear the context at natural checkpoints.

What is context compaction, and why is it a risk?

Compaction is what a harness does when the window fills: it replaces earlier conversation with a summary so the session can continue. The summary keeps what looks central and drops the rest.

The risk is that constraints you stated once look peripheral. A 2026 study measured constraint violations going from 0% to 30% after compaction, and up to 59% for some models. Keep important rules in files the harness loads on its own.

Does AI-DLC work on legacy code?

Yes, and it was designed for it. In brownfield projects Inception starts with a Reverse Engineering stage that models the existing code before any change is proposed. The quality of that stage, and of the AGENTS.md or CLAUDE.md the repo already has, decides how much the agent has to guess.

Does AI-DLC 2 support multiple repositories?

Yes. An AI-DLC 2 intent can record the repositories it spans, explicitly or by discovering sibling repositories in the workspace. Reverse Engineering runs one complete chain per repository, and Construction anchors git operations to each repository. The first version was single-repository.

How big a context window do I need for AI-DLC?

For Inception on a real service, plan on a one-million-token window. The full Inception in the rollout I followed landed around 500,000 to 600,000 tokens. Construction Units need much less, so a smaller and cheaper model often works there.

Should I use /compact or /clear?

In an AI-DLC workflow, prefer /clear at a gate, after committing the approved artifacts, and resume from the state file. Compaction summarizes the session and can silently drop constraints. Clearing throws away the noise and keeps the work, because the work is already in files.

Where to go next

Brownfield is where AI-DLC either pays or burns money, and the difference is mostly hygiene. Map only what the feature touches, map it before the meeting, give Inception the big window, clear at gates, and put every rule that matters on disk. The agent does not need to remember your system. It needs to be able to read it again.