Insights28/7/2026
What agentic coding teaches us: a real world case study
The acceleration promised by agentic AI in software development is real, measurable, and in some areas striking. But it is also partial: it manifests strongly where the work is technically well-defined, and diminishes sharply—sometimes to a standstill—wherever architectural decisions, coordination with third parties, and human validation of knowledge come into play. The competitive advantage for companies adopting these tools does not lie in having faster agents: it lies in knowing precisely where their autonomy ends and where human governance of the process must begin.
THE RISK OF MEASURING ONLY TECHNICAL PRODUCTIVITY AFTER ADOPTING AGENTIC AI
Many organizations introducing agentic AI into software development measure the initiative's success by looking only at technical productivity: lines of code generated, test cases produced, person-days saved on isolated tasks. From these figures, it is tempting to infer, by extension, that the entire project is moving faster.
This inference is risky. A real software project is not made up of technical tasks alone: it includes architectural choices with cross-cutting impact on the system, negotiations with external vendors or partners, and—in projects inheriting pre-existing code (brownfield)—human validation of the knowledge agents extract from legacy systems that are often layered and redundant. If the introduction of agents is not accompanied by a rethinking of these steps, technical acceleration collides with an unchanged bottleneck—and the project as a whole does not finish any sooner.
THE REAL NUMBERS: FROM 3X TECHNICAL ACCELERATION TO TWO WEEKS AT A STANDSTILL

A concrete case makes this dynamic evident: an enterprise project in the education sector, run with an agentic development model and partly on a brownfield basis (pre-existing components and logic to recover and refactor).
- Infrastructure setup activities estimated at 30–40 person-days were completed in 11 person-days—an acceleration of roughly 3x.
- Over 900 test cases were written and executed (front-end and back-end) since the project's start.
- Refactoring and documentation analysis saw an even greater acceleration than infrastructure setup.
- Against these figures, development came to a substantial standstill for two weeks—caused not by technical limitations, but by pending decisions with an external partner.
EXECUTION LAYER VS. SITUATIONAL LAYER: THE DISTINCTION THAT EXPLAINS THE GAP

The delivery team framed this distinction explicitly, speaking of an "execution layer" (technical, where agents work well) and a "situational layer" (decision-making, where acceleration drops to zero): the same AI initiative produces opposite results depending on which of the two layers is under observation.
The execution layer includes tasks with a closed specification: writing code, running front-end and back-end tests, refactoring on known patterns, generating documentation. These are tasks where the goal is defined in advance and the success criterion can be verified autonomously by the agent—which is why, in this project, they produced a roughly 3x acceleration on infrastructure setup and over 900 test cases in reduced time.
The situational layer, by contrast, includes decisions that require context not reducible to a specification: an architectural choice with cross-system impact, validation of knowledge extracted from a legacy system before it can be reused, a trade-off to be negotiated with an external partner. In this project, for instance, the two-week slowdown did not originate from a technical limitation of the team, but from a decision left pending on the external partner's side: the agent did not have—and could not have had—the elements needed to resolve that issue in place of the people involved.
This distinction matters because the two metrics tell opposite stories: technical productivity in the execution layer can run at full speed at the very same moment the situational layer is holding the entire project still.
THE COMPETITIVE ADVANTAGE LIES IN GOVERNANCE, NOT IN THE TOOL
This case confirms a position adesso.it has long argued in the debate on AI and software development: competitive advantage does not reside in the AI tool itself, but in the ability to govern its effects. Writing more code, faster, is not an end in itself—it becomes one only when paired with explicit governance over what can be delegated to agents and what must remain in the team's hands: architectural choices, validation of knowledge extracted from legacy systems, and coordination with external stakeholders.
FOUR PRACTICAL ACTIONS TO GOVERN AGENTIC CODING IN SOFTWARE PROJECTS
For organizations adopting or evaluating agentic AI in software development, four concrete steps:
- Explicitly distinguish, at the planning stage, which activities belong to the execution layer (delegable to agents) and which to the situational layer (decision-making, non-delegable)—do not assume that technical acceleration automatically propagates to the project as a whole.
- Measure project productivity, not only technical productivity. Track the actual time to release, not just the person-days saved on individual tasks: it is the comparison between these two figures that reveals where the real bottleneck lies.
- In brownfield contexts, impose a human validation step on knowledge extracted from legacy code before using it as input for agents, to avoid automating duplications or choices that are no longer valid.
- Treat decision-making dependencies on external parties as a first-class project risk, with dedicated owners and timelines—exactly as one would with a technical risk.
FAQ
1.What is governed agentic coding?
It is a software development model in which AI agents autonomously execute well-defined technical tasks—code writing, testing, refactoring, documentation—while a human governance framework establishes which decisions remain non-delegable: architectural choices, validation of knowledge extracted from legacy systems, and coordination with third parties.
2.Why does AI accelerate development but not always reduce a project's overall timeline?
Because a real project is not made up of technical tasks alone. It also includes architectural decisions, negotiations with external partners, and validations that are human by nature. If technical acceleration is not accompanied by a review of these steps, the project remains bound to the pace of its slowest bottleneck—which is often organizational, not technical.
3.How should productivity be measured correctly when introducing agentic AI?
Two metrics need to be distinguished: technical productivity (execution speed on defined tasks, e.g. person-days saved, test cases produced) and project productivity (actual time to release). Only the latter captures the effect of decision-making bottlenecks, and it should be tracked with dedicated tools—not inferred from the former.
4.What role does human validation play when agents work on legacy (brownfield) code?
It is a critical step. Legacy systems often contain duplications and layered logic: if the knowledge an agent extracts from that code is not validated by a person before being used to generate new code, there is a risk of building functionality that is misaligned with the client's actual expectations.
CONTACT US
fill out the form
to get in touch with us
