The phrase "software factory" has been around for decades and it was always slightly embarrassing, because software was never much like manufacturing. The analogy is getting less embarrassing, and that is worth taking seriously rather than celebrating.
The shift underneath it is a change in interface. Generative AI operated on a prompt-in, content-out model: it suggested. Agentic systems operate on goal-in, outcome-out: given an objective like adding a subscription tier, an agent plans the steps, calls the APIs, updates the schema, writes the code and runs the tests.1 Once the unit of instruction becomes a goal rather than a keystroke, the thing you are managing stops being a tool and starts being a production line.
What is actually being claimed
BCG Platinion's account of the agentic software factory is the most concrete published version I have found, and it makes a specific structural claim: autonomous agents build, test and ship around the clock while humans define business intent and review outcomes, with organisations at that level reporting average productivity gains of three to five times.2 They date the transition to 2026, on the grounds that three things converged — models got substantially more capable while inference costs fell, a new generation of coding agents achieved a step change in autonomous execution, and the understanding of how to harness them matured.2
The individual data points are more interesting than the multiplier. OpenAI reportedly built a million-line product in five months with three engineers and no manually written code.2 Spotify's engineers have reportedly not written a line of code since December 2025, with the company merging around 650 AI-generated pull requests a month and cutting large-scale migration time by about 90%.2 BCG's own five-day task force converted two business-critical enterprise applications originally estimated at hundreds of person-days, reporting 20% productivity gains per application after two days and above 50% at scale.2
Two caveats before those numbers get load-bearing. They are vendor and practitioner reports, not controlled studies — the same document also cites the earlier copilot era delivering productivity improvements of "up to 30%," which is exactly the kind of ceiling figure that reads differently once measured independently.2 And Spotify's own engineering write-up frames its background coding agent as a journey of 1,500-plus pull requests, which is a rather more incremental story than the headline implies.3
The independent evidence is more ambivalent
Set the case studies beside the closest thing the industry has to population-scale measurement. DORA's 2025 report surveyed nearly 5,000 technology professionals alongside more than 100 hours of qualitative data, and its central finding is that AI does not fix a team — it amplifies what is already there.4 Adoption is effectively universal: 90% of respondents use AI in daily work, up 14% year on year, with a median of about two hours a day.5
The direction of travel improved in one dimension and did not in another. In 2024, AI adoption was associated with reduced throughput; in 2025 that reversed, and AI now correlates with improved throughput, product performance and time spent on valuable work.5 But delivery instability remained elevated. DORA's own phrasing is unusually direct: AI adoption "not only fails to fix instability, it is currently associated with increasing instability."6 AI also had no measurable effect on friction or burnout, which stay properties of the organisational system rather than the workstation.5
Telemetry from Faros AI points the same way with sharper edges: epics completed per developer up 66.2%, and median time in pull-request review up 441% against 91% in their previous dataset.7 More work arriving, and a review stage buckling under it. If you have ever watched a queue form behind a bottleneck you know what that shape means — the constraint moved downstream, it did not disappear.
The measurement problem is worse than it looks
There is a further reason to be careful with productivity claims, and it is the most useful single study in this whole area. METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real tasks in mature repositories they had contributed to for an average of five years, randomly assigning each task to allow or disallow early-2025 AI tools.8
Developers forecast that AI would cut completion time by 24%. Measured, it increased completion time by 19% — and afterwards, having lived through the slowdown, they still estimated AI had made them 20% faster.8 That is roughly a 43-point calibration error, and a reversal in direction.9 Economics and machine-learning experts asked to predict the result overestimated the speedup even more.9
The honest reading is narrow. This measured early-2025 tools in high-context repositories, and METR itself now treats it as a snapshot and has changed the design after finding selection effects severe enough to warrant it.10 Contemporaneous trials on constrained, self-contained tasks found large speedups in the other direction, with the divergence best explained by task complexity and prior familiarity.9 The lesson is not that agents do not work. It is that a "3–5x productivity gain" sourced from practitioner self-report should be read as a hypothesis.
The case studies and the population data are not actually in conflict. Both say the same thing from different ends: output goes up sharply where the surrounding system is strong, and instability goes up where it is not. DORA's amplifier framing is the reconciliation.4
Two competencies the factory actually runs on
If the multiplier is contingent, the interesting question becomes what it is contingent on. BCG names two competencies, and I think they are the right two.
Harness engineering is the discipline of designing and continuously refining the factory itself — encoding organisational standards into agent instructions and feeding information to its assembly lines.2 The practical artefact is the agent harness: markdown rule files, tooling and automated hooks that instruct agents how to behave at each stage — the factory's operating manual, written for machines.2 And as in a physical plant, each delivery archetype gets its own line: greenfield, brownfield and legacy modernisation each need a tailored harness.2
Intent thinking is the ability to translate business needs into precise, testable descriptions of desired outcomes — explicitly not prompt engineering, because it requires business and technical depth no model substitutes for.2 The formulation I find most useful: intent thinking specifies not only what the software should do but what "correct" looks like, which edge cases matter, and which trade-offs are acceptable.2
That is a real change to the shape of the job, and the labour-market forecasts reflect it. The World Economic Forum estimates 59% of the global workforce will need reskilling; Gartner projects 80% of engineers must upskill through 2027.2
Where it breaks
Three failure modes, all of which I have watched happen in smaller form.
The first is automating chaos. Agents are only as effective as the codified knowledge they can reach, and in most enterprises the critical knowledge is precisely what is least documented — architecture decisions in Slack threads, business rules in the heads of long-tenured engineers, stale API docs. Skip the codification step and you automate the dysfunction.2 DORA reaches this from the data side: without strong automated testing, mature version control and fast feedback, an increase in change volume simply produces instability.7
The second is the review bottleneck. When humans stop writing code but remain the approval path, throughput multiplies into a queue — which is what a 441% increase in review time looks like.7 The structural answer is to stop reviewing lines and verify against intent: scenario-based behavioural tests derived from requirements and stored outside the agents' accessible codebase, plus static analysis, architecture conformance checks and red-team agents probing edge cases.2
The third is orchestration debt. Enterprises now run around a dozen agents on average, and roughly half of them do not talk to each other — which caps what any of them can accomplish.11 A dozen disconnected agents is not a factory. It is a dozen power tools.
Why this matters more outside Silicon Valley
The strategic consequences BCG draws are, I think, understated for anyone not sitting on a large engineering budget. Most enterprise IT spend is consumed by maintenance, and legacy modernisation programmes shelved as prohibitively expensive become viable when the economics shift — unlocking stranded capital.2 The build-versus-buy threshold moves, because custom solutions that were too expensive now are not.2 And competitive advantage relocates to proprietary data, domain knowledge, ecosystem and intent quality, precisely because anyone can have agents build software.2
Read that list from Dhaka rather than from Munich and it is close to an argument for our existence. If the differentiating inputs are domain knowledge and the clarity of intent rather than the headcount able to write code, then a small senior team with deep vertical understanding is not at a structural disadvantage to a large one. That is the same reasoning behind running one team across Gigabit, GigaCommerce, Top Dentistry and FormBridge: the harness, the verification scaffolding and the operating discipline are the expensive parts, and they transfer.
What does not transfer, and what I would not outsource to a factory, is knowing what correct looks like in a specific business. That is the intent-thinking claim, and it is the part of the job that got more valuable rather than less.
My position, stated plainly. The agentic factory is real as an architecture and unproven as a multiplier. The mechanisms it describes — intent-driven specification, codified knowledge, harness engineering, layered verification, auditability by design2 — are good engineering whether or not you get 3x. The productivity figures should be treated as claims awaiting the kind of controlled measurement that has, so far, embarrassed nearly everyone's intuitions.8
The metric I would actually track is not velocity. It is defect escape rate against delivery volume, because that ratio is the only one that tells you whether you built a factory or just removed the brakes.