Product StudioAI TransformationTech Due DiligenceBlogAboutBook a call →
AI Strategy · October 2026 · 9 min read

Where human judgment actually sits.

The useful question is no longer whether agents can write the code. It is where you put the humans once they can.

Human judgment does not leave the software factory. It relocates. That is the argument Addy Osmani makes in O'Reilly Radar, and we think it is the right one. Judgment moves upstream into intent and system shape, and downstream into evidence, risk and ownership. What it does not do is disappear. The hard and skilled job, as he puts it, is deciding where to put each switch. This post is our answer: which switches we leave on, why, and what we do about the one cost that this model genuinely creates.

The category has a name now, and that matters

software factoryAI software factorylight and dark factoriesagentic development

We have described what we build as an AI software factory for a while, and for most of that time it was a phrase we had to explain. It is not any more. O'Reilly Radar has run a series on software factories through 2026, and Osmani's definition is almost exactly the thing we install: many harnessed loops running at once, fed by a queue of work and drained through a review gate into production, with humans owning the whole thing from above.

Two details in that sentence are worth slowing down for. A loop is a single agent doing one job repeatedly. A harness is the environment that constrains it: sandbox, tools, memory, gates. We did not borrow that word. Week two of our eight-week Sprint has been called the Harness since before the series ran, which is either a coincidence or a sign that anyone building this seriously arrives at the same parts.

The sharpest idea in the series is the split between a dark factory and a lit one. A dark factory ships code no human has read, verified only by other machines. A lit factory is the same pipeline with the lights left on where judgment lives. Agents still do most of the building. A human reads what comes out before it ships.

A factory that ships code nobody has read is not fast. It is borrowing, at a rate nobody quotes you up front.

Where we put the switches

human in the loopapproval gatescode review AIarchitecture decision records

Saying humans stay in the loop is cheap. Every vendor says it. The question worth asking a supplier is which loop, and what happens at each gate when the answer is no. Ours are these four.

01

Intent

What gets built, and whether it should be built at all. No agent proposes work. A product owner shapes the feature with the Product Agent and signs the acceptance criteria.

02

Shape

Architecture, data model, and the decisions that are expensive to reverse. Written down as architecture decision records before an agent writes code against them.

03

Boundary

What each agent may touch. Repositories, branches, ticket states, tools, environments. Set by an admin, never by the agent, and never inside a prompt.

04

Verdict

The merge. A senior engineer reads the diff and owns the consequence. Business logic, security and anything touching money or personal data never ship on an agent's say-so.

Notice that three of the four sit before any code exists. That is the relocation Osmani describes, and it is also the cheaper place to apply attention. A bad architectural decision caught at the merge gate has already been implemented, tested and documented by an agent that was very confident and completely wrong.

Autonomy cannot exceed verification capacity

agent autonomyback pressureverification budgettest coverage agents

Osmani calls this back pressure: the rule that autonomy cannot outrun your ability to verify. We would put it more bluntly. An agent may act without a human reading the result only where something other than a human can catch it being wrong. A type checker, a test, a policy, a migration guard. Everywhere else it drafts and a person decides.

This is why autonomy is a setting per workflow rather than a promise per vendor. The same agent can be trusted to rename a symbol across a repository and not trusted to touch a pricing rule, and the difference is not the agent's capability. It is whether your test suite would notice. Most teams discover, when they draw this up honestly, that their verification is thinner than their ambition. That gap is the real project.

The agent can ship more than you can review. Decide what that means before it does, not after.

Comprehension debt is the bill nobody quotes

comprehension debttechnical debt AIcode ownershipkey person risk

Here is the part of this argument that cuts against what we sell, so it is worth stating plainly. Osmani's best term is comprehension debt: the widening gap between the code a system contains and the code anyone on the team actually understands. Generation is cheap. Reading is not. A team can accumulate a year of code and a month of understanding.

It is a real risk and it is a risk of our model specifically. In 90+ tech due diligences the most common finding that changes an investment thesis is not a security hole. It is that one or two people understand the system and everyone else operates it on faith. Agents make that failure mode cheaper to reach.

Our answer is ownership, and it is structural rather than cultural. Each agent in a Sprint is co-built with a named Champion on the client side who owns it from week one. Week six is deliberately the hardest week, the one where an agent has to break down work it cannot do in a single pass, because that is where a team either learns to steer the system or discovers it has been watching a demo. The week-by-week sequence is built around transferring comprehension, not just capability.

We are not claiming to have solved it. Comprehension debt is a cost of this way of working, the way interest is a cost of borrowing. You manage it, you budget for it, and you notice when it grows. Anyone telling you their agents eliminate it is selling you a dark factory with the lights painted on.

Instructions are not permissions

agent permissionsagent guardrailsleast privilegeAI governance

One distinction does more work than any other, and most teams collapse it. What an agent should do is an instruction. What it may touch is a permission. Instructions live in prompts and skill files and are probabilistic: a model usually follows them. Permissions live outside the model, in tokens, scopes and branch protection, and are deterministic.

Put the boundaries you actually care about in the second category. If it matters that an agent cannot push to main, cannot read the production database and cannot move a ticket to Done, none of those belong in a prompt. We write this up for every team we work with, and it is covered in more detail in what a shared AI playbook contains.

What to ask a supplier

evaluating AI vendorsAI transformation questionsagent accountabilitybuying AI services
  1. Which decisions does a human make, and at what point? If the answer is only the pull request, judgment has not relocated. It has been deleted everywhere else.
  2. Where can an agent act unreviewed, and what catches it if it is wrong there?
  3. Who on my side owns each agent, and can they change it without calling you?
  4. What is the record of what entered the system and what the agents did with it?
  5. What happens to our understanding of our own codebase over the next year?

The last one is the question almost nobody asks, and it is the one that determines whether you own a factory or rent an operating model.

Frequently asked questions

software factory FAQcomprehension debtagent autonomyhuman in the loop
What is a software factory?+

A repeatable loop around software work: a queue of tasks, agents building inside a constrained harness, and a review gate before production, with humans owning the whole thing. The term is now in general use, including across O'Reilly Radar's 2026 series.

What is comprehension debt?+

The widening gap between the code a system contains and the code any human on the team understands. It is created when generation outpaces reading, it does not show up in velocity metrics, and it is paid back during the first incident nobody can diagnose.

Does using agents mean humans stop reviewing code?+

No. A senior engineer reads and owns every merge. What changes is that judgment also moves earlier, into intent and architecture, where it is cheaper to apply. Removing the review gate is what turns a factory into a liability.

How much autonomy should an agent have?+

Autonomy should be a per-workflow setting rather than a global promise. The rule: autonomy cannot exceed verification capacity. An agent may act unreviewed only where a test, a type check or a policy can catch it being wrong. Everywhere else it drafts and a human decides.

Who owns an agent after the transformation ends?+

A named person on your side, which we call a Champion, co-builds each agent and owns it from the first week. If nobody in the company can explain what an agent does and change it without calling us, the transformation has not happened yet.

Above The Clouds runs the AI Factory Assessment and the eight-week AI Transformation Sprint. If you are about to hand a codebase to agents, the decision worth making first is where your humans stand, not which model you buy. Get in touch to discuss your company or portfolio.