DEV Community

Cover image for OpenAI's Software Factory Can Skip Human Review. Who Evaluates That Decision?
Sara Mo
Sara Mo

Posted on

OpenAI's Software Factory Can Skip Human Review. Who Evaluates That Decision?

A pull request is green.

The specialist reviewers found nothing that should block it. The risk classifier marks the change as low risk. The human-review branch disappears, and the change continues toward production.

Every component may have done exactly what it was asked to do.

The remaining question is whether the system was right to stop asking a human.

That question jumped out at me in Gergely Orosz's diagram of OpenAI's agentic software factory.

diagram of OpenAI's agentic software factory

The public workflow includes code-writing agents, CI, specialist agent reviews, risk classification, deployment agents, production monitoring, and feedback loops. The article says that areas of the codebase can opt into automatic approval for low-risk pull requests, while higher-risk changes can receive stricter review.

I have not tested OpenAI's system. I am looking at the evaluation problem the design raises.

The small diamond marked "Low-risk change?" is not merely organizing work. Its answer can determine whether human review remains in the release path.

That makes it a release decision.

CI and Risk Classification Prove Different Things

A green CI run can provide strong evidence about the change under the conditions the pipeline checked.

The code built. The selected tests passed. The linters accepted it. A performance harness may have found no unacceptable regression. Specialist review agents may have found no issue within their assigned domains.

All of that matters.

None of it automatically proves that the change belonged in the low-risk route.

The classifier is making a different claim. It is saying that the available evidence, affected surface, expected consequence, and uncertainty are compatible with less human scrutiny.

That claim needs evidence of its own.

Otherwise, a team can evaluate every component around the gate while leaving the routing decision itself mostly assumed. The code is tested. The review agents are monitored. The deployment agent watches production. But the decision that removed the human is treated as a label rather than behavior that can be right or wrong.

"Low Risk" Depends on Context

Risk is not a permanent property attached to a file or type of change.

A change that was low risk yesterday may stop being low risk after a dependency changes, a permission boundary moves, a feature becomes widely used, or a once-correct operational rule is reversed.

The code can still pass the same local checks.

The route can still be wrong.

That is why evaluating the gate requires more than collecting examples of clean deployments. Easy successes mostly confirm that obvious low-risk changes can pass through a low-risk path. The harder cases sit near the boundary.

For example:

  • A small change touches a component whose blast radius recently expanded.
  • A test suite still enforces an old constraint after the product requirement changed.
  • A deployment agent watches a healthy proxy metric while the required production state quietly diverges.
  • Several specialist reviewers pass the change because the relevant failure exists between their domains rather than inside one of them.

These are challenge cases, not claims about observed OpenAI failures. Their purpose is to ask whether the classification still holds when the surrounding conditions become less convenient.

Score the Decision Against What Happened Next

The gate should not be judged only by whether the pull request was green when it arrived.

It should be judged against the consequence of the route it selected.

Did the supposedly low-risk change preserve the required production state?

Did it trigger a rollback, incident, manual intervention, or delayed repair?

Did monitoring capture the condition that actually mattered, or only the signals the deployment agent expected to matter?

Would an informed reviewer, given the same evidence available at classification time, have selected the same route?

This does not mean every negative production event proves the classifier was wrong. Systems fail for many reasons, and hindsight can make risk look more obvious than it was. The evaluation has to preserve what was knowable when the decision was made.

But it also cannot stop at the classifier's own explanation. A confident risk label is still a claim produced by the system being evaluated.

The useful answer comes from connecting the decision to independent evidence about the resulting state.

The Gate Needs a Record Someone Else Can Inspect

If a team wants to defend an automatic approval later, the record needs to make the decision reconstructable.

At minimum, I would want to know:

  • What exact change was classified?
  • Which risk policy and version were applied?
  • What evidence did the classifier receive?
  • Which specialist reviews ran, and what did they cover?
  • What route did the gate select?
  • Was approval human, automatic, or mixed?
  • What was actually deployed?
  • Which production postconditions were checked afterward?
  • Did rollback, escalation, or repair become necessary?

That record turns "the system marked it low risk" into something another person can examine.

It also creates the foundation for replay. When a policy, dependency, or requirement changes, the team can rerun earlier decisions and ask whether changes previously sent down the easy path would still belong there.

Without that evidence, the gate can report that it made the correct decision, but the organization cannot show why the decision was defensible.

The Nugalaxy Evaluation Cases examine this broader gap between visible success and the behavior the evidence can actually support.

More Automation Makes the Gate More Important

An agentic software factory can make code creation, review, deployment, monitoring, and repair much faster.

That speed changes the role of the risk classifier. It is no longer sorting a small queue for convenience. It is controlling where scarce human judgment enters a high-volume system.

False positives have a cost. If the gate sends harmless changes to humans too often, the bottleneck returns and reviewers learn to distrust the alerts.

False negatives have a different cost. The system removes review precisely where a human might have challenged the assumptions shared by the surrounding agents.

Both errors matter, but they are not equally consequential in every part of the codebase. The evaluation has to reflect the real cost of the wrong route, not merely report overall classification accuracy.

This is where the harness contract matters. The system needs to capture the decision inputs, expected route, executed route, and production outcome in a form that can be compared consistently. Read more about harness engineering here.

Test the Gate, Not Only the Code

The public diagram shows an ambitious feedback system. Code is written, checked, reviewed, classified, deployed, observed, and repaired through connected agentic loops.

The more capable that factory becomes, the less useful it is to ask only whether each pull request passed CI.

The factory is also deciding when its own evidence is sufficient to proceed without a person.

That decision should be evaluated as seriously as the code it lets through.

A green build can show that the change passed the checks.

It cannot, by itself, show that removing the human was safe.

Top comments (16)

Collapse
 
nark3d profile image
Adam Lewis •

I used to think the evaluation set was the thing worth defending. A green suite only ever tells you the change holds together. The question your post leaves me with is who owns the judgement that no human review is needed, and on what evidence that judgement gets revised.

Collapse
 
jo-do profile image
Jo Do •

The important metric is not simply the classifier false-negative rate. It is the counterfactual value of the human review the gate removed. I would keep a small random sample of low-risk changes in the human path, then compare discoveries, reversals, and downstream incidents against the auto-approved cohort. Without that control group, a clean deployment can look like evidence for the gate even when a reviewer would have caught a near miss that monitoring never measured.

Collapse
 
sara_mo profile image
Sara Mo •

This is a sharp way to expose what the clean deployments hide. I’d also want the sampled human reviewers blinded to the gate’s low-risk label, otherwise that label could anchor them toward approval. Then the comparison measures what review adds, not merely how often reviewers agree with the classifier.

Collapse
 
mateo_ruiz_6992b1fce47843 profile image
Mateo Ruiz •

The important distinction here is between evaluating the change and evaluating the decision to bypass evaluation. A green CI result is evidence about the code under tested conditions; it isn’t evidence that the risk classifier chose the correct review path.

I’d make the routing decision itself an auditable state transition: capture the exact diff, policy/version, evidence presented to the classifier, specialist coverage, selected route, and then connect that decision to independently observed production postconditions. At IT Path Solutions, this kind of separation is especially useful in agentic workflows because otherwise the same system ends up generating the evidence, interpreting it, and declaring that its own evidence was sufficient.

The replay idea is particularly powerful. Historical decisions can be re-evaluated against newer policies and expose boundary cases that aggregate accuracy would hide. The gate needs its own test suite, not just the code it governs.

Collapse
 
sara_mo profile image
Sara Mo •

The replay distinction matters. Re-running an old diff under today’s policy can expose missed boundary cases, but it answers a different question from whether yesterday’s gate made a defensible decision with what it knew then. I’d preserve both: an exact historical replay and a counterfactual replay under the newer policy.

Collapse
 
alikhatersaibreakroom profile image
Ali Khater •

Strong point. One additional failure mode is correlated confidence: the code-writing agent, specialist reviewers, and risk classifier may inherit the same stale policy or context, so agreement is not independent evidence. I would log evidence provenance and require a deliberately different evaluator or random human audit near the boundary. The gate should know when its reviewers are correlated.

Collapse
 
sara_mo profile image
Sara Mo •

This is a good catch. Three reviewers can still amount to one opinion if they inherit the same premise. I especially like making that correlation visible to the gate, so shared provenance near the boundary triggers an independent evaluator or human audit instead of inflating confidence.

Collapse
 
zira125 profile image
Zira •

The distinction between evaluating the change and evaluating the route is the key one. I would make the gate measurable with a weighted confusion matrix: false negatives should cost more when the affected component has a large blast radius or weak rollback path.

I would also keep the classifier inputs and policy version alongside the selected route, then join that record to post-deploy evidence such as invariant checks, rollback/escalation, and delayed incidents. That makes it possible to measure boundary cases rather than just celebrate clean low-risk passes.

One extra safeguard: when a dependency, permission boundary, or production invariant changes, freeze automatic approval for the affected slice until the gate is recalibrated. Otherwise the system can keep making internally consistent decisions against an outdated definition of risk.

Collapse
 
sara_mo profile image
Sara Mo •

The freeze after a dependency or permission change is the part I’d adopt immediately. A stored policy version tells us what the gate used, but not whether that policy was still valid at decision time. Freezing the affected slice turns recalibration into a release condition instead of cleanup after a miss.

Collapse
 
nazmul_himel profile image
Nazmul Himel •

This is a really interesting distinction. A green CI pipeline can tell us that the code passed the checks we defined, but it doesn't necessarily prove that skipping human review was the right decision.

As a frontend developer, I’ve seen how a seemingly small change can have unexpected effects once it interacts with other parts of an application. That makes the “low-risk” classification itself something worth evaluating, not just the code that was classified.

I especially liked the idea of connecting the risk decision with what actually happened after deployment. That seems much more meaningful than judging the classifier only by clean CI results.

Great perspective on an often-overlooked part of AI-assisted engineering. 👏

Collapse
 
sara_mo profile image
Sara Mo •

Frontend is a great example because a tiny diff can have a large user-facing effect that conventional CI never sees. That’s exactly why “low risk” can’t be inferred from change size alone; it has to account for the surface the change can affect.

Collapse
 
micheypico profile image
Micheal Heypico •

From operating a multi-model routing layer: the underrated variable here is provider variance over time — model behavior shifts on the vendor side without any change in your code. Testing the same prompt across several models (we keep 32 behind one key at heypico.ai) is how you tell real signal from one model's quirk.

Collapse
 
sara_mo profile image
Sara Mo •

Provider-side drift is easy to miss. Comparing models can show whether a behavior is model-specific, but detecting change over time still requires rerunning the same pinned eval and recording the model version actually used. Otherwise, 32 models give you a wider snapshot, not a baseline.

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

Your record list has one field missing, and it is the field the two upstream suggestions need in order to be answerable at all: the score the classifier produced and the cut value in force at that moment — not the route. "Was approval human, automatic, or mixed?" reconstructs the decision, but a record that holds only the route cannot support any estimator of the gate's effect, because every case below the cut looks like "a human reviewed it" regardless of how close it sat to the line.

With the score logged, there is a design that answers your own question from data the pipeline already generates. Cases just above and just below the cut differ by a hair in the score but take different routes, so comparing their downstream outcomes is close to a controlled comparison — a regression discontinuity on the routing rule. That is the only unconfounded quantity in the system: a plain comparison of incidents in the auto-approved cohort against the human-reviewed cohort is selected on risk by construction, so it measures the selection, not the gate.

The reason I'd push the score rather than a bigger sample is arithmetic. If the audited sample returns zero misses, the 95% upper bound on the mis-route rate is 3/n, so:

claim                        audits needed (0 misses)
mis-route rate <= 5%                 60
mis-route rate <= 2%                150
mis-route rate <= 1%                300
mis-route rate <= 0.1%            3,000
Enter fullscreen mode Exit fullscreen mode

and to detect a miss rate rather than bound it, at 90% power, to see a true 0.5% you need 460 audits. At 50 auto-approvals a day a uniform 1% holdback yields half an audit a day, so the 1% bound arrives in 600 days and the 0.5% detection in about 2.5 years. That is not an argument against the sampled human path — it is the argument for stratifying it: the miss rate is not spread uniformly, it lives where the score is near the cut, because that is where the classifier's own evidence was thin. The same arithmetic restricted to that stratum is 45 audits to see a 5% rate, 22 for 10%. Nine days of review versus two and a half years, for the same question, if the sample is drawn where the decision is close.

Three things I would attach to that before anyone acts on it, including as caveats on my own suggestion:

  • The score must be written before the route is taken. A score recomputed after an incident has absorbed the outcome into it, and no design recovers from that. Same for the postconditions: which production state was checked has to be fixed before the change ships, otherwise "monitoring captured the condition that mattered" is chosen in hindsight from the cases that went wrong.
  • Do not tune the cut on the window you evaluate on. RD's near-cut comparability disappears if the threshold was moved while the sample was being collected — freeze the policy version for the window, which is the discipline you already propose for the dependency/permission slice, applied to time rather than to a subsystem.
  • It estimates the effect near the cut only. That is a feature here, since your four challenge cases sit near the boundary, but it must not be quoted as the portfolio-wide rate for auto-approval.

One line on scope: I have no data from OpenAI's system and none of this tests it — it is the arithmetic the design implies, so it reads as a sizing exercise, not a finding. But it is also the second half of why the record is worth building: a record that cannot support one of these estimators has answered the audit question with a summary.

Collapse
 
sara_mo profile image
Sara Mo •

That analysis depends on the gate producing a stable scalar score and routing against a fixed cutoff, which the public design does not establish. My record list is intentionally architecture-neutral: it needs to preserve the decision inputs, policy version, and route well enough to reconstruct the decision. If a score and cutoff are part of the actual routing mechanism, they belong in that record; if not, adding them would invent a model the system may not use.

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

You're right, and the correction is exact: "the record list has one field missing" can only be true of a mechanism that produces the field. A scalar score against a cutoff is one router shape, not a property of gates, and I reached for it because it is the shape I can write an estimator for. That is the same mistake as reading a rule off an enumeration of cases — I did that version of it this week, so I won't soften this one.

Where the two halves separate, since that is the useful part:

Mechanism-dependent, conceded: (1) the field; (2) RD on the routing rule; (3) "stratify near the cut" — with no cut there is nothing to be near; (4) my "the only unconfounded quantity in the system", which is false generally and true only given a threshold on a scalar. The general form of (4) is weaker and survives: any estimate of a gate's value needs a comparison group that differs in the routing decision and nothing else. A threshold band is one way to get that group; a rule-based router gets it from cases the rules don't decide — a rule that applied and was overridden, or two rules that disagree. The requirement is the group, not the band.

Mechanism-independent, and I'd defend these two:

(a) The 3/n arithmetic. It needs no score. It is the rule of three on a binary outcome, so it sizes any audit that draws n decisions and observes zero harms: 0 misses gives a 95% upper bound of 3/n — 60 audits for a 5% bound, 300 for 1%, 3,000 for 0.1%. What it does need from the record is the population the sample was drawn from. "Route" gets you there only if every decision is recorded; if the record exists for some routes and not others, the denominator is the subset you kept, and 0-of-n over a population you can't enumerate is the number I would trust least. That is the one I'd check against your own list: does it say which decisions are enumerated, not just what is stored per decision?

(b) A record that reconstructs the decision can't evaluate it. Inputs + policy version + route is what you need to re-derive what the system did. To say whether the routing was right you need the outcome of the action that was taken — for an approval, what happened; for a denial, what the denied action would have done. That is a second record or a join key to one, and it is independent of score and cutoff. It is the difference between "the decision is reproducible" and "the decision was correct". Without it an audit measures the reproducibility of the record, and "0 harms in n" becomes a statement about decisions rather than about outcomes.

So the amendment to my earlier note: strike "one field missing", replace it with two properties — the audit population is enumerable from the record, and the record joins to the outcome. Both extend your neutral list without committing to a router, which is what I should have written the first time.

One honest note on the sizing, in case there is no score: the near-cut stratum does not exist, so the choice is between a uniform sample and strata that your rules actually produce — rule id, review depth, deviation flag, whatever the route names. Those are strata; they just aren't ordered by the classifier's confidence. The 3/n line is unchanged either way.