About this text. Preprint, Part 1 of 2 (4 October 2026). Citable version with the Russian translation: doi:10.5281/zenodo.23242704. Materials (Part 2 protocol, counting script, contribution log): keryx repository. The research and the text were produced jointly by the author and AI agents; who did what is described in the Contributions section at the end.
Part 1 of 2 · observations, hypothesis and test protocol
A practitioner's research essay: observations, a formalization of the problems, a working hypothesis and the protocol for testing it in Part 2
Author: a product development practitioner who works on team projects with a conventional division of roles and also develops the open-source product keryx with agents. The author and the agents worked jointly on the research and the article; their respective contributions are described at the end.
Abstract
In the process described here, AI agents have taken over most of the production of artifacts: requirements, acceptance criteria, code, tests, reviews and releases. Checking that the result matches the intent and actually helps has stayed with humans, and we hypothesize that its cost has dropped far less. This is a research essay by a practitioner who works in teams with a conventional division of roles and in parallel develops the open-source product keryx with agents. The main material is the history of one project (383 units of work and 3,175 acceptance criteria recorded over three months, showing how the approach changed), together with anonymised data on the work of a team with a conventional division of roles. The initial process was set up the way many are, in our experience: criteria were written and confirmed, but nothing distinguished criteria that had been checked from those merely marked as confirmed, the confirmation signature was assigned automatically, and the question "did it help?" had no place in the process. That is what made the problem visible and let us make it observable step by step: first the check type per criterion, then the origin of the request, the expected outcome, the cause of rework, and the recommendation journal.
In this paper we formalize six problems and propose a working concept of residual judgement. For a given set of tools, budget and acceptable level of risk, a residual remains for humans: criteria whose verification that set does not provide, and decisions reserved for humans on normative grounds. The unit of observation becomes a claim record: what is claimed, how it is checked, what state the check is in, who is responsible. Five mechanisms make the residual visible inside the process; four of them are already implemented in keryx. Seven predictions, with the conditions under which they would be refuted, will be tested in Part 2. Here we claim only what the records show: how the documentation of decisions and checks changed; conclusions about costs, about people's behaviour and about boundaries between roles remain hypotheses.
Keywords: agentic automation, product development, role blurring, acceptance criteria, automation bias, authorship, default effect, advice taking
Figure 1. The paper's frame. Artifact production at every stage has mostly passed to agents; the human residual of judgement is spread across stages and is invisible until recorded. Five mechanisms (M1–M5) make it observable; the unit of observation is the claim record.
1. Introduction
I work in two modes at the same time. On team projects, roles are still divided along conventional lines: there's a product owner, an analyst, developers, testers. In parallel I'm building keryx, a system for managing development with agents, and there the agents and I perform all of these roles between us.
It was the comparison between the two modes that caught my attention. In teams I see role boundaries getting less and less clear: the designer generates code, the developer writes requirements with a model, an agent does the review, and responsibility for the result matching the intent gradually stops belonging to anyone in particular. In my own project the same process is taken to its limit: I barely write anything by hand and mostly state the intent and accept the result.
keryx is not a research testbed but an open-source product I develop in my own time as part of my professional development: agent orchestration, durable project memory, frozen acceptance criteria and an evidence ledger in one repository. The research and the product grow together: an observation from the paper becomes a mechanism in the product, and the product's data are used to test the paper's claims. The examples below are therefore not illustrations but records from a working repository, identified by their flow numbers.
This paper is not an account of my experience. That experience raised the question; the answer is built from the project's records, industry data and the literature.
Research question. Which parts of product development remain for humans when formalization, checking and coordination are taken over by agents and automation, what problems arise as a result, and how can that part of the work be made visible so that it is not lost among finished artifacts?
The work has two parts. In the first we describe observations, formalize the problems, formulate a working hypothesis together with the proposed mechanisms, and set out the test protocol. In the second the hypothesis will be tested on the keryx project, where every step of work leaves a measurable trace. This is neither a survey nor a general theory: where we rely on our own observations, we say so.
2. Industry background
Briefly, because the reader has most likely seen these numbers. Production has become cheap: in our practice a specification, a task, a prototype and a first version of code take minutes rather than person-days; as we have observed, when an attempt is cheaper than the discussion about it, the discussion disappears, and the discussion before work starts goes first. Checking, judging by the available data, has not become cheap in the same way: according to LinearB (8.1M pull requests, 2026) AI-written pull requests wait 4.6 times longer before review begins, although the review itself is twice as fast once picked up; in the Sonar survey (2026, 1,149 developers) 96% did not fully agree that they were confident in the correctness of such code, and only 48% fully agreed that they always check it; pull requests commented on only by an agent reviewer are merged in 45% of cases compared with 68% for those with human review (Chowdhury et al., 2026, an observational sample). DORA (2025) reports that higher AI adoption is associated with both higher delivery throughput and higher delivery instability. Functions are increasingly combined in one person: the share of solo founders among startups on Carta grew from 31% to 36% in a year and doubled over ten years (Carta, 2026); among C corps incorporated through Stripe Atlas from the start of Q2 2026 to the end of May, it is 63% (Stripe, 2026).
All of these sources report waiting times, self-reported assessments and company composition, not effort spent or how decisions are distributed inside a team. They set the background but do not answer the paper's question. We seek the answer in our own project's records.
3. Formalizing the problems
Six problems follow from these shifts. Each rests on external data or a known mechanism; in my own project I see their manifestations, which we will measure in Part 2. The "grounds" column says where each problem comes from; it does not establish that it occurs here.
| # | Problem | Grounds |
|---|---|---|
| 1 | Checked and unchecked criteria are indistinguishable | the self-reported gap between distrust and checking (Sonar); in our artifacts, confirmation without a stated check |
| 2 | Reduced critical scrutiny across roles | in an observational study of 14 developers, 56% of LLM-related actions (456 of 808) showed at least one of the coded cognitive biases, among them adopting a suggestion without checking (Zhou et al., ICSE 2026); one mechanism is automation bias (Parasuraman and Manzey, 2010); extending this to other roles is our extrapolation |
| 3 | Loss of intent during formalization | named as a risk in spec-driven approaches (SDAD, 2026, §8) but not measured |
| 4 | Blurring of intent authorship | attribution laundering in dialogue with a model: a mechanism proposed in the essay Dead Cognitions (2026), not measured |
| 5 | Judgement drifts towards approving the recommendation | AI advice shifts human choice: a covertly biased adviser raised the odds of preferring the inferior option by a factor of 5–8 in financial and emotional scenarios (Sabour et al., PNAS 2026); in an experiment with simulated resume-screening recommendations, preference for the favoured group reached up to 90% under strong bias (Wilson et al., 2025); in keryx the recommended option was pre-selected (the conditions of the default effect, Johnson and Goldstein, 2003) and labelled (advice as an anchor, Bonaccio and Dalal, 2006); blind questions were introduced as part of this study |
| 6 | The outcome loop does not close | AI adoption is associated with growth in both releases and delivery instability (DORA, 2025); self-assessed speed-up diverges from measured time (METR, 2025); most ideas do not improve the chosen metric, in experimental data (Kohavi et al., 2009) and in practitioner observation (Cagan, 2013) alike; our process had no place for the question "did it help?" |
Table 1. Six problems and their grounds (sources and rationale).
3.1 What it looks like in the project records
The keryx repository as of 4 October 2026 contains 383 units of work (flows) created since 9 July and 3,175 acceptance criteria, of which 3,062 are confirmed; 159 of the flows have undergone 368 rounds of agent review. Each unit leaves a description, frozen criteria, a journal and review records, so the problems in the table can be shown rather than described. The repository is public; the snapshot, the counting script and the other materials for this part live in the research directory of the keryx repository. The initial approach was not exemplary, and we do not present it as such: it was ordinary, and that is exactly why it shows what becomes observable once judgement has a place in the process.
Problem 1, checked and unchecked criteria are indistinguishable. The description of flow 379 (28 September) opens: "A confirmed-but-unverifiable acceptance criterion and a confirmed-and-tested one are rendered, gated and signed identically. Roughly a quarter of the 2,705 frozen criteria cannot be verified in principle, and nothing distinguishes them" (2,705 is the count frozen by 28 September; by 4 October it is 3,175). "The habit of claiming verification emerged unprompted; the habit of naming the check did not, because no field exists for it." A check-type field was introduced in flow 379. The process does not require it, and at the snapshot 166 of the 223 criteria written since then carry one, about three quarters; six units of work never filled it in. Earlier criteria remain the baseline. Flow 397 shows why the field is needed: "the call sites are confirmed by reading only. Deleting one call leaves every test green."
Problem 2, reduced critical scrutiny. Over 368 rounds the agent review produced 2,544 findings, about seven per round, 175 of them marked as blocking. We did not measure whether a human reads such a stream in full; we only see that the process answered not with "read more carefully" but with a separate re-verification step and a recorded verdict. The full breakdown: 1,967 findings carry a verdict and 577 do not; among the verdicts, 329 are "confirmed" (the problem was reproduced and is still open), 1,623 "refuted" and 15 "unverifiable". The label "refuted" means the problem can no longer be reproduced when checked again: in 1,546 cases a fix is recorded alongside, in 49 the finding is recorded as wrong, and in 28 no decision is recorded. A fix confirms the state after the change, not the truth of the original finding, and we do not conflate the two. By these records, agent review in our project does not look predominantly like noise; what is visible is something else: without a recorded re-check, "fixed" and "false positive" would have been indistinguishable.
Problem 3, loss of intent. Flow 396 was born from my question: "Why such restrictions? I said Telegram should be an interface." The agent imported the shell's default restrictions into the channel without formally violating any requirement. Flow 376, criterion AC5: the review found that the criterion's text and the behaviour disagreed, and the decision was mine: "AC wording follows the behavior". The criterion was fitted to the code rather than the code to the criterion, and the journal of that decision is the only thing that tells a deliberate choice from a capitulation.
Problem 4, blurring of authorship. For a long time the signature under a criterion confirmation was taken from git config by default, and the journal could not say whether a human had confirmed the criterion or an agent working under that person's account. The signature now distinguishes an explicitly named confirmer from one whose identity was inferred automatically, and since flow 387 the expected outcome has an author: since then the agent has been the author in 18 cases and the human in 11. This does not make the earlier confirmations false, but it makes visible whose they are.
Problem 5, approval instead of decision. Short confirmations such as "390, ok" or "1 accepted" in reply to the agent's report are normal practice where there is trust, but the journal did not distinguish these brief approvals from decisions made after reflection. Flow 381 says it plainly: "design decisions of mine, overrulable" (verbatim from the journal), that is, the agent decides first, the human approves after the fact. This is where the recommendation journal came from: in keryx the recommended option was always marked and pre-selected, and we hypothesize that this presentation makes it easier to miss something important while agreeing with the ready-made option than when choosing among equal ones; this is a hypothesis, not a measured effect. The journal's first numbers for 2–4 October: 47 decisions, 33 of them with a recommendation; in the ordinary mode the choice matched it in 16 cases out of 22, in the blind mode in 6 out of 11. This is a description of a small series by one operator, in which the ordinary and blind modes differ in several factors at once, and no conclusion is drawn from it yet. What matters more: of eleven deviations, a reason was given for two. The journal sees that the human chose differently, but almost never why. Retrospectively reconstructed records of polls conducted before the journal was introduced give a baseline match rate of about 77% across 53 decisions with a recommendation, but they were not recorded at the moment of answering and are not directly comparable with the new ones.
Problem 6, the outcome loop. Until late September the process had no place to record whether what was built had helped. Flow 380 is the first run of the tool that creates that place; its first report on the repository says candidly: "It says the instrument did not exist, not that the intents failed to work." From that point every new unit of work has a field for the expected effect and how to observe it, and everything built before remains the baseline for P5.
3.2 What it looks like in a team
The second source is an anonymised dataset covering a year of work by a team with a real division of roles, October 2025 to September 2026: the working repositories (PRs) and the task board. An agent collected the data using a brief written in advance; only numbers were included in the report: no names, no texts, no account identifiers. Participants are split only into "human" and "agent". The team has two kinds of agent that differ in how easily they can be identified: a bot using a GitHub bot account is identified by that account; an agent such as Claude or Codex, working on a person's behalf under that person's account (the same agents as in keryx), is recognised only indirectly: through co-authorship markers in commits, templates in descriptions, and "heavy formatting", a structuredness heuristic that classifies a description as structured when at least three of the following markers are present: length of 250 words or more, headings, lists, bold text, tables, emoji, code blocks. The second kind is now the main one, and it is harder to distinguish from a human contributor. We validated the heuristic on two classes with known authorship: 2,338 PR descriptions from the period before agents appeared in the repositories (February 2024 – July 2025, human by construction) and 1,689 agent-era descriptions carrying a formal marker. On human descriptions the heuristic fires in 0.3% of cases (never above 0.5% in any half-year; the median human description is 12 words), on marked agent descriptions in 54% (62% on those with a text marker). A structured description without a marker is therefore almost never human, but the heuristic misses about half of agent descriptions; below it is a lower bound on unsigned agent authorship, not an upper one. The limitation: the human class comes from an earlier period, and if people have since started writing more structured descriptions on their own, the false-positive rate is now higher than 0.3%; the PR template appeared in one repository in April 2026 and does not affect that class.
There is one turning point, and it is visible in the repositories and on the board at once, in June–July 2026. Before that point, the median code-change description contained only a few words and the median task description sixty to seventy; afterwards, both were two to three hundred words long, and the number of changes per month increased by a factor of 2.5. Structured descriptions without a formal marker grew from a few percent to a third; corrected for the heuristic's sensitivity, this means that in June–September 2026 roughly nine in ten PR descriptions were written by an agent (the naive estimate from markers and structure is seven in ten), and most of them carry no signature. This is problem 4 in its observable form. The correction assumes that unmarked agent descriptions look like marked ones; in one of the repositories unmarked descriptions are more structured than marked ones, and there the estimate hits 100% and is uninformative. Acceptance criteria appeared in tasks: from about two percent to thirteen; a check is named for four criteria in ten, and for the rest it is not recorded, so the records cannot say how they were checked. This is problem 1 in its observable form. The share of changes without a recorded review grew from a third to a half; at the same time the absolute number of changes with a recorded review grew from roughly 190 to 300 per month, so the lag concerns coverage, not the volume of checking. Approval without a single comment remained the norm, more than half of approvals, though the share fell over the year; an approval without a comment does not mean there was no review, it means no trace of it was recorded. Within the team the response to the stream was uneven: in some places recorded findings increased, in others they did not. Links between changes began to be recorded: the share of changes followed within two weeks by a related follow-up grew from two or three percent to a fifth; whether there is more rework, these figures do not say. The effect section stayed empty almost everywhere: a non-empty "effect" block exists in a few percent of changes and tasks, and the share of tasks closed with a substantive comment did not change over the year. Whether outcomes were assessed outside these records, in analytics or in conversation, the data cannot show; what they show is that the process has no place for it, and that is problem 6.
| Measure | Working repositories (PRs) | Board (tasks) |
|---|---|---|
| Items per month | ≈280 → ≈690 | ≈250 → ≈340 |
| Median description length, words | 6 → 315 | 68 → 249 |
| Empty description | 45% → 12% | 7% → 3% |
| Formal agent marker | 19% → 43% | 0% → 2% |
| Structured description without a marker | 8% → 28% | 5% → 34% |
| Tasks with acceptance criteria | — | 2% → 13% |
| Criteria with a named check | — | 16% → 39% |
| Changes without a recorded review | 32% → 56% | — |
| Approval without a comment | 80% → 62% | — |
| Linked follow-up within 14 days | 3% → 23% | — |
| Non-empty "effect" section | 1% → 2% | 0% → 5% |
Table 2. Team data: October 2025 – May 2026 versus June – September 2026. Shares are weighted by the number of changes or tasks in the period; length is the median of monthly medians. "Agent" in the formal marker means a bot account or a co-authorship trailer in a commit. Structuredness heuristic: 0.3% false positives on 2,338 pre-agent descriptions, 54% sensitivity on 1,689 marked ones.
Figure 2. Team data by month, October 2025 – October 2026. The June 2026 turning point is visible in all four series; the dotted fill marks the period after it, the last point is a partial month.
What these data show and what they do not. They show that the change in records is not a peculiarity of one project: documented output grew severalfold, the share of changes with a recorded review fell, most agent-written descriptions carry no signature, criteria appeared before recorded ways of checking them, and there is no place to record an outcome. These are measures of documentation, not of behaviour: they measure neither effort, nor attention, nor the quality of decisions. Nor do they show causes: we know first-hand that the turning point coincided with team members starting to use agents under their own accounts, but in the data it is one event without a control; the structuredness heuristic was validated on the pre-agent era, not on current human descriptions. A breakdown by role, needed to examine the blurring that motivated this analysis, requires a role map and remains the next step.
4. Three processes behind the six problems
The shift of cost from production to verification is a hypothesis. Industry data support it indirectly (review queues, self-reported distrust), and the records show only its documented side: in keryx, before the field was introduced no criterion specified a check type, and about three in four of those added since do; in the team, documented output grew severalfold while the share of changes with a recorded review fell from two thirds to a half even as their absolute number grew. Neither source measures active human time, the cost of agents or the cost of rework; without that, "the cost has shifted" remains an assumption. If the shift is real, the bottleneck becomes the verifier's attention rather than the speed of production.
The blurring of roles is the author's observation: in teams the product owner, analyst, developer and tester share functions among themselves and with agents; in keryx they all converge in one person. Describing work through roles no longer explains who is responsible for what. The records show only part of this: in our project, confirmation signatures that until September did not distinguish a human from an agent using that person's account; in the team, a third of descriptions in a structured format without any signature, which after correcting for the heuristic's misses means a majority of descriptions. The transfer of functions, authority and responsibility between people is not measured by these data.
Distortion of judgement. Three mechanisms, two known and one proposed, may operate under the same conditions, and each calls for its own intervention: automation bias, that is, adopting a suggestion without checking (in Zhou et al., 2026, 56% of LLM-related actions showed at least one of the coded biases), countered by a separate re-verification step with a recorded verdict; the default effect and advice as an anchor, that is, pre-selection and labelling of the recommended option (Johnson and Goldstein, 2003; Bonaccio and Dalal, 2006; in Sabour et al., 2026, covertly biased advice raised the odds of the worse choice by a factor of 5–8), countered by separate arms in the recommendation journal; attribution laundering (Dead Cognitions, 2026, unmeasured), countered by recording the origin of work. We do not show whether they reinforce one another, nor whether their magnitudes carry over to a development interface. None of the development tools we know shows these effects to the human at the moment of decision; that visibility is what we build in keryx.
5. Hypothesis: the concept of residual judgement
Functions observed through stage-specific criteria. Instead of roles, the concept looks at functions (Jesuthasan and Boudreau, 2022, make the same move from jobs to the tasks that make up work), and each function is observed through a criterion at its stage: intent ("what should exist and for whom"), architecture and design ("how it is structured"), formalization ("how to write it down"), acceptance ("what counts as done"), checking ("what confirms it"), effect ("did it help"). Human judgement is not tied to one stage: it may be needed when checking code as well as when reading a metric.
Residual. The residual is defined not in general but relative to conditions: the set of models and tools, the budget for checking, and the acceptable level of risk. Under given conditions it has three parts, which we account for separately: the technical residual, criteria whose verification that set does not provide (not yet automated, or automation unreliable); the economic residual, criteria whose verification could be automated but is uneconomic within the budget; the normative residual, decisions reserved for humans by rule, even where a machine could make them. The unknown is counted separately: criteria whose check type is not recorded. Operationally, the residual is measured as the share of criteria labelled "judgement" or "not checked" among units of work created after the label was introduced, with the original wording of a criterion preserved across rewordings. We do not claim the residual is irreducible in principle. We hypothesize that, with the conditions held fixed, the technical residual does not shrink below some level through more careful wording of criteria alone; the normative residual by definition does not depend on wording and is not part of that test.
Claim record. Every claim about the product is described not by one status but by several independent fields: what is claimed; the check type (a command, an automated test, human judgement, or "not checked" with a stated reason); the progress of the check (not started, machine-generated evidence available, judged by a human; the latter two can coexist) and separately its result (confirmed, refuted, data contradictory); the limits of the evidence (what exactly was checked, in which environment, against which version of the claim; when the meaning of a claim changes, the applicability of earlier evidence is decided afresh rather than inherited); who is responsible for the decision. We deliberately do not write "proven by a machine": a test confirms behaviour under given conditions, and if the agent wrote both the code and the tests from the same misunderstanding, the resulting shared error passes the checks undetected. The limits of the evidence therefore note separately whether the test oracle is independent of the generation of the implementation. The unit of analysis becomes not a position or a function, but a claim record.
Managing the agent. Once the human steps back from execution, the human manages the agent as well as the product: deciding what value to pursue and in which direction remains part of product management. An agent's manageability has several components. The first is what we call addressability of the residual: a property of the agent's representation of its work that makes clear what it understood from the request, what it promises, what has been checked, and where a human is needed, that is, exactly where human judgement is required. The others are authority (what the agent is allowed to do), the ability to intervene and stop the agent, and the reversibility of the result. This work deals with the first component and makes no claim about the others.
Recommendation as a channel of blurring. The share of decisions in which the human departed from the agent's recommendation does not by itself measure the quality of judgement: the better the recommendations, the fewer reasons to depart. The deviation rate therefore remains a descriptive measure; reasons are collected symmetrically, for agreement as well as for deviation, on a fixed subsample, and telling "accepted good advice" from "accepted the default" needs an independent, blinded assessment of the quality of both the recommendation and the choice. It is planned as the next addition to the measurement framework (flow 400) and is part of the Part 2 protocol.
Figure 3. One status versus a claim record, on criterion AC5 of flow 376. The test is green in both schemes; the difference lies in the record: who judged, by what means, under which limits, and who is responsible.
Example claim record · flow 376, criterion AC5 (criterion text verbatim from the repository)
-
claim: before review: "a second
keryx serveon the same bot token does not poll and exits with the reason"; after: "…meets a 409, stops polling for good, says why in its log and status, and keeps serving its other routes" -
check type: test
single-poller.test.tsfor the polling behaviour; but the review (finding F-011, agent reviewer) showed the test asserts only the poller state, the code never "exits" as the criterion says, and whether to exit or keep running is a product decision that requires human judgement - state: judged by a human: operator decision (poll 25, 1 October): "AC wording follows the behavior"
- limits: the criterion was reworded to match the actual behaviour, rather than the behaviour being changed to match the criterion; test and implementation were written by the same agent; after the rewording all five criteria were reconfirmed automatically, "evidence unchanged"
- responsible: the operator; the confirmation signature was derived automatically, the explicit confirmer field came later
Closest related work and our contribution. Bainbridge (1983) described the ironies of automation: the human is left with the tasks automation did not take, and those are the hardest. Simkute et al. (2024) carried the argument over to generative models: the user shifts from producing to evaluating the result. SDAD (2026, §8) discusses role transformation, intent verification and human release authorization in spec-driven development. Dead Cognitions (2026) proposes the mechanism of attribution laundering. The closest technical predecessors of the claim record are the SACM structured-assurance standard (OMG, 2023), which separates claim, argument, context and evidence, and the W3C PROV-DM provenance model (2013), which distinguishes activity, attribution and delegation; our record borrows that separation and adds the state of the check and the independence of the oracle from generation. The recommendation journal is a particular case of cognitive-forcing interventions in work with AI: Buçinca et al. (2021) showed that such interventions reduce overreliance but carry a cost for the user; we measure that cost separately. On the team side, the closest material is the field experiment of Dell'Acqua et al. (2025) on the boundaries of expertise in product tasks with AI; it does not study lasting changes of authority in agentic development. We did not conduct a systematic review and do not claim the model as a whole is new. Our contribution is narrower and more concrete: a claim-record schema with independent check fields, a recommendation journal designed to distinguish the effects of the label, the order and pre-selection (the arms themselves are the next step), and a protocol for testing seven predictions with explicit conditions under which they would be refuted, built into one working agentic process where they can be tested.
6. Proposed solution
The solution introduces no prohibitions. This is a design premise, not a result: we assume that people find ways around imposed rules, and we test the premise through the completeness of records and the load on the user. Where possible, the mechanisms record data automatically; whatever requires a human action (naming the reason for an answer, say) is optional in routine use, although the research protocol may require it on a subsample. The mechanisms are intended only to make the residual visible and to show the human which questions need their attention. The cost of running such a process is measured separately.
The mechanisms were not derived from the list of problems; they came out of observations, in both directions. The check-type field (M1) arose from the observation that a quarter of the criteria could not be verified and nothing distinguished them; why it was needed was shown by the very next case, flow 397, where deleting a call left every test green. The recommendation journal (M5) grew out of short replies such as "390, ok" that were indistinguishable from decisions. Problem 2 is not closed by a mechanism: it is addressed by the re-verification step, which existed in keryx before this work. Below, for each mechanism, we state what is implemented, what is being piloted and what is only planned.
-
M1. A check type for each criterion. A command; an automated test, including a test asserting the absence of unwanted behaviour (that the old path is no longer called, say); human judgement; or "not checked" with a reason. The limits of the evidence are recorded alongside the result. Criteria that cannot be checked no longer look like criteria that have been checked. In keryx, implemented: the
[verify: exec|invariant|judged|none]tag since flow 379 (28 September); the tag is optional, and at the snapshot about three in four new criteria carry it; those written before it are deliberately left untagged as the baseline. Planned: making the tag mandatory when criteria are frozen. Planned: a field for the limits of the evidence and the version of the claim. - M2. Origin of work. The record distinguishes the author of the message, whoever first proposed the idea, whoever substantially transformed it, and whoever made the decision; mixed or unknown origin is allowed. The "human request" label is set only together with an exact quote of the original message, and the original wording is kept next to the agent's. A quote shows the origin of a message, not of a thought: the human may have repeated an earlier proposal by the agent, and that limitation is acknowledged. In keryx, implemented: the origin field since flow 390; nine units of work carry "human request" with a verbatim quote, one carries "agent finding"; for five units the origin was backfilled by operator decision on 2 October, and this is recorded in the history. The field records the origin of the record, not of the thought: what the person remembers about where a proposal came from, it does not measure.
- M3. Outcome recorded alongside the release. The human's original expectation in their own words, the agent's interpretation confirmed by the human, a metric, an observation window, an owner and the actual result, with the option to record "not measured" and a reason. "Not measured" is counted as its own category in the report so that it does not become a way to mark the question as resolved without assessing the outcome. In keryx, partly implemented: the template "request verbatim / effect as formalized by the agent / how to observe" since flow 390, narrower than the full schema above; the observation window, the owner and the actual result as separate fields are planned. What was built before the template remains the baseline for P5, where the absence of the field counts as unknown, not as zero assessments.
- M4. Rework as a signal. A follow-up change to a feature is a signal observable without asking anyone, but not a diagnosis: the cause is coded separately (new knowledge, implementation defect, integration constraint, new requirement, lost intent, deliberate iteration), and several causes are allowed. The absence of rework does not mean the intent was preserved. In keryx, planned: the titles of 59 units of work identify them as fixes, remaining work or follow-ups to earlier units; detecting rework beyond titles and the cause-coding scheme are introduced in Part 2.
- M5. Recommendation journal. For every decision with options, the journal records the agent's recommendation, the order of options, whether an option was pre-selected, the human's choice, time to answer and, when the human chooses a different option, an optional reason. In keryx, pilot: running since flow 392 (2 October). A third of questions are asked blind (the recommendation label is hidden, the order is shuffled, nothing is pre-selected, the recommendation is revealed after the answer), and the report is built without a model. At this stage the interface package as a whole is compared, not the separate influences. Questions about irreversible actions currently always go to the ordinary mode, which makes the groups unequal in risk; in the next version they are excluded from both comparison groups and described separately. Planned (flow 400): deterministic assignment of arms, a field recording the rationale for the recommendation, the remaining interfaces, separate experimental arms for the label, the order and pre-selection, and an independent assessment of recommendation quality.
Figure 4. How the same question is shown in the four arms of the recommendation journal, and what is written before it is shown and after the answer. A and D are running now; B and C separate the contributions of the label, the order and pre-selection.
7. Testable predictions
The test will be carried out in Part 2 on the keryx project: each unit of work in it (a flow) leaves in the repository a description, frozen criteria, a journal and review results, so the work can be reconstructed from artifacts. The data as of 4 October 2026: 383 units of work since 9 July and 3,175 acceptance criteria, 3,062 confirmed; the baseline was fixed on 28 September, its records are kept unchanged, and 44 units of work have been created since. The observations are nested in one project by one author and are interdependent by time and task type, so 3,175 criteria are not 3,175 independent observations; the unit of analysis will be the unit of work or the randomly assigned question, not the artifact line. The recommendation journal (M5) has been running since 2 October; its first three days contain 33 decisions with a recommendation (16 of 22 matches in the ordinary mode, 6 of 11 in the blind mode). These records are a pilot used in designing the protocol; they will not enter the confirmatory sample.
The protocol (definitions, inclusion units, denominators, windows, thresholds, stopping rules, assignment of conditions, handling of missing data and the permitted third outcome "insufficient data") was fixed on 4 October 2026 and lives in the same directory of the keryx repository as a dated document with a version log; the confirmatory sample starts with units of work created after its publication. The main parameters: P1 is refuted if at least 30% of human decisions fall at the formalization and checking stages; P2, if after rewording with independently confirmed equivalence of meaning the technical residual falls below 10% of criteria; P3 is rated on a three-level rubric (cannot tell how to check / method clear but not reproducible / reproducible from the text) with the label and the period hidden; P4 uses one primary cause per rework item, a tie counts as refutation, rework is detected beyond titles, with a 30-day window; P5 counts an outcome as assessed if, within 14 days, an observation on a pre-named metric and a verdict are recorded, with "not measured" as a separate category; P6 codes avoidability from information available before the work began; P7 uses a minimum meaningful difference of 15 percentage points, an equivalence margin of ±10 points and about 150 randomly assigned reversible questions per arm (α = 0.05, power 0.8). Each prediction allows three outcomes (supported, refuted, insufficient data), and the absence of a significant difference is not treated as equivalence; intervals are computed with clustering by unit of work, and at least 20% of the records in each set are coded blind by an independent rater. Part 2 will be published when the journal and the rework coding yield enough records for these thresholds.
| # | Prediction | Refuted if |
|---|---|---|
| P1 | Human decisions are concentrated at the intent and acceptance stages | the share of all human decisions that fall at the formalization and checking stages (denominator: all human decisions coded from artifacts across all six stages) is not below a pre-set threshold. A hypothesis about this project, not a test of the residual's existence |
| P2 | Under fixed tools the technical residual is substantial and does not shrink from rewording alone | with the same tools and budget the share of criteria labelled "judgement" or "not checked" falls below a pre-set minimum after a rewording that preserves the original commitment (equivalence of meaning assessed independently); the normative residual is counted separately and is not part of the test |
| P3 | A visible check type raises the checkability of criteria | independent ratings of checkability on a pre-set rubric do not differ between criteria with and without a check-type label, with the label and the creation period hidden from the rater |
| P4 | Lost intent is the most frequent cause of rework | under blind coding with the full taxonomy (one primary cause per rework item; rework detected beyond titles) at least one other cause is as frequent as, or more frequent than, loss of intent, allowing for the uncertainty interval. A hypothesis about this project |
| P5 | M3 increases the proportion of releases whose outcomes are assessed | the proportion of releases whose outcome is assessed, by a pre-set sufficiency criterion and within the window, is not higher after M3 than in the baseline; assessment is established the same way in both periods, the absence of the field before M3 counts as unknown rather than zero; the proportion assessed is not the proportion that helped |
| P6 | Making the origin of work visible reduces the proportion of avoidable rework | in comparable units of work with the same observation window the proportion of avoidable rework (avoidability coded from information available before the rework) is not lower after M2; the effect of M2 is not identified separately from the other mechanisms, the package is assessed |
| P7 | Presenting options in blind mode lowers agreement with the agent's recommendation | on randomly assigned reversible questions the difference in match rates between the blind and ordinary modes lies within the ±10-point equivalence margin; at the first stage the interface package is assessed, with the label, the order and pre-selection assessed in separate experimental arms later; the quality of the decision is assessed separately from agreement |
Table 3. Predictions and the conditions under which they would be refuted. The conditions for P2, P3 and P5 were recorded before the first measurements; the rest were formulated while preparing Part 1; all will be tested on work created after the protocol is published. Support for P3, P5, P6 and P7 provides evidence of the mechanisms' usefulness; P1, P2 and P4 are hypotheses about this project, and their failure limits the framework without invalidating the definition of the residual.
8. Limitations: what cannot be concluded
- About behaviour: only from records. Every measure in this paper measures documentation: whether a review is recorded, a check type stated, an outcome field filled in, a reason given. They do not measure whether a person read, whether they thought, or whether an outcome was assessed outside the record. The sensitivity and specificity of these traces are unknown; on a small subsample they will be compared with observation and with interviews about specific episodes.
- About costs: no data. Active human time, the cost of agents and the cost of rework are measured neither in keryx nor in the team. The shift of cost to verification remains a hypothesis.
- About roles: observation only. The team data distinguish human from agent, not roles, and identify an agent's contribution under a human account only indirectly; the structuredness heuristic was validated on pre-agent descriptions (0.3% false positives) and on marked agent descriptions (54% sensitivity), but not on current human descriptions, so its estimate of agent authorship is a lower bound with a caveat. We found no empirical studies that directly measure how role boundaries within a team change in AI-assisted development; the closest field experiment (Dell'Acqua et al., 2025) studies the boundaries of expertise, not authority. The conclusion about blurring between roles rests on the author's observations.
- About the mechanisms: only for an individual agentic process. The test is carried out on a single project where the author is at once the researcher, the only participant, the operator who confirms the criteria, and the owner of the tool that measures them; some records were backfilled by the author's decision, and this is marked. A solo project is a limiting case of the transfer of functions to agents, not of team role blurring. Part 2's findings about the mechanisms concern this process; its claims about teams are limited to what the aggregates in section 3.2 can show.
- Agents measure themselves. Agents proposed the metrics, collected the team data and may code their own work; the reviews of the drafts were also produced by models. To limit this, at least 20% of the records in each set are coded by an independent rater under blinded conditions, without seeing the label, the period or the author, and disagreements are resolved by a third rater. Inter-rater agreement is reported, coding schemes are published before analysis, and the author is responsible for checking the agents' computations.
- Before/after does not identify a single mechanism. M1–M3 were introduced almost at once, together with the author's learning and model changes; the absence of a field in the past is not a zero. The package of changes is therefore assessed as a whole, the historical period is coded the same way as the new one, and the unknown stays unknown.
- The literature review was not systematic: we did not define a search method and cannot rule out relevant work beyond that discussed in section 5. Some industry data come from tool vendors and report waiting times or self-reported assessments rather than costs; they are cited as context.
9. Conclusion
The records of two processes, one individual and one team, show the same thing: the production of artifacts grew severalfold while the process had no place for human judgement, so a checked criterion was indistinguishable from a claimed one, a signature did not distinguish a human from an agent, and there was nowhere to record whether it helped. We hypothesize that behind this lies a shift of cost from production to verification, a blurring of roles and the joint action of known cognitive mechanisms, and that the consequence is a human who decides less and approves more; but these are hypotheses, and this part does not test them. What has been shown is that the residual can be made observable without new prohibitions: the check type per criterion, the origin of work, the expected outcome, the cause of rework and the recommendation journal already change the records of keryx, and for the first time those records show exactly where judgement was required. The concept of residual judgement is a working measurement framework, not a proven theory. Part 2 will show which of the seven predictions hold up under a protocol published in advance, and which of them speak to the usefulness of the mechanisms and which to the framework itself.
Contributions
The research and the article are the joint work of the author and the agents. The author contributed the problem statement and its framing: the observation of role blurring on team projects, the division into theoretical and practical parts, looking at functions instead of roles, the principle "intent from the human, formalization from the agent", the principle of managing the agent, the role of the recommendation, architecture and design as a separate stage, and corrections, each of which changed the direction of the work. The agents reviewed the sources, collected and processed the data, proposed the measurement framework and the wording of the concept, and prepared the drafts. External reviews of the drafts were produced at the author's request by models: GPT-6 Astra (drafts 2 and 5: an independent scientific review, a linguistic review of the English version and a comparison of the drafts, with sources checked against the originals) and MiniMax 3 (draft 5). Their quality differed: the Astra reviews found errors in how results had been carried over from the sources, which we accepted; the MiniMax review, alongside useful comments on the protocol, also contained mistaken ones (a misreading of the denominator in Zhou et al., a claim that another model does not exist), and each of its comments was checked against the primary source before entering the text. None of these reviews is independent peer review in the scientific sense. The contribution log, with dates and verbatim quotes, is in the research directory of the keryx repository together with the protocol and the counting script. We arrived at the authorship problem from our own data; we learned of the essay "Dead Cognitions" later and use it as the source of the term.
References
- Bainbridge L. Ironies of Automation. Automatica 19 (6), 1983, 775-779.
- Becker J., Rush N., Barnes E., Rein D. (METR). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089
- Bonaccio S., Dalal R. S. Advice taking and decision-making: An integrative literature review. Organizational Behavior and Human Decision Processes 101 (2), 2006, 127-151.
- Buçinca Z., Malaya M. B., Gajos K. Z. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. Proc. ACM Hum.-Comput. Interact. 5 (CSCW1), 2021, 188.
- Cagan M. The Inconvenient Truth About Product. SVPG, 2013. svpg.com
- Carta. Founder Ownership Report, 2026. carta.com
- Chowdhury K., Banik D., Ferdous K. M., Shamim S. I. From Industry Claims to Empirical Reality: An Empirical Study of Code Review Agents in Pull Requests. MSR 2026. arXiv:2604.03196
- Dell'Acqua F., Ayoubi C., Lifshitz-Assaf H., Sadun R., Mollick E. R., Mollick L., Han Y., Goldman J., Nair H., Lakhani K. R. The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise. NBER Working Paper 33641, 2025. nber.org
- DORA. State of AI-assisted Software Development 2025. Google Cloud, 2025. dora.dev
- Jesuthasan R., Boudreau J. Work Without Jobs. MIT Press, 2022.
- Johnson E. J., Goldstein D. Do Defaults Save Lives? Science 302 (5649), 2003, 1338-1339.
- Kohavi R., Crook T., Longbotham R. Online Experimentation at Microsoft. Data Mining Case Studies, 2009.
- LinearB. 2026 Software Engineering Benchmarks Report. linearb.io
- Nguyen H., Nguyen T. SDAD: Spec-Driven Agentic Development for the AI-Native SDLC. 2026. arXiv:2608.20341
- OMG. Structured Assurance Case Metamodel (SACM), version 2.3, 2023. omg.org
- Parasuraman R., Manzey D. H. Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors 52 (3), 2010, 381-410.
- Sabour S. et al. Human preferences are susceptible to covertly misaligned AI advice. PNAS 123 (38), e2600684123, 2026. doi:10.1073/pnas.2600684123
- Simkute A. et al. Ironies of Generative AI: Understanding and Mitigating Productivity Loss in Human-AI Interaction. arXiv:2402.11364, 2024.
- Sonar. State of Code Developer Survey Report, 2026 (n = 1,149, surveyed October 2025). sonarsource.com
- Stripe. Top solo founder traits. Stripe blog, 28 May 2026: share of solo-founded C corps incorporated through Atlas from the start of Q2 2026. stripe.com
- Tuor A., claude.ai (as listed by the authors). Dead Cognitions: A Census of Misattributed Insights. 2026. arXiv:2604.10288
- W3C. PROV-DM: The PROV Data Model. W3C Recommendation, 2013. w3.org
- Wilson K., Sim M., Gueorguieva A.-M., Caliskan A. No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy. arXiv:2509.04404, 2025.
- Zhou X., Saghi Z., Sabouri S., Pandita R., McGuire M., Chattopadhyay S. Cognitive Biases in LLM-Assisted Software Development. ICSE 2026, doi:10.1145/3744916.3773104. arXiv:2601.08045




Top comments (0)