DEV Community

Cover image for I Replaced a Gate That Accepted Everyone With a Gate That Accepted No One. My Tests Couldn't Tell the Difference.
Self-Correcting Systems
Self-Correcting Systems

Posted on AI-assisted

I Replaced a Gate That Accepted Everyone With a Gate That Accepted No One. My Tests Couldn't Tell the Difference.

Run this in a terminal, then run it again under script(1):

python3 -c 'import sys; print(sys.stdin.isatty(), sys.stdout.isatty())'
Enter fullscreen mode Exit fullscreen mode

If your code decides whether a human is present by calling isatty(), and then reads that
human's confirmation from /dev/tty, those are two different questions, and there are
processes that answer them differently. This is the story of finding that out three times in
one file, each time because the previous fix was wrong in a way my tests could not see.

Here is the probe. Standard library only, writes no files and mutates no state; full source at the end.

$ script -q /dev/null python3 gate_probe.py --redirected

  isatty(stdin) and isatty(stdout) : False
  /dev/tty openable                : True
  detail                           : opened
  open("/dev/tty", "r+") -> UnsupportedOperation: File or stream is not seekable. (errno=None)

  THE TWO PREDICATES DISAGREE for this process.
  A gate on isatty and a confirmation read on /dev/tty will not agree here.
Enter fullscreen mode Exit fullscreen mode

That process has an openable controlling terminal even though its redirected stdin and stdout
are not TTYs. For three days that was the shape of my gate.

One precision, because a reader will check. Default pytest capture replaces both streams,
which I had wrong in an earlier version of this paragraph. On pytest 9.0.2, launched from a real
terminal:

default capture   stdin=DontReadFromInput (isatty False)   stdout=EncodedFile (isatty False)
pytest -s         stdin=TextIOWrapper     (isatty True)    stdout=TextIOWrapper (isatty True)

/dev/tty under default capture: OPENS
Enter fullscreen mode Exit fullscreen mode

So plain pytest from a terminal is the disagreement, on its own — no wrapper needed. My gate
required stdin.isatty() and stdout.isatty(), false under default capture for two reasons at
once, while the confirmation path could open the terminal the whole time. Under pytest -s the
two predicates agree, so don't look for a mismatch there.

The system, briefly

An agent that can run three frozen commands against an exported copy of a repository. It has
never been authorized for normal execution — the policy file forbids it — and I temporarily
flipped that flag during the test run described below, which is how the rest of this post
exists.

$ python3 -c "import sys; sys.path.insert(0,'.')
import stage2a_preflight as PF; print(PF.check_policy())"

(False, 'policy forbids execute_shell; owner has not authorized Stage 2a')
Enter fullscreen mode Exit fullscreen mode

Before it runs anything, a human is supposed to be shown what will happen and type two
identifiers back.

v0 — the gate was a parameter

def execute(job, approval, typed_job_id=None, typed_procedure_id=None):
    ok, note = confirm_approval(approval, typed_job_id, typed_procedure_id)
    if not ok:
        return "REFUSED_APPROVAL_MISSING", note
Enter fullscreen mode Exit fullscreen mode

Read it and it looks like a human gate. typed_job_id is a parameter, and a parameter is
something any caller supplies. A test fixture is a caller with two correct strings.

I flipped the policy flag that permits execution and ran the suite. Afterward:

$ ls -lT state/attempts | awk 'NR>1 {print $6, $7, $8, $10}'
Sep 24 16:36:41 TEST-JOB-1-1790282200
Sep 24 16:36:42 TEST-JOB-1-1790282202
Sep 24 16:36:44 TEST-JOB-1-1790282204
Sep 24 16:36:45 TEST-JOB-1-1790282205
Sep 24 16:37:42 TEST-JOB-1-1790282262
Sep 24 16:37:43 TEST-JOB-1-1790282263
Sep 24 16:37:45 TEST-JOB-1-1790282265
Sep 24 16:37:46 TEST-JOB-1-1790282266
Sep 24 16:36:43 TEST-JOB-2-1790282203
Sep 24 16:37:44 TEST-JOB-2-1790282264
Enter fullscreen mode Exit fullscreen mode

Sorted by name, not time. Read the clock column and it is two runs of five, about a minute
apart
— 16:36:41–45 and 16:37:42–46 — eight for one fixture job and two for a second. Each
directory holds a 266,240-byte tar of my own repository, all ten identical in size, extracted
into a 27-file tree.

The gate opened exactly as written.

What those ten do and do not prove

They contain a third entry, runtime-tmp, and it is empty. Its existence is not proof the
three commands started, because of where it is created:

def frozen_env(attempt_root):
    """addendum v1 B1. An allowlist, not a filtered copy. HOME absent. No PYTHON*."""
    tmp = attempt_root / "runtime-tmp"
    tmp.mkdir(parents=True, exist_ok=True)

def run_frozen_commands(export_root, attempt_root):
    env = frozen_env(attempt_root)
    results = []
    for entry in PF.FROZEN_COMMANDS:
Enter fullscreen mode Exit fullscreen mode

frozen_env() is the first statement of run_frozen_commands(), and the for loop below it is
where commands actually launch. So the directory is created before anything runs. So runtime-tmp proves the runner was entered. Whether any command
started is not established by these directories, and I am not going to round that up.

What the ledger says is stranger. The ten exports left no receipt at all:

$ python3 -c "
import json
for l in open('state/receipts.jsonl'):
    r = json.loads(l)
    print(r['final_state'], r['job_id'], len(r['commands'] or []), r['written_at'])"

REFUSED_NO_HUMAN_PRESENT tg_912616161 0 2026-09-24T21:29:28.666303+00:00
REFUSED_NO_HUMAN_PRESENT tg_912616161 0 2026-09-24T21:30:40.369098+00:00
REFUSED_NO_HUMAN_PRESENT TEST-JOB-1 0 2026-09-28T01:31:55.420602+00:00
Enter fullscreen mode Exit fullscreen mode

Three rows, zero commands. None of them is the export run. Rows 1 and 2 land fifty-three
minutes after it; row 3 is from this repair session, below.

There is also a quarantine file from that day, receipts_TEST_POLLUTION_QUARANTINED_2026-09-24,
and the export run is not in that either — its ten rows are all REFUSED_COMMAND_BOUNDARY
written between 15:02 and 16:06 UTC, while the exports are 20:36-20:37 UTC. A different event,
four to five hours earlier.

So: ten real exports in production state, and no receipt for them in either ledger. Where that
receipt went I cannot establish. execute() writes one in every terminal case, so either it was
written to a redirected path and discarded with a temp directory, or something raised before the
write. I did not preserve the test configuration from that run and there is no version control on
that tree, so absence here is absence — not evidence of a destination.

Those rows read REFUSED_NO_HUMAN_PRESENT, which is the same overclaim I rename a function for
further down. That value is now REFUSED_CONTROLLING_TERMINAL_UNAVAILABLE and the receipt schema
went v1 -> v2 to say so. The three rows above were not rewritten — an append-only ledger
edited to match new vocabulary is not an audit trail — so the file holds both values, and
schema is what tells a reader which vocabulary a row was written under.

The thing that was actually protecting me

Not the typed confirmation. A boolean in a config file that I had left set to false.

v1 — the fix could not open a terminal

Move the read off the parameter list and onto the controlling terminal. There is no argument
to fill:

def read_confirmation(approval):
    with open("/dev/tty", "r+") as tty:      # <- this line
        ...
Enter fullscreen mode Exit fullscreen mode

On this macOS machine, open("/dev/tty", "r+") fails on both CPython 3.9.6 and 3.13.9. The
update-mode I/O stack buffers through BufferedRandom, which wants a seekable raw stream; this
terminal is not one:

$ script -q /dev/null python3 gate_probe.py </dev/null

  isatty(stdin) and isatty(stdout) : True
  /dev/tty openable                : True
  detail                           : opened
  open("/dev/tty", "r+") -> UnsupportedOperation: File or stream is not seekable. (errno=None)
Enter fullscreen mode Exit fullscreen mode

The last line is the probe opening the same terminal it just opened successfully, in the mode
I had used. Reproduced identically on CPython 3.9.6 and 3.13.9. And then the detail that turned a broken
line into a broken control:

$ python3 -c "import io; print(io.UnsupportedOperation.__mro__)"
(<class 'io.UnsupportedOperation'>, <class 'OSError'>, <class 'ValueError'>, ...)
Enter fullscreen mode Exit fullscreen mode

io.UnsupportedOperation subclasses OSError, and carries errno=None. My handler was:

except OSError as exc:
    return "REFUSED_NO_HUMAN_PRESENT", "no controlling terminal"
Enter fullscreen mode Exit fullscreen mode

So a human sitting at a real terminal was told they were not there, and the receipt recorded
the wrong cause. It failed closed, which is the good direction to fail — but the gate now
refused the only caller it was built for.

In v1, every test that exercised approval replaced read_confirmation by name. Not one
exercised the real tty read. v0 had the earlier version of the same blind spot: there was no
read_confirmation to replace, because its tests supplied the ids directly as arguments.
Different bypass, same hole — neither broken version had a human-interaction boundary under test,
so the suite was green for both. It was never testing the gate; it was testing the bypass.

v2 — and the third defect, which the third test found

Separate handles, and an except narrow enough to mean something:

CTTY_UNAVAILABLE_ERRNOS = frozenset({errno.ENXIO, errno.ENODEV, errno.ENOTTY, errno.ENOENT})

def open_controlling_terminal():
    out = open("/dev/tty", "w")
    try:
        inp = open("/dev/tty", "r")
    except BaseException:
        out.close()
        raise
    return out, inp
Enter fullscreen mode Exit fullscreen mode

Then a test I wrote for this failed, and I nearly patched the test. That would have been the
fourth version of the same mistake. What it had found:

require_human()        asked  "is sys.stdin a tty?"
read_confirmation()    asked  "can I read /dev/tty?"
Enter fullscreen mode Exit fullscreen mode

Two different predicates for one control, and the stale one ran first. They are not
ordered — one asks whether two particular streams are terminal devices, the other asks for the
process's controlling terminal. In the captured pytest configuration I reproduced, isatty
was false while /dev/tty stayed openable, so the stale check refused before the confirmation
path could run. Fail-closed, but the real control was shadowed by a different predicate.

I first wrote that this shadowing was why the r+ version survived the suite. That is a
wrong-reason claim in a post about wrong-reason claims, and the probe disproves it:

no controlling terminal  ->  OSError errno=6 (ENXIO)
terminal attached        ->  UnsupportedOperation (errno=None)
Enter fullscreen mode Exit fullscreen mode

r+ fails two different ways, and except OSError swallows both. Delete the shadow and
run again: with no terminal you get a real ENXIO, which is the verdict every refusal test
expects; at a terminal you get UnsupportedOperation relabelled into the same refusal. Either
way, green.

So the shadowing is a second independent reason nothing could have caught it, not the cause.
The causes are the two already named: every test patched read_confirmation, and the except
was wide enough to swallow a defect. What the shadow did do is stop the real path from being
exercised by that integration route — which is why the fix needed a test that calls it
directly.

So there is now one predicate, in one function, used by both:

def require_controlling_terminal():
    try:
        out, inp = open_controlling_terminal()
    except OSError as exc:
        return False, controlling_terminal_unavailable(exc)
    out.close()
    inp.close()
    return True, "controlling terminal present and openable at /dev/tty"
Enter fullscreen mode Exit fullscreen mode

And controlling_terminal_unavailable() re-raises anything outside the set, so a programming
fault can no longer be reported as an absent human. The refusal carries the evidence rather
than the conclusion:

controlling terminal unavailable: /dev/tty open failed with errno 6 (ENXIO): ...
Enter fullscreen mode Exit fullscreen mode

Two renames went with this, both the same correction. require_human() asserted something no
check in that file can establish — it cannot prove a person is present, only that confirmation
is obtainable from a terminal, and naming it after the stronger claim is what licensed a second
predicate to grow beside it. And the errno bucket was called NO_CTTY_ERRNOS while containing
ENOENT (/dev/tty does not exist here) and ENOTTY (that fd is not a terminal). Neither
literally means "this process has no controlling terminal." A set named for its strongest member
is the same overclaim one level down, so it is CTTY_UNAVAILABLE_ERRNOS now. The bucket is a
decision about what to refuse on; the errno is the fact.

Passing tests prove nothing, so I broke it on purpose

Four mutations against the gate, each run as the full 131-test suite:

mutation caught by
restore open("/dev/tty", "r+") 4 tests
treat any OSError as an absent human 3 tests
restore the isatty predicate 9 tests and subtests
drop the explicit isinstance(io.UnsupportedOperation) guard nothing. 131 passed

The fourth is the useful one. That line is redundant: errno is already None for
UnsupportedOperation, so the errno test re-raises it anyway. I kept it, and labelled it in
the source as documentation rather than a control, because a line the suite cannot
distinguish from its own absence is not protecting anything.

Population, since it is the whole point of this post: 5 of 131 tests run under a pty, and one
more opens /dev/tty in-process, for six that exercise a terminal-backed path. A pseudo-terminal
is deliberately not evidence of a physical terminal or a person, which is why that phrasing is not
"six tests prove a human."

The count in order, because it is the argument and a reader should be able to check the ordering:

version what the gate was tests opening a terminal
v0 a function parameter 0
v1 /dev/tty with r+ 0
v2 separate handles 2 — written as part of this repair
v3 one predicate 5 under a pty, +1 in-process

The two terminal tests did not exist while r+ was in the tree. They were written to fix it, and
they are what caught it — restore r+ today and four tests go red, which is the first row of the
mutation table above. Zero is the number that carried v0 and v1, and zero is why both shipped
green.

$ python3 -m pytest -q tests/
131 passed, 25 subtests passed
Enter fullscreen mode Exit fullscreen mode

The honest claim about what this technique is worth

sudoers(5), verbatim from this machine:

requiretty   If set, sudo will only run when the user is logged in to a real tty.
             When this flag is set, sudo can only be run from a login session and
             not via other means such as cron(8) or cgi-bin scripts.
             This flag is off by default.
Enter fullscreen mode Exit fullscreen mode

Off by default, and note what the description is about: cron and CGI, not adversaries. Which
is the correct amount of credit to give this. It is not a security boundary:

$ python3 -c 'import sys; print(sys.stdin.isatty())'
False
$ echo "" | script -q /dev/null python3 -c 'import sys; print(sys.stdin.isatty())'
True
Enter fullscreen mode Exit fullscreen mode

script(1) allocates a pseudo-terminal. Code running with my user's permissions and access to
PTY facilities can do that, including the assistant I write most of this code with. Its default
tool calls fail the check — I have watched them fail, and one of my own tests is that wrapper
passing on purpose:

def test_a_pty_wrapper_satisfies_the_gate(self):
    """Stated as a limit, not a defect. This is what the post claims."""
Enter fullscreen mode Exit fullscreen mode

So the claim is narrow: this converts an accidental bypass into a more explicit one. My
current noninteractive automation path no longer gets past it merely by supplying function
arguments. An environment that already controls a PTY can still satisfy the terminal path
without writing anything new, so this is not identity and not authentication. It is worth
having, and it is not the same as being safe.

The second defect, which came back in three days

The ten exports landed in live state because my tests wrote to production paths. I fixed that
on the 24th by redirecting three module constants in the setUp of each class that needed it.

Row 3 of that ledger is the 27th local — 01:31 UTC on the 28th, which is why the timestamp
you read above looks like a different day. A test I wrote during this repair reached
write_receipt() without a redirect and appended to the real ledger. Its refusal_detail is
the old isatty message, which is how I know which version wrote it.

The lesson is not "remember to redirect." Isolation was opt-in, so correctness depended
on every future author choosing to comply, and I was the future author who did not. It is now
opt-out:

@pytest.fixture(autouse=True)
def isolate_agent_state(request):
    if request.node.get_closest_marker("live_state"):
        yield
        return
    ...redirect RECEIPTS, NONCES, ATTEMPTS to a temp dir...
Enter fullscreen mode Exit fullscreen mode

A test must now declare @pytest.mark.live_state to touch real state, and that declaration
is visible in the test source. Verified load-bearing by flipping autouse=False: four
failures, from tests that exist only to prove the fixture fires.

If a rule in your project is enforced by someone remembering it, it is a request, not a
control. Mine took three days to prove that.

The probe

The tool at the top. Four modes; the one that matters is --redirected, which reproduces the
disagreement:

python3 gate_probe.py                                    # pipeline / CI shape
script -q /dev/null python3 gate_probe.py </dev/null     # terminal shape
script -q /dev/null python3 gate_probe.py --redirected   # ctty intact, stdio redirected
script -q /dev/null python3 gate_probe.py --detach       # no controlling terminal at all
Enter fullscreen mode Exit fullscreen mode

And a cheap smell check, labelled as what it is — a string search, not proof:

grep -rn 'open(["'"]/dev/tty' tests/
Enter fullscreen mode Exit fullscreen mode

A bare grep -rc '/dev/tty' tests/ is worse than useless: it prints one count per file rather
than a single number, and it counts docstrings. On my own suite 9 of its 15 hits are prose
about the terminal. It can only err in the flattering direction, which is the exact failure
mode this post is about. A zero would not prove much either — a test can reach /dev/tty
through production code without the string appearing in tests/ at all.

The evidence that actually settles it is above: a test that exercises the real path without
patching it, plus a mutation showing the suite goes red when that path breaks.

I cannot show you that number for my own broken versions — that tree is not under version
control. What I can show is structural and needs no count: v1's approval tests replaced
read_confirmation by name, and v0 had no terminal boundary to exercise at all.
A suite that
never opens /dev/tty cannot report anything about a gate that reads from it, whatever its
total.

gate_probe.py, in full

#!/usr/bin/env python3
"""Which question is your human-in-the-loop gate actually asking?

Run it four ways and compare the two verdict columns:

    python3 gate_probe.py                                  # a pipeline / CI shape
    script -q /dev/null python3 gate_probe.py </dev/null    # a terminal shape
    script -q /dev/null python3 gate_probe.py --redirected   # ctty intact, stdio redirected
    script -q /dev/null python3 gate_probe.py --detach       # no controlling terminal at all

`isatty` and `/dev/tty` are different predicates, and a gate built on the first while it
reads from the second will disagree with itself. The row that matters is the one where the
two columns differ: that is a process your gate classifies one way and your confirmation
code classifies the other.

Standard library only. Writes no files and mutates no state; it prints to stdout.
"""

import errno
import io
import os
import subprocess
import sys

# Named for what the errnos establish, not for the strongest one in the set. ENOENT means
# /dev/tty does not exist here; ENOTTY means that fd is not a terminal. Neither literally
# asserts "this process has no controlling terminal."
CTTY_UNAVAILABLE = {errno.ENXIO, errno.ENODEV, errno.ENOTTY, errno.ENOENT}


def probe():
    """Return (isatty_verdict, dev_tty_verdict, detail)."""
    stdio = sys.stdin.isatty() and sys.stdout.isatty()

    try:
        out = open("/dev/tty", "w")
        inp = open("/dev/tty", "r")
        out.close()
        inp.close()
        return stdio, True, "opened"
    except io.UnsupportedOperation as exc:
        # Not an unavailable terminal. Reported separately because it subclasses OSError with
        # errno None, so an `except OSError` upstream will call this "no human present."
        return stdio, False, f"UnsupportedOperation: {exc} (errno={exc.errno}) -- A BUG, NOT AN ABSENCE"
    except OSError as exc:
        kind = ("controlling terminal unavailable"
                if exc.errno in CTTY_UNAVAILABLE else "UNEXPECTED")
        name = errno.errorcode.get(exc.errno, "?")
        return stdio, False, f"{kind}: errno {exc.errno} ({name})"


def also_show_the_broken_open():
    """The mode that looks correct and is not. A tty is not seekable; text update mode
    wants it to be. Reproduced on CPython 3.9.6 and 3.13.9."""
    try:
        open("/dev/tty", "r+").close()
        return 'open("/dev/tty", "r+") -> opened'
    except OSError as exc:
        return f'open("/dev/tty", "r+") -> {type(exc).__name__}: {exc} (errno={exc.errno})'


def main():
    me = os.path.abspath(__file__)

    if "--detach" in sys.argv:
        # setsid() leaves the session, so the child loses the controlling terminal. Only
        # differs from a plain run if the PARENT had one -- run it under script(1) to see it.
        r = subprocess.run([sys.executable, me], preexec_fn=os.setsid,
                           capture_output=True, text=True)
        sys.stdout.write(r.stdout + r.stderr)
        return 0

    if "--redirected" in sys.argv:
        # The case that produces the disagreement, and the shape of `pytest` launched from
        # a terminal: stdio replaced, session intact. Run this one under script(1).
        r = subprocess.run([sys.executable, me], stdin=subprocess.DEVNULL,
                           capture_output=True, text=True)
        sys.stdout.write(r.stdout + r.stderr)
        return 0

    stdio, ctty, detail = probe()
    print(f"  isatty(stdin) and isatty(stdout) : {stdio}")
    print(f"  /dev/tty openable                : {ctty}")
    print(f"  detail                           : {detail}")
    print(f"  {also_show_the_broken_open()}")
    if stdio != ctty:
        print("\n  THE TWO PREDICATES DISAGREE for this process.")
        print("  A gate on isatty and a confirmation read on /dev/tty will not agree here.")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

The question

Not "can a test satisfy your gate" — I had that as the ending for two drafts and my own suite
disproves it. test_a_pty_wrapper_satisfies_the_gate exists on purpose. A test that can drive the
real boundary deliberately is what good integration coverage looks like.

The question is:

Can ordinary automation satisfy your human-in-the-loop gate without crossing the interaction
boundary you intended?

If it can pass merely by supplying the right values into the same function call, then the gate has
established knowledge of those values, not human involvement. v0 established that a caller knew
two strings. It never established anyone read them.

And the follow-up that took me three versions to reach: does any test exercise the real
boundary, or do they all replace it?
Mine all replaced it — which is how one suite certified a
gate that accepted everyone and a gate that accepted no one, and reported both as correct.

Mine held for exactly as long as the flag next to it was set to false.

Top comments (1)

Collapse
 
nomad-link-id profile image
Igor Eduardo •

The line that should scare anyone shipping an agent with a “human gate” is structural: the same suite certified a gate that accepted everyone (parameters as approval) and a gate that accepted no one (broken /dev/tty path), because both versions replaced the interaction boundary in tests.

Green did not mean the control worked. Green meant the bypass still compiled.

I would steal two checks from your repair story for any approval / policy / “human present” control:

  1. Which tests open the real boundary? If every case patches read_confirmation (or injects the typed ids as arguments), the suite is grading the stub.
  2. Mutation on the control, not the happy path. Restore the broken open mode, widen the except, reinstate the stale isatty predicate — if nothing goes red, that line was documentation, not a gate.

Narrow claim, same as yours: converting an accidental bypass into an explicit one is worth having, and it is still not identity. The eval contract for the gate is “does ordinary automation pass without crossing the interaction you intended?” — not “did 131 tests pass.”