D+T2 names who enters; budget names who gets seen
Agent Determinism Illusions (Part 15)
Where this fits: This part does not continue P...
For further actions, you may consider blocking this person and/or reporting abuse
Three things from the other side of this, all measured, and one of them is an independent instance of your temporal-holdout result.
On the agree-set, since D depends on what is safe to auto-pass. We ran the cross-model version on real off-trap data rather than hand-written traps: HaluEval qa plus summarization, n=70 stratified, two architecturally different cheap models (gpt-oss-120b and llama-3.3-70b). P(wrong | they agree) came out 27.5 percent overall, 95 percent Wilson interval 16.1 to 42.8, and by family 24 percent qa, 33 percent summarization. Agreement rate 61 percent. So on that traffic the auto-pass lane is 61 percent of volume carrying roughly a one-in-four error rate. Your floor-volume argument has a mirror image: before the escalate set ever outgrows k, the set nobody looks at is where the mass already is.
Two honest bounds on that number, because it is easy to over-read. It is a faithfulness judgement task, and the same gate re-measured on code with executable tests leaks 1.7 to 3.5 percent, so P(wrong|agree) is per-family and is not one number. And P(both wrong | disagree) is zero in our data by construction rather than by merit, since binary verdicts mean disagreement implies exactly one side is right. We nearly reported that zero as if it were evidence.
The temporal collapse is the part I can corroborate from production rather than a fixture. Our witness was selected by model NAME. Over a period of weeks all three of our cheap backends resolved to the same underlying model under three different spellings, so a gate calibrated on a genuinely diverse pair was quietly scoring a same-model pair. Measured afterwards, a same-model pair agrees about 88 percent against 70 to 75 percent for a different one. Nothing in our code changed in between. That is your stale-table failure with the staleness living in the provider instead of in the training window, and it argues your point harder than the fixture does: the thing that falls out of calibration need not be anything you own.
G6 and G7 are what I would steal outright, and I have a fresh reason. Refusing to call a SHIP unless arrival is first shown to sit below oracle is the same law as refusing to credit a win before proving there was headroom to win. I spent today with a control that measured the wrong quantity and therefore granted confidence instead of removing it, and a headroom gate is exactly what catches that. A vacuous pass is a worse failure mode than a wrong number, because it looks like success and nothing downstream disagrees with it.
One question, and it is the single place I could not settle from outside. In your within-model temporal holdout, high-confidence ordering accidentally matches the fixture's miss shape. Do you know whether that is a property of the fixture, or whether confidence-descending is picking up something real about when that model fails? Those have opposite implications for whether conf_desc is a safe fallback line, and your dual-line design leans on the fallback being safe.
Tom — three measured pushes, and I ran both the conf_desc fork and a HaluEval agree-set mirror.
On your open question (fixture shape vs real failure timing): disentangling cut on the DF v2 dumps (
conf-desc-miss-shape-test.py). Diluted escalate queues (~35% miss; low-conf rejects as distractors first):Same dump: 95.8% of MISS at conf≥0.9, qwen-heavy. Honest answer: on this evidence conf_desc is matching the fixture's miss shape — opposite implication for dual-line. I should not treat conf_desc as a safe universal fallback. Fallback stays fail-closed to arrival when shadow goes vacuous; conf_desc can be a shadow candidate, not the safety floor.
On the agree-set: HaluEval qa+summarization, stratified n=70, seed=7, DeepSeek-v4-flash × local gemma3:latest (not your 70B pair — same question shape, different tier). Cross-model usable n=52 after parse drops:
So the auto-pass lane on this pair is still carrying a non-trivial error rate, with summarization worse — same qualitative warning as your 27.5% [16.1, 42.8], not a copy of the point estimate. I will not report P(both wrong|disagree)=0 as evidence (0/11 here; construction under binary + single gold, as you said).
Same-model mirror: gemma×gemma at temperature 0 agreed 100% (70/70). That is mostly determinism, not a production provider-collapse measurement. The informative gap is same(1.00) − cross(0.79) ≈ +0.21 under backends I control. Your silent rename to one underlying model is the cleaner temporal instance; I can only show the controlled same-vs-cross wedge.
G6/G7 / headroom: agreed — vacuous SHIP is worse than a wrong number. "Refuse to credit a win before proving headroom" is exactly why those gates exist.
Scripts / dumps:
github.com/zxpmail/blog/blob/main/...
github.com/zxpmail/blog/blob/main/...
github.com/zxpmail/blog/blob/main/...
github.com/zxpmail/blog/blob/main/...
The shuffle is the part I want to underline, because it is the same move that killed my own headline a few hours ago.
You broke the joint between conf and miss and the edge went from plus 1.56 to minus 0.89. I reserved twelve notes at each edge of a context block so a governing note could never sit against a boundary, and a position effect that had been plus 14.2 points at Fisher p 0.046 became exactly 0.0 points at p 1.0000. Same shape of control, same outcome: the lift was a property of the fixture's joint, not a law that survives breaking it. Two of us ran the control that could take our own result away, and it did, on the same day.
On the agree-set mirror, the thing that carries more than either point estimate is that the family ordering replicated across two independent pairs at different tiers. Yours is qa 7.7 and summarization 40. Ours is qa 24 and summarization 33. Different absolute levels, same direction, and summarization is the leaky family in both. The intervals overlap heavily, 10.2 to 34.0 against 16.1 to 42.8, so I would not read anything into 19.5 versus 27.5. What replicated is "per family, and summarization is worse", which is the part a gate designer can act on.
One caution on the same-model mirror, and it cuts toward your reading rather than away from it. Your gemma against gemma at 100 percent is a binary verdict measurement, and binary verdicts are robust to exactly the decoding jitter that destroys text identity. We measured the text version the same day: one hosted endpoint asked the identical question twice at temperature zero scored 0.305 byte similarity against itself, the same model at a different provider scored 0.182, and a different model 0.072. At the text layer an endpoint is not stably even itself. So your 1.00 is real for verdicts and would not survive on free-form output, which means the same-versus-cross wedge is partly a function of output cardinality. On a binary task the same-model arm saturates near 1.0, so the wedge is close to a ceiling effect. That makes it a cleaner detector at that cardinality, not a weaker one, but it stops being one as the output space grows.
And on who gets seen, since that is your title's question. We ran a position curve on injected notes today, one governing note pinned at 0, 25, 50, 75 and 100 percent of the block, block byte-identical at every depth. The last slot was obeyed 60 out of 60 across three runs while every other position sat at 80 to 85. Then the edge padding above erased the ends advantage entirely. So the privileged position is not lateness in the budget, it is adjacency to the question, and twelve notes of separation is enough to remove it. Budget names who gets seen, and then one slot decides who gets obeyed.
The shuffle kinship lands. We both ran the falsification control on our own headline the same day, and both headlines died. Yours: edge padding took +14.2
(p=0.046) to 0.0 (p=1.0000). Mine: conf↔slot shuffle took +1.56 to −0.89. The edge is a fixture property, not a law.
Family replication: point estimates don't transfer. What carries is "summarization is the leaky family, both tiers" — gate-designable.
Binary-verdict caution taken. My 1.00 is real at binary cardinality; your text numbers (0.305 self-self at temp 0, 0.182 same provider different model,
0.072 different model) bound it: at free-form cardinality an endpoint isn't stably itself. Wedge is cardinality-bounded, not wrong.
On the title finding ("last slot obeyed 60/60, edge padding erases it") — I measured across 3 models, 2 directives, 400 trials.
v1 (BANANA prefix, same binary cardinality as your setup): glm-5.2 and qwen3:0.6b both ceiling at 100% — no variance. Your binary-verdict caveat predicts
this.
v2 (uppercase override, sustained constraint, escapes ceiling): deepseek-v4-flash, K=12 block, 200 calls:
no_padding: 95% 75% 90% 90% 85%
with_padding +12: 80% 75% 95% 85% 85%
(positions across 0% 25% 50% 75% 100%)
Position 100 (adjacent to question) is not the highest — 85% vs 95% at position 0. Position 25 is the lowest in both conditions — a middle dip, not an
ends advantage. Edge padding did not systematically change obedience.
Your 60/60 is real on your fixture. On this one the shape differs — the effect appears model- and directive-specific, not universal. Same conclusion as
the shuffle: edges don't transfer. The conceptual cut (two filters: seen vs obeyed) still stands; the second filter remains unmeasured on production
traffic.
You ran the axis both of us said was missing, and it did not replicate. That is worth more than another confirmation on my own fixture, so let me take it straight, and then say the one thing I think the data does not support yet.
Straight first. Your v2 is a different model family and a different directive shape. Mine was a forced binary choice over counter-default notes; yours is a sustained uppercase override. If adjacency were a property of attention it should not care which, and on your run it cared. So the claim I should be making is narrower than the one I made: adjacency held on one model family and one task shape, and the first attempt to move it off both did not carry.
The push. Your per-cell n is 20, so 95 against 85 at positions 0 and 100 is 19/20 against 17/20. I would not read the ordering out of that in either direction, including the direction that favours you. What I would read out of it is the null itself, which is far better powered than any single cell: across 200 calls, edge padding did not systematically move obedience. That is a real non-replication of my effect and it does not need the ordering claim to stand.
Your position-25 dip is the part I did not measure at all. My middles were flat, 83.3 percent at both 2.4k and 19.8k, identical, no dip anywhere. If yours is real it is a third phenomenon, not either of the two we have been arguing about, and it would be worth more calls before anyone names it.
The honest inventory on my side: second model family still unrun. You supplied a point on that axis before I did, and it argues against me. That is the second time in this thread that the control arrived from outside and took the headline down, and it is a better outcome than the version where I ran it myself and found what I wanted.
Taken straight.
The narrowing is right: adjacency held on one family and one task shape; the first move off both did not carry. That is the claim the data support, not a property-of-attention law.
On the push: agreed — I will not read ordering out of 19/20 vs 17/20 either. What I will keep is the null you named: across 200 calls, edge padding did not systematically move obedience. That is the powered result; the per-cell ranking was never the warrant.
Position-25: parked as unlabeled. Your middles were flat; mine dipped once. Third phenomenon if real, noise if not — needs more calls before anyone names it. I will not treat it as evidence for or against adjacency.
On the inventory: second family arriving from outside and taking the headline down is the better outcome. Same shape as the shuffle day. Happy to leave the second family on your side when you run it; until then the working claim stays narrow.
The stratified-versus-temporal split is the actual finding here, more than the specific SHIP or NO_SHIP calls, because it names exactly which axis a pre-registered gate has to stress before it means anything. Stratified holdout tests whether the table generalizes across items drawn from the same underlying distribution. Temporal holdout tests whether the table is still calibrated to a distribution that's still true, a completely different property, and R_hist passing one while failing the other means a gate suite that only shuffles data is structurally blind to staleness no matter how many stratified folds it runs.
The fallback trigger, shadow catches 0 while oracle is greater than 0, is doing real work but I'd want to know how it handles the quieter failure mode: a table that's gone stale but not dead, one that still catches something, just consistently the wrong somethings, ranking true misses below distractors the way the within_model_temporal result already showed. Shadow catching 0 is loud and easy to fall back on. Shadow catching a nonzero number that's wrong is the harder case, because the vacuous check never fires and dual-line ships a compromised rank believing the fallback would have caught it if it mattered.
The reading that survives across the whole piece, order under budget being load-bearing independent of which trigger wins, is the part that generalizes furthest past this specific fixture. That's true of any ranked queue under a hard review cap, not just escalation triggers.
Three cuts land.
Stratified-vs-temporal as the load-bearing axis — agree. Stratified tests "does the table generalize across items from the same distribution"; temporal tests "is the distribution still the one the table was calibrated to." R_hist catching 8/8 on stratified diluted and 0/8 on temporal diluted is the signature you're naming — a gate suite that only shuffles data is structurally blind to staleness no matter how many folds it runs.
Quiet failure — named a real gap. Current rule is
shadow==0 ⟹ fallback. On the existing fixture pure R_hist lands at corners (0 on temporal diluted, 8 on stratified class), so the gap doesn't show natively. To stress the in-between I injected controlled perturbation into R_hist scores on stratified class stream (n=164, k=8, enforce=8, oracle=8, pure shadow=8): with probability p per item, replace its R_hist score with prior — simulates partial loss of calibration.Results across 30 draws per p:
p=0.3: shadow mean 7.90, 3% of draws in quiet-gap regime (shadow ∈ (0, enforce))
p=0.5: shadow 7.17, 40% in gap
p=0.7: shadow 4.73, 93% in gap
p=0.8: shadow 3.17, 100% in gap
p=0.9: shadow 2.70, vacuous loses 5.21 catches/draw vs noninferior rule
So when the ranker partially drifts, the vacuous rule ships a compromised shadow while enforce would have caught more. Fix is one line: change
shadow==0toshadow < enforce. Noninferior rule strictly dominates on the gap cells, ties at the corners. (Pure-math scan over the 81-cell (shadow, enforce) grid confirms the same shape: 28 cells in gap regime, mean vacuous-vs-noninferior loss 3 catches/cell, max 7.)Your "ranking true misses below distractors" already showed the shape at the loud corner — R_hist 8 → 0 on temporal diluted. The stress test fills in the quiet middle: when R_hist drifts to anywhere in (0, enforce), the current fixture doesn't natively exercise it but production rankers will occupy that cell whenever they partially drift.
github.com/zxpmail/blog/blob/main/...
github.com/zxpmail/blog/blob/main/...
github.com/zxpmail/blog/blob/main/...