The rule is simple to explain and surprisingly hard to evaluate.
For the 2026–27 season, English professional football is trialling a response to suspected goalkeeper “tactical timeouts.” When play is stopped because a goalkeeper is injured, the coach must nominate an outfield player to leave the pitch. That player remains off for at least one minute after play restarts. The official protocol includes medical exceptions, such as bleeding or a collision in which both the goalkeeper and an outfield player need treatment.
The policy changes an incentive. It does not automatically create a measurement system.
As of 14 August 2026, the Premier League season has not started. The current Xtra-Stats snapshot contains 37 scheduled league matches between 21 August and 12 September, all still to come. That is an incomplete early schedule, not a full-season sample. There is no honest result to report yet.
What we can do now is define a protocol that will tell us, later, whether the trial changed behaviour — without mistaking correlation, missing data or a small sample for proof.
Start with the causal question, not the dashboard
The tempting question is: “Does the new rule work?” That is too broad for a first analysis.
A useful evaluation breaks the mechanism into a chain:
- Do non-exempt goalkeeper injury stoppages become less frequent?
- Does their duration change?
- Is the nominated player kept off for at least one minute after the restart?
- Does the interruption still create a tactical reset for the defending team?
- Do any downstream match outcomes change?
The first three are process questions. They are close to the intervention and can be observed directly. The fourth requires a defensible definition of “tactical reset.” The fifth is far downstream and exposed to almost every confounder in football: team strength, score state, substitutions, venue, red cards, game phase and simple randomness.
That ordering matters. A trial could reduce suspicious interruptions without moving goals, shots or possession enough to be detectable. Calling it ineffective because the final score distribution did not change would be a category error.
My proposed primary endpoint is therefore narrow:
Non-exempt goalkeeper injury stoppages per 100 matches.
Everything else should be secondary or exploratory until the incident data are mature.
Build the incident layer the existing data cannot supply
Xtra-Stats already provides match schedules, results and team-level performance measures. It does not currently identify goalkeeper injury stoppages, the player nominated to leave, the exception applied, or the one-minute period after the restart.
That gap cannot be filled by treating a missing record as “no incident.” It requires a separate observation layer, created from match video, official match reports or another source that records the stoppage itself.
For every possible incident, the annotation guide should capture:
| Observation | Why it matters |
|---|---|
| Match, team and minute | Locates the incident and its game phase |
| Score state | Teams may value a pause differently when leading, level or trailing |
| Stoppage start and restart | Measures the interruption rather than estimating it |
| Reason shown or reported | Distinguishes visible treatment from an unsupported assumption |
| Official exception category | Keeps exempt incidents out of the primary endpoint |
| Nominated outfield player | Tests whether the trial procedure was applied |
| Time the player left and returned | Measures the one-minute compliance window |
| Evidence source and confidence | Makes uncertain annotations auditable |
The codebook must be written before analysts start labelling matches. Terms such as “injury stoppage,” “restart” and “return to play” need operational definitions. For example, does the clock start when the referee stops play, when treatment begins, or when the goalkeeper first sits down? Any choice can work if it is consistent, justified and frozen before the results are inspected.
At least a sample of incidents should be labelled independently by two reviewers. Agreement should be measured on the incident decision, exception category and timestamps. Disagreements are not noise to hide; they reveal where the rule or the codebook is ambiguous.
Keep genuine treatment visible
The trial is intended to discourage tactical use of alleged injuries while preserving legitimate care. That means the analysis must not equate “non-exempt” with “fake” or “exempt” with “genuine.” The official categories determine how the trial is applied, not what was in a player’s mind.
The safest public labels are descriptive:
- exempt under the published protocol;
- non-exempt under the published protocol;
- unclear from the available evidence.
The unclear category is essential. Forcing every ambiguous case into a binary outcome would manufacture precision and could unfairly imply deception.
The official wording also matters. The exception list is introduced as including specified situations, so it should not be presented as necessarily exhaustive. If competition guidance or referee instructions add detail during the season, the codebook should be versioned and historical labels re-audited.
Use three datasets, with three different jobs
A credible evaluation needs three connected but separate layers.
1. The incident log
This is the new observational dataset. It answers whether the event occurred, how the protocol was applied and how long the relevant periods lasted.
2. Match context
This supplies the minute, score, home/away status, substitutions, cards and other events needed to compare like with like. It prevents a stoppage in the 92nd minute with a team defending a lead from being treated as equivalent to one in the 12th minute at 0–0.
3. Match-level outcomes
Xtra-Stats data can add shots, shots on target, possession, goalkeeper saves, cards and results. These measures describe the match around an incident. By themselves, they do not identify the incident and cannot prove why a goalkeeper required treatment.
The 2025 Premier League season shows why coverage has to be audited before modelling. Xtra-Stats contains 380 completed matches. Team-level match statistics are available for 358 of them, or about 94.2%, producing 716 team-match rows. Total shots and possession are present in all 716 rows. Goalkeeper saves are present in 712 rows, or about 99.4%; four rows are missing and must not be converted to zero.
That is strong coverage for contextual outcomes, but it is not complete coverage. It also says nothing about the missing intervention. A separate audit of the recorded event labels found no injury, timeout or medical event category in the 2025 league data. The correct statement is “the feature is not captured,” not “there were no such incidents.”
Define the metrics before seeing the season
The measurement plan should be registered before enough 2026–27 matches exist to tempt analysts into selecting the most dramatic result.
I would use these process measures:
- Incident rate: non-exempt stoppages divided by matches observed, reported per 100 matches.
- Exception rate: exempt incidents divided by all annotated goalkeeper injury stoppages.
- Stoppage burden: median seconds from the referee stopping play to the restart, with the full distribution rather than only an average.
- Compliance exposure: seconds the nominated outfield player is unavailable after the restart.
- Repeat rate: teams with more than one non-exempt incident in a defined rolling period.
- Uncertainty rate: possible incidents that cannot be classified from the evidence.
The denominator must remain stable. Switching from “per match” to “per goalkeeper treatment” midway through the analysis can reverse the apparent direction of a trend.
For exploratory match effects, define fixed windows around the restart — for example the next five minutes of active play — and compare shots, entries into the penalty area or possession sequences. Full-match totals are too coarse for this purpose. If the data source cannot produce time-windowed measures, say so and keep those outcomes out of the main claim.
Goalkeeper saves are context, not a timeout detector
Goalkeeper saves are an attractive metric because they describe workload. In the audited 2025 sample, the 712 available team-match values range from 0 to 10, with an average of about 2.79.
That does not make saves a proxy for tactical timeouts.
A goalkeeper facing repeated shots may need genuine treatment more often, may be under more pressure, or may simply play in a match with an unusual tactical profile. Saves can be used as a contextual variable when comparing matches, but a relationship between saves and stoppages would not establish causality.
This is also where zero and missing data must stay separate. Zero saves is an observed football value. A missing save value means the measure was unavailable for that team-match row. Combining them would bias both the distribution and any model that uses workload as a control.
Choose comparisons that can fail honestly
No single comparison will remove every source of bias. A useful analysis should combine several views and make their assumptions explicit.
Before and after within the Premier League
Compare 2026–27 with a pre-trial baseline using the same incident codebook. This is intuitive, but season-to-season changes in teams, refereeing, added time and playing style can contaminate the result.
A comparison competition
Track a similar competition that is not using the trial, if one can be identified with comparable incident evidence. Compare the change over time in the trial league with the change in the comparison league. This is stronger than a simple before/after chart, but only if the leagues followed similar trends before the intervention.
An interrupted weekly series
Aggregate the incident rate by matchweek and examine whether there is an immediate level change or a gradual trend. This can reveal adaptation, but early weeks will have wide uncertainty intervals.
Matched incident analysis
For post-restart effects, match incidents on game minute, score state, home/away status, team strength and goalkeeper workload. Treat this as exploratory unless the sample becomes large and the matching quality is demonstrated.
Across all four designs, report counts and uncertainty intervals. A percentage based on three incidents is not a stable signal, however impressive it looks in a chart.
A practical first-six-weeks workflow
The opening phase should prioritise data quality over a weekly verdict.
Before matchweek one
- freeze the incident codebook and primary endpoint;
- record the official protocol and known exceptions;
- test the annotation form on historical video;
- define the comparison period and any comparison league;
- publish an internal data-quality checklist.
During weeks one and two
- review every match for possible incidents;
- double-label every candidate;
- resolve disagreements without looking at team-level outcomes;
- track missing video or unclear evidence separately.
During weeks three to six
- double-label a random quality-control sample;
- report only process metrics with counts and intervals;
- audit whether codebook changes altered earlier labels;
- keep downstream football outcomes exploratory.
Each snapshot should have a clear cut-off time. The current schedule illustrates why: 37 future fixtures are present today, but that number is not the league’s final season total. A live feed grows and changes. Reproducible analysis requires stating what was available and when.
What would count as evidence?
A persuasive early pattern would be a sustained reduction in non-exempt stoppages, accompanied by stable recording of exempt medical cases, high protocol compliance and similar evidence quality across periods.
Several weaker patterns would require more caution:
- shorter stoppages but no change in incident rate could indicate faster management rather than deterrence;
- fewer annotated incidents alongside more missing video could be a data-quality artefact;
- a change concentrated in one team or referee group would not establish a league-wide effect;
- a shift in shots or results without a change in stoppage behaviour would not validate the proposed mechanism;
- no visible change in the first few weeks could simply reflect low statistical power.
The trial was approved for 2026–27; it is not yet a permanent change to the Laws of the Game. The evaluation should be equally provisional.
Limits that should stay attached to every result
The method still depends on human observation because the defining event is absent from the current structured data. Video availability, camera cuts and commentary can affect classification. Official guidance may evolve. Teams may adapt at different speeds. Promoted and relegated clubs change the season composition. A comparison competition may face different incentives or refereeing practices.
Most importantly, the trial acts on behaviour that is difficult to observe directly: the strategic purpose of a stoppage. The protocol can measure what happened on the field and how the rule was applied. It should not claim to read intent.
The best first result is therefore not a headline about whether the rule “worked.” It is a transparent incident dataset, a frozen definition, an audited denominator and a claim narrow enough to survive replication.
Trial protocol details: The Football Association. Authorization context: IFAB Circular 32. Match and coverage figures: Xtra-Stats data, audited 14 August 2026.
Disclosure: This article was prepared with AI assistance and reviewed against official sources and audited Xtra-Stats data.
Top comments (0)