DEV Community

Jordan Huang
Jordan Huang

Posted on

FAQ: A Drafted Retry Block Is Not a Flake Policy

A red job came back green after a single retry.
Someone pasted a model draft and called the flake gone.
I do not buy that story without a real job log.

The YAML looked tidy, and the chat looked confident.
The pipeline contract was still missing from the diff.
Have you merged a draft like that in the last month?

Why this FAQ exists

I review GitLab CI drafts more often than application code.
People ask whether a generated retry block is enough.
My answer is no, and I want the reasons on paper.

This FAQ is not a product tour or a benchmark post.
I am not reporting a flake rate from your estate.
I am correcting five claims I keep hearing in review.

Myth 1: A retry key means the flake is gone

The claim says a retry key made the job stable.
The evidence says retry only queues another attempt.
The corrected model says retry spends a failure budget.

A second attempt can pass for very boring reasons.
The network blip ended, or a cold cache warmed up.
The same bug can still sit in the same script.

GitLab documents retry as rerun control, not as diagnosis.
Read the keyword page for the version you actually run.
I will not pin a numeric max your docs may have changed.

Would you approve a deploy because the second try passed?
I would not, even when that second attempt was green.
A green retry is a clue, not a closed incident.

Myth 2: allow_failure keeps the gate honest

The claim says allow_failure lets you move without hiding risk.
The evidence says a true flag can hide a failed job.
The corrected model says that flag is a product decision.

Ask who must read the failure before you merge the draft.
If nobody must act, call the job advisory in the name.
If the suite must pass, that flag does not belong.

I keep seeing drafts mark the test job optional by habit.
That habit can ship a green pipeline over a red suite.
Is that the story you want on the merge request?

GitLab also documents failure-exit forms of the same flag.
Confirm the form against your docs before you copy one.
Search your CI YAML reference for allow_failure before merging.

Myth 3: A rules block means the job is protected

The claim says a rules block means the job is protected.
The evidence says rules only decide whether the job is created.
The corrected model says protection lives outside that snippet.

Models like a branch variable because it looks precise.
That rule can still create a deploy job on the wrong branch.
Did the draft name a protected environment, or only a branch?

I do not treat a rules snippet as an approval policy.
Approvals and environment tiers do that work in GitLab.
The YAML should point at those controls, not replace them.

Search the CI YAML reference for rules on your version.
Do not trust a copied if clause without that check.
A precise-looking rule can still be the wrong gate.

Myth 4: needs means the files will be there

The claim says a needs list guarantees the upstream files.
The evidence says optional needs can start you without that job.
The corrected model separates the graph edge from the tarball.

Look for optional true before you trust the drawn graph.
Look for artifacts on the producer, not only on needs.
A graph edge is not a tarball sitting in the job.

Would your job fail closed if the producer was skipped?
If the answer is no, the draft is optimistic.
Optimistic graphs are how empty reports reach main.

I reject optional needs on release jobs by default.
You can waive that if the job truly stands alone.
Write the waiver in the merge thread, not only in chat.

Myth 5: A chat pass is a pipeline receipt

The claim says the model ran the script, so CI will too.
The evidence says a scratch shell is not your job image.
The corrected model says only your runner log counts.

I am not re-teaching runner registration in this FAQ.
The narrower point is that plausible YAML is not a receipt.
A receipt is a pipeline URL on the commit you merge.

Do not paste CI secrets into a scratch shell to prove it.
A secret in chat is not a masked CI variable.
If the check needs secrets, run it inside the project.

A review table, not a verdict

I keep this table next to the diff, not in the chat.
It does not call GitLab, and it does not score quality.
It forces one hard question for each risky key.

Draft pattern Repeated claim Question I ask Reject when
retry as a bare number Flakes are handled Which failure class may rerun? Script bugs burn the same budget
allow_failure: true on tests The gate stays honest Who must read the failure? The suite can fail and main stays green
Broad rules:if on a branch The job is protected Which environment tier applies? Deploy is created with no protected target
needs with optional: true The graph is safe What if the producer is skipped? The job starts with missing reports
Chat transcript as proof The script already passed Where is the project job log? No pipeline URL on the commit

Use the table as a review aid, not as a verdict.
Your policy may be stricter than the rows I wrote.
Write the stricter rule down before you waive a row.

A proposed checker you can run

The script below is a proposal you can run yourself.
I am not claiming I ran it on your repository today.
I am not claiming a pass count, a runtime, or a delta.

It is a text scan, not a real YAML parser.
Anchors, references, and includes can hide the same risk.
A clean scan still owes a pipeline on your runner.

#!/usr/bin/env bash
# Proposal only. Not a GitLab CI linter.
# Not run against a private project for this article.
set -euo pipefail

file="${1:-.gitlab-ci.yml}"
if [[ ! -f "$file" ]]; then
  echo "missing file: ${file}" >&2
  exit 2
fi

echo "scanning ${file}"

if grep -nE '^[[:space:]]*retry:[[:space:]]*[0-9]+[[:space:]]*$' "$file"; then
  echo "REVIEW: numeric retry has no when filter"
fi

if grep -nE '^[[:space:]]*allow_failure:[[:space:]]*true[[:space:]]*$' "$file"; then
  echo "REVIEW: allow_failure true can hide a failed job"
fi

if grep -nE '^[[:space:]]*optional:[[:space:]]*true[[:space:]]*$' "$file"; then
  echo "REVIEW: optional needs may start without artifacts"
fi

if grep -nE 'CI_COMMIT_BRANCH' "$file"; then
  echo "REVIEW: branch rule is not an environment gate"
fi

echo "done: silence is not a merge approval"
Enter fullscreen mode Exit fullscreen mode

Grep inside an if does not abort the script on a miss.
Silence can also mean the regex missed a multiline form.
Read the REVIEW lines, then read the YAML yourself.

  1. A numeric retry line means you still owe a when filter.
  2. An allow_failure true line means you still owe an owner.
  3. An optional true line means you still owe a fail-closed test.
  4. A branch variable hit means you still owe an environment tier.

Commands I want beside the diff

Start from the diff, not from the chat transcript.
Then run the proposal script against the fixture and your file.
Compare REVIEW lines with the five questions before you approve.

git diff -- .gitlab-ci.yml
bash -n check-ci-draft.sh
bash check-ci-draft.sh fixture-retry.yml
bash check-ci-draft.sh .gitlab-ci.yml
Enter fullscreen mode Exit fullscreen mode

bash -n only checks shell syntax, not YAML meaning.
A syntax-clean script can still flag the wrong lines.
Read one flagged line in the file before you waive it.

Fixture that should trip the rows

The fixture is a teaching sample, not a customer file.
I wrote it to trip four review rows on purpose.
You can delete it after the questions make sense.

# Example fixture only. Not a production pipeline.
test_job:
  script:
    - pytest
  retry: 2
  allow_failure: true

deploy_job:
  script:
    - echo deploy
  rules:
    - if: $CI_COMMIT_BRANCH
  needs:
    - job: test_job
      optional: true
Enter fullscreen mode Exit fullscreen mode

The test job retries with a bare number and no when.
That matches myth one, because no failure class is named.
A script bug would burn the same budget as a blip.

The same job sets allow_failure to true on the suite.
That matches myth two, because the suite can fail quietly.
Would you want pytest optional on the branch you release?

The deploy job rules on a branch variable alone.
That matches myth three, because creation is not protection.
A branch name is not an environment tier.

The deploy job needs the test job and marks it optional.
That matches myth four, because the producer may be skipped.
Optional plus allow_failure is a hole, not a belt.

A narrower retry sketch

Here is a narrower sketch, not a config from a live project.
It names a small max and two infrastructure failure classes.
Confirm those when values in the docs for your GitLab version.

# Sketch only. Confirm when values for your GitLab version.
fetch_deps:
  script:
    - ./fetch-deps.sh
  retry:
    max: 1
    when:
      - runner_system_failure
      - stuck_or_timeout_failure
  allow_failure: false
Enter fullscreen mode Exit fullscreen mode

I would still want a pipeline URL before I trusted even this.
A better draft is still a draft until your runner executes it.
Does your review thread have that URL yet?

Questions I paste on the merge request

Five questions take less time than one bad deploy.
I still want them answered in the thread, not in chat.
A model can draft the answers, and you still own them.

  1. Which failure classes may retry, and which must stop?
  2. If this job fails, does the pipeline still go green?
  3. Which protected environment, if any, receives this deploy?
  4. If the producer is skipped, does this job fail closed?
  5. Where is the pipeline URL for this commit on our runner?

The corrected mental model

Separate four objects in your head before you merge.
The draft is text, and the scratch run is a foreign shell.
The project pipeline is evidence, and the approval is a person.

Where a free draft fits

MonkeyCode's free model access can help you draft the checker.
MonkeyCode's free server option can host a scratch run of it.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.

I do not know which model name you will be offered.
I do not know the quota, the image, or the lifetime.
Those gaps are why a scratch run cannot close review.

Draft the checker on that free server if a scratch shell helps.
Then run the same file on a registered runner you control.
Paste the job URL on the merge request, and stop there.

Who should not use this

Skip this checker if policy lint already blocks these keys.
Skip it if your real risk is secrets leaking in logs.
Skip it if you need full simulation of includes and child pipelines.

Do not use a free server as the place you unlock production variables.
Do not use a model draft as the only reviewer on deploy.
Do not treat this FAQ as compliance or vendor documentation.

Limits I will say out loud

The scan misses multiline maps and hidden YAML anchors.
A branch-variable hit can be a comment, so read the line.
Keyword behavior can change, so your docs beat my memory.

I did not invent a pass rate or a time box.
If a product page states those facts, read that page instead.
If the page is silent, you do not have the promise.

Open the CI/CD YAML reference for the GitLab version you run.
Search retry, allow_failure, rules, and needs before you copy.
I am not pinning a docs URL that may move out from under you.

Before you call it fixed

A drafted retry block can be a decent starting note.
It is not a flake policy, and it is not a receipt.
Ask for the job log before you call the myth fixed.

Top comments (0)