DEV Community

Jordan Huang
Jordan Huang

Posted on

I Do Not Trust a Free Bench With CI Secrets

I Do Not Trust a Free Bench With CI Secrets

I keep seeing one shortcut in merge review threads.
A failing job gets pasted into a free chat.
The suggested patch then runs on a free host.

The author called that scratch pair a proof.
Did that host ever hold your masked variables?
Would you ship if the honest answer is no?

The smell in the thread

Speed feels like evidence when the clock is loud.
A free model answers before your runner queue moves.
A free server can compile while you wait.

That comfort is still not a security boundary.
I want the shortcut as much as anyone.
I still refuse to treat it as approval.

What I refuse to invent

No private quota sheet sits on my desk.
I will not name models I cannot verify.
I will not invent hardware, duration, or scores.

The operator supplied two availability claims only.
Free model access is available for this workflow.
A free server option is available as well.

Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode is only the bench inside this method.
Remove that name and the checks still stand.

Myth: the chat box is a vault

People repeat this after one lucky session.
They say the free model is private enough.
Then they paste the full failing job log.

Is a chat transcript really your variable store?
Did you read retention before you pasted tokens?
Can you delete a prompt already stored upstream?

I treat every prompt as a one-way copy.
A mask is a log filter, not a vault.
Pasting the text outside GitLab drops that filter.

The corrected model treats prompts as export surfaces.
If a secret shape appears, I rotate it.
I do not argue with the log after that.

Evidence I can actually check

  • Read the vendor retention note before pasting secrets.
  • Search the pasted text for TOKEN, KEY, and PASSWORD.
  • Rotate any value that showed up in the paste.

Myth: the free server is your runner

A free host can compile a branch snapshot.
That does not make it your project runner.
Who registered that host inside your group?

This free server is not a GitLab shared runner.
Please do not register it as a group runner.
Which protected tags can that host actually reach?

GitLab binds protected variables to protected refs only.
Confirm that rule in the current variables docs.
A vendor bench does not inherit those settings.

A scratch host outside that ref should see nothing.
A bad export can still hand it everything.
I don't point production jobs at a bench.

I don't store deploy keys on that disk.
The corrected model calls the free server disposable.
It may sketch a build and then vanish.

It must never hold your real release credentials.
Shared runners follow your project's own CI settings.
A disposable bench follows whatever you upload today.

Myth: a scratch success is a recorded run

The scratch build printed a friendly success line.
Your real pipeline never recorded that scratch run.
Where is the job id a reviewer can open?

Where is the artifact checksum in the registry?
A hosted success is only a debugging clue.
A clue is not an attestation you can audit.

I want a job URL before I approve anything.
I want the same script text committed in git.
The corrected model separates clues from records.

Myth: generated YAML inherits protection

Models love to emit a tidy job block.
They rarely emit your protection rules with it.
Did the sketch limit itself to a scratch branch?

Did it mask every variable it just invented?
I diff generated YAML against the living file.
I reject any job that echoes env wholesale.

Bare printenv can spill unmasked job variables fast.
The corrected model says protection is repo state.
A generated block does not inherit your rules.

Myth: you can promote the bench quietly

Someone copies the working command straight into main.
No second person reads the variable scope first.
Would you merge a runner you cannot name?

Would you merge a prompt you cannot delete?
Promotion needs a reviewed commit, not a vibe.
The bench stays behind the reviewed branch boundary.

The corrected model keeps promotion boring and visible.
Boring is exactly what I want near secrets.
Quiet promotion is how leaks skip review.

A check you can run on dummy input

This script is a proposal, not a lab score.
I have not published any timings for it.
Run it on a fake env file first.

Do not point it at real production secrets.
Dummy values belong in the sample file only.
Real tokens stay inside protected CI variables.

Copy the proposal below into a local script file.
Keep the file name stable inside your repo.
Mark the example values as dummy before you run.

#!/usr/bin/env bash
# Proposed local check. Not a recorded benchmark.
# Strict on purpose. Innocent curl lines fail closed.
set -euo pipefail

env_file="${1:?need a dummy env file}"
job_file="${2:?need a job script path}"

if grep -E 'TOKEN=|SECRET=|PASSWORD=|KEY=' "$env_file" | grep -vE '=dummy|=changeme|=example' >/dev/null; then
  echo "exit 2: env file looks like real secrets"
  exit 2
fi

if grep -nE 'printenv' "$job_file"; then
  echo "exit 3: job script can print the environment"
  exit 3
fi

if grep -nE 'curl ' "$job_file"; then
  echo "exit 4: review every curl before a hosted run"
  exit 4
fi

echo "exit 0: dummy scan passed only"
Enter fullscreen mode Exit fullscreen mode

Here is a dummy env file for the first run.
Do not replace dummy with a live token.
A live token does not belong in git.

APP_ENV=scratch
DEPLOY_TOKEN=dummy
CI_JOB_TOKEN=changeme
Enter fullscreen mode Exit fullscreen mode

This job sketch is the kind I reject.
It prints the environment and a token value.
Nothing in this sketch is safe to promote.

leak_job:
  script:
    - printenv
    - echo "$DEPLOY_TOKEN"
Enter fullscreen mode Exit fullscreen mode

This tighter sketch still needs a human diff.
The branch name is a label, not a lock.
Rules in the repo still have to match.

scratch_job:
  rules:
    - if: $CI_COMMIT_BRANCH == $SCRATCH_BRANCH
  script:
    - bash scripts/boundary-check.sh .env.scratch scripts/job.sh
    - make test
Enter fullscreen mode Exit fullscreen mode

How I read the exit codes

Exit two means the env file looks real.
Stop and move those values out of the file.
Exit three means the job can print secrets.

Rewrite that script before any hosted run.
Exit four means a secret may leave the host.
Delete that URL line and review the diff.

Exit zero only means the dummy scan passed.
It does not mean your pipeline is safe.
Say that out loud before you celebrate it.

Commands I actually type

I run these commands from the repo root.
I expect a non-zero exit on the bad sketch.
I expect exit zero only after the rewrite.

cp examples/env.scratch .env.scratch
bash scripts/boundary-check.sh .env.scratch scripts/job.sh
git status --short
git diff -- .gitlab-ci.yml scripts/job.sh
Enter fullscreen mode Exit fullscreen mode

Walk the change in eight moves

  1. I reproduce the failure with dummy variables only.
  2. I paste the error text, never the secret values.
  3. I ask the free model for a patch sketch.
  4. I run that sketch on the free server.
  5. I copy only the reviewed diff back into git.
  6. I run the boundary check on the new job.
  7. I open a merge request against the real pipeline.
  8. I stop using the bench after the sketch lands.

If a check fails, I rotate before I retry.
I do not paste the new log into chat.
I file the rotation note on the merge request.

A small decision table

I use this table when a review gets noisy.
Each row is a claim, a check, and a model.
If I cannot run the check, I do not merge.

Repeated claim Check I can run Corrected model
Chat is a private vault Read retention, then search the paste Prompts are export surfaces
Free server is my runner Name the registration and protected refs The bench is disposable
Scratch success is recorded Find a pipeline job id A clue is not a record
Generated YAML is protected Diff rules, masks, and echo lines Protection lives in the repo
Promote the working host Name the runner and the secret scope Promotion stays reviewed

Who should skip this

Skip this if your job needs production data.
Skip this if policy forbids third-party prompt uploads.
Skip this if the fix touches deploy keys.

Skip this if you cannot rotate a leaked value.
A regulated release needs your own named runners.
A public demo repo can use the scratch path.

Limits I keep taped to the monitor

Free access can change without a blog post.
I do not treat today's bench as permanent.
The script catches obvious leaks, not clever ones.

A base64 wrap can still hide a secret.
Reviewers still have to read the full diff.
No tool fully replaces that careful human pass.

GitLab mask rules are not a prompt filter.
Confirm current mask behavior in the product docs.
I will not cite a version I have not opened.

Where the bench still earns its keep

I still use a free model for the sketch pass.
I still use a free server for the scratch compile.
The value is speed before the real boundary.

The real boundary stays in your reviewed pipeline.
If you already keep a MonkeyCode bench, use it here.
Keep that bench on the scratch side only.

So which boundary failed in your last green run?
Was it the prompt, the host, or the missing job?
I check those three before I click merge.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to