Part 1 explained how PPO, GRPO, DAPO and GDPO turn several rewards into one learning signal. Part 2 benchmarked the trainers on a 14B model for 50 steps and found a trap: CISPO won the training curves and lost on held-out tasks. This part asks what happens when the same ideas get a bigger model and twelve times more steps.
We trained Qwen3.8-27B for 600 steps on a single-answer task scored by 12 reward channels, and climbed a controlled ladder: DAPO, then CISPO, then CISPO with REPO-R, then fewer but better prompts. Then we moved the winning recipe to a harsher environment and found a guard in our own advantage pipeline that silently threw away 96% of the learning signal.
The result in one line: same model, same rewards, same 600 steps. Held-out score went from −0.47 with DAPO to −0.02 with CISPO, to +1.33 with CISPO + REPO-R, and to +1.80 with a smaller, filtered prompt set. The starting model scores −1.37.

One training step in our trainer. Stages 2–5 are published methods or a small guard of ours; none of them changes the verifier or the rewards.
Start here
See what GDPO, CISPO and REPO-R each change inside one training step, how much each moved the held-out score of a 27B model, why identical training curves can hide a 2× holdout gap, and how a sensible-looking advantage guard can stop learning completely when the environment gets harder.
- Advantage (A): How much better an answer scored than the other answers to the same prompt. Positive A makes its tokens more likely; negative A makes them less likely.
- GDPO: Normalizes each reward channel inside the prompt group before adding them with priorities, so a channel with a wide numeric range cannot drown the others.
- CISPO: A policy loss that caps each token's importance weight at 1.2 instead of switching the token off when its probability has already moved.
- REPO-R: Rescales each token's advantage by how rare the token is, with a strength ζ set by an entropy thermostat.
- Advantage floor: A guard that forbids positive advantage for answers whose core score is at or below a threshold.
- Holdout score: The score of a frozen checkpoint on prompts and data the trainer never saw. Higher is better; zero is break-even.
1. Where Part 2 left us
On a 14B model and 50 steps, CISPO produced the highest training reward and one of the worst held-out scores. Adding REPO-R recovered a little (+0.0175), not enough to beat the untrained thinking baseline. Two questions stayed open: does the picture change with a larger model and a longer run, and which part of a combined recipe actually does the work?
So this time we changed one thing per run and judged every run the same way: a frozen held-out probe, and a checkpoint chosen without looking at it.
2. The setup, abstracted
We keep the domain out of this post on purpose; the reward structure is what matters. The policy reads a short specification (one of 500 prompts) and writes one structured answer of at most 128 tokens. A deterministic verifier parses it, runs it on a frozen dataset and returns twelve reward channels:
- Core score: a cross-validated estimate of answer quality. This is the objective; GDPO priority 3.0.
- Stability check: penalizes answers whose in-sample and out-of-sample ranks disagree; priority 1.0.
- Action-cost penalty: every change in the answer’s output costs a fee.
- Nine guards and shapers: a secondary predictiveness score, complexity, format, syntax and behavior novelty, a neutrality gate, a minimum-activity guard, a pattern guard and a stress test.
Held fixed in every run: Qwen3.8-27B in NF4 with a rank-32 adapter (alpha 64, FP32 master weights; LoRA unless a row says DoRA), no thinking, 4 prompts per step with 12 answers each, temperature 1.0, learning rate 1e-5, KL β 0.08, GDPO, 600 optimizer steps, one H100, seed 42. The REPO-R rows also change two loss settings, described in section 5.
Holdout: every 100 steps, a frozen probe of 50 unseen prompts is scored on a later slice of data that training never touched. Each run “ships” the checkpoint with the best train-side core score on the same probe, so the holdout number is never used to pick it.
3. GDPO in one minute
With twelve channels on very different scales, plain GRPO adds them first and normalizes the sum, so the widest channel decides almost everything. GDPO standardizes each channel inside the prompt group first, then adds them with explicit priorities. Every run in this post uses GDPO, with priorities core 3.0, stability 1.0, pattern guard 0.3, complexity 0.2 and 1.0 for the rest; Part 1 covers the math and an interactive demo.
GDPO decides which answers count as better. The next two ideas decide how hard each token moves once we know.
4. CISPO: cap the weight, don’t drop the token
An answer was sampled by a slightly older copy of the policy. For each token, the ratio r = π_new / π_old says how much more likely the current policy makes it. PPO, GRPO and DAPO keep updates safe by clipping: once a token’s ratio moves past the band in the direction its advantage pushes it, its gradient becomes zero for that step. CISPO (from the MiniMax-M1 report) keeps every token in the gradient and only caps its importance weight; we use a cap of 1.2.
DAPO: loss = −min(r·A, clip(r, 0.8, 1.28)·A)
CISPO: loss = −stopgrad(min(r, 1.2)) · A · log π(token)

Gradient weight of one token as its probability moves. The clipped objective switches the token off past the band; CISPO keeps it in the gradient and only caps its importance weight at 1.2.
Why this matters for short structured answers: a handful of “fork” tokens decide what the answer does. Those are exactly the tokens whose probability moves fastest when a good answer is reinforced, so a clipped loss switches them off first. CISPO keeps teaching them, just more quietly.
Measured effect: swapping DAPO for CISPO, with everything else equal, moved the held-out score from −0.47 to −0.02, a gain of +0.45. In Part 2, CISPO overfit at 14B; here, with checkpoint selection by train-side core score and a 27B model, it generalized better than DAPO.
5. REPO-R: an entropy thermostat
RL tends to collapse a policy onto a few answers it already likes. Entropy, a measure of how many alternatives the policy still considers, falls, and exploration stops. REPO-R counters this per token: in a better-than-average answer, rare tokens get more credit; in a worse-than-average answer, rare tokens get less blame.
A > 0: A′ = A · (1 − ζ · log p)
A < 0: A′ = A · (1 + ζ · log p)
where p is the token’s probability and log p < 0

REPO-R's advantage multiplier: rare tokens get more credit in good answers and less blame in bad ones. At the 0.05 cap a 1-in-1,000 token gets ×1.35 or ×0.65; at our median ζ the effect is under 1%.
The strength ζ is not a fixed knob. A controller records the average entropy of the first five updates as its target (the “w5” in our run names). When entropy falls below target, ζ doubles; when it rises above, ζ halves, and it never goes below zero. Here is the real log of our best run:

Real controller log of the best run (1790958861). ζ sleeps near zero, doubles up to 0.05 while entropy is below target, then halves back. Median ζ = 0.00078; above 0.01 on 27% of steps.
Measured effect: adding REPO-R to CISPO moved held-out from −0.02 to +1.33, a gain of +1.35, the largest single step in the study. A caveat we take seriously: the REPO-R configuration also switches on full-token loss (every token enters the loss, not only the 20% with the highest entropy) and two optimizer passes per rollout batch. With ζ this small most of the time, part of the gain may come from those two changes. The ablation that would separate them is still on our list.
6. Results at 27B
CISPO beat DAPO
One-line loss change, everything else fixed. Part 2’s 14B, 50-step benchmark had favored DAPO.
REPO-R bundle: biggest step
Includes full-token loss and two passes per batch, which we have not yet isolated.
Fewer, better prompts
Keeping only prompts where the starting model sometimes, but not always, succeeds gave the best run.
Same curves, 2× holdout gap
With REPO-R and DoRA on both, the training curves overlap. Only the holdout separates them.

Holdout score of the shipped checkpoint at each rung, chosen by train-side core score, not by holdout. One seed per configuration.
| Step of the ladder | Configuration | Shipped step | Holdout | Change | Versus |
|---|---|---|---|---|---|
| Start: DAPO loss | DAPO · LoRA · 500 prompts | 600 | −0.47 | — | starting model: −1.37 |
| Swap the loss for CISPO | CISPO · LoRA · 500 prompts | 600 | −0.02 | +0.45 | DAPO · LoRA · 500 prompts |
| Add REPO-R (with full-token loss, 2 passes) | CISPO + REPO-R · LoRA · 500 prompts | 600 | +1.33 | +1.35 | CISPO · LoRA · 500 prompts |
| Train on 226 selected prompts instead of 500 (best holdout) | CISPO + REPO-R · LoRA · 226 prompts | 500 | +1.80 | +0.47 | CISPO + REPO-R · LoRA · 500 prompts |
| Swap LoRA for DoRA | CISPO + REPO-R · DoRA · 226 prompts | 500 | +1.33 | −0.47 | CISPO + REPO-R · LoRA · 226 prompts |
| Swap CISPO back to DAPO (keeps REPO-R, DoRA) | DAPO + REPO-R · DoRA · 226 prompts | 500 | +0.69 | −0.64 | CISPO + REPO-R · DoRA · 226 prompts |
DoRA did not help once REPO-R was on: +1.33 against +1.80 for the same recipe with LoRA. Without REPO-R it gave DAPO a small lift (−0.29 against −0.47). The best run peaked at step 500 and sagged slightly by step 600, which is why we ship by the train-side selector rather than the last step.
Training curves are not a verdict

Same adapter, prompts and REPO-R; only the loss differs. The training curves overlap, the held-out scores do not: +1.33 for CISPO, +0.69 for DAPO.
This is Part 2’s lesson again, at a larger scale: a trainer can look identical on the training reward and still learn something quite different. Every decision in this post is made on a frozen holdout, and every run uses the same probe and the same selection rule.
Every training curve
The holdout chart is what we decide on; the training curves are what the trainer saw. In Environment A, the REPO-R runs stay close on training reward and core score while their held-out scores spread from +0.69 to +1.80. Every run’s per-step values are in the JSON export linked at the end.
7. The advantage floor trap
Next we moved the winning recipe to Environment B: different data, the same verifier, and per-action costs about 19× higher. Same 27B model, same GDPO, CISPO and REPO-R, and the same prompt filter, which this time found so few prompts with any successful answer that it kept the 100 closest to break-even. The stress-test channel is off there: its world model failed its own quality gates on that data. The runs started, the curves wiggled, and almost nothing happened.
The cause was a guard we had added to GDPO ourselves. It keeps invalid answers, and answers whose core score is at or below a floor (0), from receiving positive advantage, while preserving the rule that advantages inside a group sum to zero. In Environment A this is a sensible safety net: good answers score above zero. In Environment B, the fresh policy almost never does, and only about 1% of answers clear the floor. A group in which all twelve answers are below the floor gets twelve zeros.
The logs made the damage plain. With the floor, 4% of answers carried any learning signal; KL to the starting model stayed around 0.02. Removing the floor, and changing nothing else, gave 100%, a KL rising to about 0.3, and a training reward that climbed from about −10 to −2.5.

Environment B, first data slice, one change: the advantage floor. Ten-step means; the no-floor run was still training at 339 of 600 steps.
All seven Environment B runs, across four data slices, tell the same story: every floor-at-zero run stays flat, every no-floor run climbs.

All seven Environment B runs across four data slices: with the floor at 0, 3–4% of answers carried any signal and reward stayed flat; without it, nearly all did and reward climbed.
Turning the guard off did not turn it off
Our first fix was to disable the guard. The trainer then falls back to a second safety net that clamps any positive advantage to zero when the core score is at or below zero: the same trap with a different name. What worked was keeping the projection, which still forces invalid answers to non-positive advantage, and moving the floor far below the environment’s score range. A floor is an assumption about the reward distribution. When the environment changes, the assumption has to be checked again.
Even in Environment A the floor was not free: 10–41% of advantages were zero across runs, more in the runs that learned slowly. Whether a lower floor would also help there is one of our next experiments. The held-out results for Environment B will follow once those runs finish their probes.
8. What we would do again
- Log the share of non-zero advantages from step 1. In a healthy run it sits well above half. At 4% the trainer is idling while the GPU bill keeps running.
- Treat every guard threshold as a property of the environment. Re-derive floors and clamps when costs, data or reward ranges change, and check what the fallback path does when you disable one.
- Change one thing per run and judge on a frozen holdout. Ship the checkpoint chosen by a train-side selector, and use the holdout only to verify it.
- Prefer fewer, informative prompts. Prompts the starting model always or never solves give little signal; filtering them beat training on all 500.
- Name the recipe by its parts. GDPO + CISPO + REPO-R is a combination of published methods plus one guard of ours, and it should be credited that way.
Limits: one seed per configuration; one held-out slice; REPO-R measured as a bundle with full-token loss and two passes; Environment B holdout pending. Treat the gaps as strong signals, not final effect sizes.
Evidence and method references
The JSON export holds every logged optimizer step (reward, core score, cost, KL, entropy), the share of non-zero advantages per step, the REPO-R controller log and the holdout score of every checkpoint, with run IDs. The Markdown companion has the tables.
Methods: DeepSeekMath / GRPO; DAPO; GDPO; MiniMax-M1 / CISPO; REPO-R.
Applying this to your own models? We post-train open models on teams' own tasks with SFT and GRPO, and hand back the weights. LLM post-training with g factor.
Originally published at g-ftech.com.
Top comments (0)