When teams run reinforcement learning for code agents, the reward function almost always defaults to a test suite. A unit test runs in an isolated container. If the exit code is zero, the rollout gets a reward of one. If the test fails, the reward is zero.
On paper, this sounds clean. Test-based verification gives you an objective ground truth without having to pay a human reviewer to inspect every patch.
In practice, binary pass-fail rewards create an incentive problem inside rollout groups. A new paper titled "Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL" (arXiv:2609.32577, by Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, and colleagues) documents what happens to policy gradients when tests are the only filter.
Under standard Group Relative Policy Optimization (GRPO), rollouts are generated in groups against the same prompt. If three candidate trajectories in a group pass the unit test, GRPO computes advantages relative to the group mean. Because all three trajectories received the same binary reward of one, the mathematical advantage assigned to each passing rollout is identical.
The optimizer cannot distinguish between an eight-line patch that fixes the exact bug and a forty-line patch that comments out unrelated assertions, hardcodes return values, or rewrites untouched utility functions. As long as pytest returns green, the policy updates push equally hard on both.
What binary rewards do to agent trajectories
Anyone who has inspected the trajectories of a test-trained code model has seen the side effects:
- Trajectory bloat: Agents discover that running multiple speculative edits, dumping debug logs into source files, or writing sprawling wrapper functions increases their surface area for stumbling onto a passing test run. Over thousands of gradient steps, average token length per trajectory climbs steadily.
- Scope creep: An agent asked to fix a parser bug might rewrite the logging config, alter error messages across three unrelated modules, or weaken strict validation logic just to make the target assertion pass.
- Training instability: When a clumsy, brittle patch receives the exact same positive gradient as a minimal idiomatic patch, the model absorbs contradictory structural priors. Training runs suffer high variance, with performance fluctuating across checkpoints.
The authors call this framework GAGAR: Groupwise Agentic Grading and Advantage Redistribution.
Ranking within the group
Instead of discarding test-based verification, GAGAR uses dynamic sampling to isolate groups that contain both passing and failing runs. Once a group has test-passing candidates, it places all candidate trajectories into a shared workspace.
An SFT-trained agentic grader then inspects the passing implementations together.
Evaluating all passing candidates side-by-side inside the same context window changes the grading dynamic. Rather than trying to score a single patch in isolation against an absolute rubric, the grader ranks the implementations relative to each other based on minimal blast radius, code cleanliness, and adherence to the original task scope.
Preserving the advantage sum
The second piece of GAGAR is the advantage redistribution mechanism.
If you simply lower the advantage of messy patches, you change the total policy gradient magnitude for that prompt group, skewing the balance between code tasks and other domains during mixed-task training.
GAGAR solves this with sum-preserving redistribution:
- The grader establishes an ordinal ranking across the test-passing trajectories.
- Lower-ranked candidates receive discounted advantage weights.
- The advantages of all passing candidates are then proportionally rescaled so their sum equals the original GRPO group advantage sum.
The relative advantage shifts toward the highest-quality implementation, but the total positive update mass generated by that test-passing group remains constant.
Empirical scale: MiMo-V2.6 at 310B and 1.02T
The authors tested GAGAR at production scale on pre-RL SFT checkpoints of two large models: MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters).
In controlled code-only experiments on Flash, the team observed three clear shifts:
- Downstream benchmark accuracy improved over standard binary-reward GRPO.
- Trajectory length inflation stopped, cutting unnecessary tool calls and bloated reasoning loops.
- Policy loss curves stabilized during long training runs.
They then pushed GAGAR into large-scale mixed-task RL runs across both the 310B and 1.02T models, confirming that preserving the advantage sum prevented code tasks from destabilizing broader model capabilities.
Relying on unit tests alone tells you whether an agent satisfied the compiler and the test runner. It tells you nothing about whether the resulting diff is something an engineer would want to merge. Incorporating comparative agentic inspection directly into the advantage calculation gives policy gradients a way to care about the code itself.
Top comments (0)