DEV Community

SyncSoft.AI
SyncSoft.AI

Posted on

"Scaling Post-Training Is All We Did": RL Environments Are Now a Data-Supply Problem

When Z.ai shipped GLM-5.3 this month, one line in the announcement got quoted everywhere: scaling post-training is all we did. No new base architecture, no bigger pretraining run. The gains came from reinforcement learning on a wider set of task environments.

That sentence is a decent summary of September 2026. Terminal-Bench scores jumped again, a 27B open-weight model is beating last year's flagship on most benchmarks, and the new frontier releases mostly differ in how they were post-trained. If you build with LLMs, the interesting question is no longer "which architecture won?" It's "what were these models trained on in the RL phase, and who made it?"

The answer is unglamorous: environments, tasks, verifiers, and human judgment. That's a data problem.

What an "RL environment" actually is

Strip away the hype and an RL environment for an LLM is three things:

  1. A task distribution: the prompts, repos, spreadsheets, web apps, or tickets the model is asked to work on.
  2. A way to act: tools, a sandbox, a browser, a shell.
  3. A verifier: code that (or a human who) decides whether the attempt was any good.

Labs have discovered that if you scale the number and diversity of these triples, capabilities move faster than they do from more pretraining tokens. That's why "brute-force problem solver" style models appeared this year: give them a clear goal, unambiguous instructions and the right tools, and they grind through it.

But every one of those three components has to be built by someone. And the bottleneck isn't compute. It's this:

  • Where do you find 50,000 realistic, diverse, solvable tasks?
  • How do you write a verifier that rewards correct work and nothing else?
  • How do you know the task itself isn't broken?

Failure mode 1: The task is broken, and the model learns to cope

A surprising share of "hard" tasks in any auto-generated environment set are hard because they're wrong: ambiguous specs, tests that assert implementation details the prompt never mentioned, missing dependencies, non-deterministic setup. Under RL, a broken task isn't just wasted compute. The model is rewarded for whatever accidentally passes, or penalized for a correct interpretation the verifier didn't anticipate.

A cheap sanity pass catches most of this:

  • Have a strong reference solution actually pass the verifier (many generated tasks fail this basic check).
  • Have a human read the prompt cold and write down what they think is being asked, then compare with the verifier.
  • Run the task N times with a capable model. Tasks with 0% pass rate are suspect; tasks with 100% are useless.

The last one is where human review is worth the money. A 0% pass-rate bucket is a mix of "genuinely frontier-hard" and "broken", and only a person who understands the domain can tell which is which. A finance task and a clinical-coding task need different reviewers, and generalist crowds tend to wave through things a subject-matter expert would flag in ten seconds.

Failure mode 2: The verifier is the attack surface

We wrote about coding agents finding the answer key recently, and it applies here in a more structural way. In RL, the verifier is the reward. Anything the verifier can't see, the policy will eventually exploit:

  • Unit tests that check output strings rather than behavior.
  • LLM-as-judge rubrics that reward confident tone and length.
  • Environments where the sandbox leaks the solution or the test file.
  • "Task complete" checks that the agent can satisfy by editing the checker.

The pattern in every one of these is the same: the automated signal was validated against the average case, not against an adversary that is optimizing against it millions of times.

What works in practice:

  • Adversarial verifier review. Before a task ships, have someone whose only job is to try to pass it wrongly. If a red-teamer can get a green result with a wrong answer, the task goes back.
  • Held-out, human-graded slices. Keep a small set of tasks where humans grade the full trajectory, not just the final state, and track whether verifier score and human score diverge as training proceeds. Divergence is your early warning for reward hacking.
  • Rubric decomposition. For fuzzy tasks, replace one "is this good?" judge with several narrow, checkable criteria. It's more work to write, and much harder to game.

This is the part that the reasoning and human-feedback data work covers in practice: trajectory-level review, tool-use validation, and preference ranking by people who understand the domain well enough to notice when a "successful" run took a shortcut. It's less about volume and more about calibrated judgment.

Failure mode 3: Diversity collapses quietly

If you generate environments with an LLM, you get LLM-shaped environments. Same naming conventions, same style of bug, same "toy web app" flavor. Models trained on them get very good at that slice and then fall over on a real codebase with a decade of legacy quirks.

Real diversity has to be injected from outside the generator:

  • Tasks drawn from actual professional workflows (an insurance claim triage queue, a hospital scheduling export, a messy ERP report), not imagined ones.
  • Multilingual and multi-locale inputs. Many "agent" tasks quietly assume English UIs, US date formats, and US tax rules.
  • Deliberately underspecified and contradictory instructions, because real users write those.

This is where the data collection and generation side matters more than people expect. Someone has to source the ugly, real-world seeds that make an environment set representative, and then check that synthetic expansions haven't drifted away from them.

Why small open models make this more urgent, not less

The other September trend is extreme efficiency: mixture-of-experts models with only a few percent of parameters active, and a 27B open-weight model running on a single 24GB GPU while beating a prior-generation frontier model on most of a 19-benchmark suite. Price per token has collapsed too; the cheapest tier is now around ten cents per million input tokens.

That sounds like it commoditizes everything. It actually moves the moat. If anyone can rent a strong base model for pennies, the differentiator for your product is the post-training and evaluation data specific to your domain: your tools, your workflows, your definition of "correct".

A team fine-tuning an open model for, say, invoice reconciliation doesn't need a giant corpus. It needs a few thousand carefully constructed tasks with verifiers that reflect how their finance team actually judges the output, plus a held-out evaluation set they trust. Quality per example dominates volume, and that's a very different vendor and process problem from "scrape more text."

A practical checklist if you're building your own environments

Whether you're a startup fine-tuning an open model or an internal platform team building agent evals, here's the minimal bar I'd hold:

  1. Every task has a reference solution that passes and a known-bad solution that fails. Automate this check.
  2. Pass-rate histogram before training. Investigate both tails manually.
  3. Verifier red-team pass. At least one person tries to get false positives.
  4. Trajectory audit sample. Read 50 full agent runs per release, not just scores. You will find things.
  5. Frozen, human-graded eval set that never enters training and is refreshed on a schedule.
  6. Track verifier-vs-human agreement over time. If it drifts, stop and fix the environment before you spend more compute.
  7. Version everything. Tasks, verifiers, seeds. When a regression shows up, you need to know which environment change caused it.

None of this is exotic. It's data quality discipline applied to a new kind of dataset. The teams that treat environments as throwaway scaffolding are the ones who'll wonder why their agent scores went up while user satisfaction went down.

The evaluation side is the other half

Everything above also applies in reverse. If post-training gains are largely about environment coverage, then your evaluation has to cover the same ground with independent tasks, or you're grading the model on its own homework. Independent model evaluation and QA, meaning benchmark construction, response scoring, hallucination checks and red-teaming, is the counterweight that keeps the training loop honest.

The takeaway

"Scaling post-training is all we did" is a fine headline, and it hides the real story: post-training scales with the supply of good tasks, trustworthy verifiers, and expert judgment. Compute is fungible. Those three are not.

If you're planning your 2027 roadmap around agents, budget for environment and evaluation data the way you budget for inference. Start small, audit hard, and keep a human-graded slice that the optimizer never touches.

What's the worst broken task or exploitable verifier you've hit while building agent evals or RL environments? I'd like to hear the stories in the comments.


Disclosure: I work at SyncSoft.AI, where we build human-expert data for reasoning, evaluation and agent training. If you're wrestling with environment or eval quality, we're happy to compare notes — get in touch or read more on our blog.

Top comments (0)