Open RL environments are great for research. They also open another route for benchmark overlap. Here is a small, reproducible case, plus one pitfall that will bite anyone checking for it.
(Short version: this does not show that any reported score is inflated; caveats below.)
The finding
Xiaomi released MiMo-V2.6-RL-oss, a set of open RL environments. Its cyber config has 1,000 tasks in which an agent must craft an input that triggers a specific known bug in real open-source code. CyberGym is a benchmark for the same skill on 1,507 bugs.
Both draw on ARVO, a public collection of reproducible OSS-Fuzz bugs: 1,368 of CyberGym's tasks come from ARVO, and the rest are newer bugs taken from OSS-Fuzz directly. So I checked the overlap:
- 223 of CyberGym's 1,507 tasks (14.8%) are bugs in the MiMo cyber set. They correspond to 278 of its 1,000 tasks.
- As a check that these are the same crashes: for 215 of the 223, the crash type and function in MiMo's prompt match CyberGym's ground-truth sanitizer report. For same-project pairs of different bugs, that happens 7.7% of the time.
What this does not show: that any reported score is inflated. The task rows contain no reference PoCs. In the 26 of the 1,000 task Docker images I sampled, the task-specific files are just the vulnerable binary and source (plus a fixed build in 5, see below), and none outside the source tarball is named like a PoC. I didn't list the files in the other 974 images, or look inside the source trees. I don't know whether or how these tasks were used to train the MiMo-V2.6 models. The tech report's RL experiments on these released environments (§7.2, Table 6) evaluate cyber on its in-house MiMo Cyber Bench (mini), not CyberGym. The count also doesn't cover every overlap: 223 counts identical OSS-Fuzz issues only. 20 more CyberGym tasks are separate issues related to a MiMo bug, by a shared fix commit or an identical crash signature (10 of them through two signatures that are common in one project); I list them but don't count them.
What it does mean: if you RL-train on this config and then report CyberGym, about one in seven test bugs is one your model practised on.
Is that more than chance? All of CyberGym's ARVO tasks come from ARVO's first release (4,993 bugs), and 710 of MiMo's distinct bugs fall in that pool. A uniform random draw of 710 bugs would share about 195 ± 11 with CyberGym. I see 219: z = 2.2, one-sided p = 0.015, so modestly above chance. Drawing within each OSS-Fuzz project instead gives 199 ± 9 (z = 2.2, p = 0.017), so it isn't just both sets favouring the same projects. The other 4 of the 223 are in CyberGym's newer OSS-Fuzz slice. Some overlap is what you get when neither set is filtered against the other. The modest excess could come from both sets favouring similar bugs within a project; I don't know, and it doesn't show how either set chose its bugs.
The task images
Each task ships as a Docker image, and Docker Hub lists every image's build steps. In 135 of the 1,000 images, the build also copies a second binary, fix_binary, which by its name is built from the fixed code. An agent that could read it could compare it with the vulnerable build to find the fix. These are exactly the 135 tasks whose ID is also a CyberGym arvo: task ID. None of the other 865 has one. That includes the other 139 overlapping tasks, which carry the bug's new tracker ID while CyberGym uses the old one (explained in the next section; not the same 139 as the plain ID join there), so the pattern follows the ID, not the bug. All 135 are already among the 278 overlapping tasks, so this adds no overlap.
Fixed builds are a common grading input for this kind of task: CyberGym, for example, counts a PoC only if it crashes the vulnerable build and not the fixed one. In the 26 images I listed (5 of them with a fixed build), the grader inside the image doesn't use it: it is the same file in all 26, and the only binary it names is the vulnerable one. I didn't read the grader in the other 130 images with a fixed build. The fixed build sits in a root-only directory, and the agent's files belong to a non-root agent user, so the images look meant for a non-root agent. But the image doesn't set which user the agent runs as, and I didn't check what runs outside the image. If you use these images, run the agent as a non-root user. I don't know why these 135 were built differently, and I've asked the dataset authors. This doesn't show that these tasks were picked because they are in CyberGym: of the 445 tasks with old IDs, the 135 with a CyberGym ID are not significantly above chance (128 ± 7.5 expected within projects, z = 0.9, p = 0.19; 122 ± 9 under a uniform draw, z = 1.5, p = 0.08). Nor does it show that anything used the fixed build.
The pitfall: bugs with two IDs
The obvious check, matching bug numbers across the two ID lists, finds 139 of the 223 CyberGym tasks. That is only 62% of the overlap. (4 of the 139 match only if you ignore CyberGym's oss-fuzz: prefix.)
OSS-Fuzz moved to a new issue tracker and its bugs got new IDs, so the same bug can be arvo_54839 or arvo_42519792. CyberGym's ARVO tasks use old IDs, while MiMo mixes both (445 old, 555 new). You need ARVO's old→new mapping (arvo/oss_fuzz_mappings.csv) to match them.
The same translation shows that the 1,000 MiMo cyber tasks cover 836 distinct bugs: 164 appear twice, once under each ID.
A 13-gram filter catches none of it
A common decontamination step drops a training item if it shares a 13-word sequence with a benchmark item. I ran that filter, comparing MiMo's full task prompts with CyberGym's vulnerability descriptions. It flags none of the 278 overlapping MiMo tasks, and none of the other 722. For the 223 shared bugs, the two texts share at most 5 words in a row (median 2). A looser 8-word filter flags 10 tasks, all through the same unrelated CyberGym description (it shares a long C++ type name with them), never their own bug.
That is not surprising. MiMo's one-line crash spec (sanitizer, crash type, function, file) looks taken from a sanitizer report. CyberGym's description is an LLM rephrasing of the fix commit's message (CyberGym paper, §3.3). Even against CyberGym's own sanitizer reports, which agents see only at higher difficulty levels, only 6 of the 223 shared bugs share a 13-word sequence with their own bug's report, through long C++ signatures and file paths. (I compared each bug with its own report; I didn't run a filter against all 1,507 reports.) Same bug, mostly different words. I only tested exact word n-grams. Fuzzy or identifier-based matching might catch part of it: for 61 of the 219 shared bugs where MiMo's prompt names a function, CyberGym's description names the same one.
How this relates to CyberGym's own contamination check
CyberGym's authors do test for contamination: for four models, they split the benchmark by whether a bug was disclosed before or after the model's knowledge cutoff, and find no significant difference (§4). That test is designed around pretraining cutoffs. RL environments built from the same bugs are a separate route, and an ID-level check finds them directly.
The same pitfall in two more pairs
I ran the same kind of check on two more pairs. Both are in the repo.
Two benchmarks: SEC-bench vs CyberGym. SEC-bench added 100 OSS-Fuzz bugs to its dataset in November 2025. It names them by their new tracker IDs, while CyberGym's ARVO tasks use the old ones. A plain ID join finds 3 shared bugs. After translation it is 33 of the 100, all with the same fix. Two of SEC-bench's CVE instances also share a fix with a CyberGym task, and one fix covers two CyberGym tasks: 35 SEC-bench instances and 36 CyberGym tasks in all.
- Benchmarks sharing bugs is not a defect, and within each project the overlap is consistent with chance (z = 0.6, p = 0.31). But results on the shared bugs are not independent evidence. CyberGym also ships each task's fix patch, which for these bugs contains SEC-bench's gold patch.
- The text-filter result depends on which text you compare against. Against CyberGym's descriptions, a 13-gram filter flags none of the 35 shared instances. Against CyberGym's sanitizer reports, the same kind of text SEC-bench ships, it flags 29 of the 35 (after dropping sanitizer boilerplate), plus 40 of SEC-bench's 265 other instances.
An SWE training set: SWE-rebench-V2. SWE-rebench-V2 is an open set of 32,079 training tasks built from GitHub pull requests. 348 of them are the same PR at the same base commit as a task in Multi-SWE-bench, SWE-PolyBench or SWE-bench Multilingual: 5.1%, 11.3% and 4.0% of those benchmarks.
- Here a text filter works. V2 keeps each PR's original issue text, so a 13-gram filter flags 347 of the 348. For the companion V2-PRs set, where an LLM writes the problem statements, none of the 20 overlapping PRs shares a 13-word sequence with its benchmark task (the longest shared run is 6 words).
- Names cause a smaller version of the ID problem: SWE-bench Multilingual stores repo names in lowercase, so an exact name join misses all 12 of its overlaps.
- V2 doesn't claim to exclude these benchmarks. Multi-SWE-bench PRs appear in V2 at about the same rate as PRs that Multi-SWE-bench's own pipeline discarded (83 vs 77.4 ± 4.1 expected, z = 1.4; one-sided p = 0.93 for a shortfall), though that test has little power.
- V2's authors replied the same day. They had excluded the original SWE-bench's repos and did not explicitly filter against these benchmarks. They added a warning to the dataset card that links to the overlap lists, so users can filter before training.
Takeaways
- When training data and a benchmark come from the same upstream source, check overlap on source IDs (bug IDs, fix commits, PR plus base commit), not only text. A 13-gram filter against the benchmark's task descriptions caught none of the overlap in the MiMo and SEC-bench cases, and nearly all of it for SWE-rebench-V2, where both sides keep the original issue text.
- Normalize IDs first. Upstream systems get renumbered, and repo names change case.
- Compare against a chance baseline, and hold the project mix fixed: SEC-bench vs CyberGym looks above chance under a uniform draw (z = 2.1, p = 0.025) but not within projects (z = 0.6, p = 0.31). For MiMo vs CyberGym, the modest excess stays either way.
Everything is reproducible with one uv run per case, from pinned dataset revisions and image digests. The exception is Docker Hub's build steps, which can only be read from current tags; their digests are recorded. The repo has the overlap lists and filtered splits for all three cases, including a 594-task MiMo cyber split filtered against CyberGym only, not against SEC-bench, which shares 15 bugs with the MiMo cyber set: https://github.com/raimondasl/agentleak
I raised each case with the maintainers first: on the MiMo dataset's discussion page (https://huggingface.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss/discussions/6), in a SEC-bench issue (https://github.com/SEC-bench/SEC-bench/issues/4) and on SWE-rebench-V2's discussion page (https://huggingface.co/datasets/nebius/SWE-rebench-V2/discussions/5). SWE-rebench-V2's authors replied the same day and added the warning to their card; I'll add any other replies here.
Top comments (0)