Our agents' tests went from 64 passes to 1,334 while the renderer stayed broken. Five checks that passed and proved the wrong thing.
In Favur's 296-hour attempt at a Terminal-Bench WebGL renderer, local test passes rose from 64 to 1,334 while the last internal conformance results stood at 13 of 672 WebGL1 cases and 20 of 2,598 WebGL2 cases. The agents were not faking tests. Code review caught real bugs, the repairs still reproduce, and a fresh build of the source draws a triangle through its own shader compiler and rasterizer.
The gap came from the checks around the code. We wrote a full postmortem of the run, linked at the end. This article covers what the run built, what it can draw, what worked, and the five checks that passed while proving something else. Anyone who leaves agents running on a long job probably has the same five.
The job
The job is the WASM WebGL Renderer challenge from Terminal-Bench. Terminal-Bench Challenges are single tasks that ask an agent to build a whole codebase from scratch, with no time limit and no one stepping in. We did not write this one. It asks for a renderer that runs WebGL1 and WebGL2 programs, written from nothing. It could not use a GPU or the browser's own graphics code. That means the agents had to write the parts a graphics driver normally hides. Those are the API state machine, a compiler and interpreter for shader programs, and a rasterizer that turns triangles into pixels.
Favur split the work by role. One model mapped the architecture. GLM 5.3 coordinated, planned the sprints, managed implementation and reviewed each sprint against its plan. Muse Spark 1.3 Contributor wrote the code, reviewed it, and handled build, test and platform work. Gemini 3.8 Flash wrote pseudocode and watched agent conduct. Grok 4.6 supplied the independent critics.
What the run built
The record covers 296 wall-clock hours across nine sessions, about 119 of them working time once idle and stalled stretches are removed. The work reached its fourteenth sprint.
By Sprint 8 the source had all three main subsystems in place, the state layer, the shader compiler and interpreter, and the rasterizer. The local test count kept climbing after that, from 64 reported passes in Sprint 1 to 1,334 in Sprint 13.
Those 1,334 are local assertions the agents wrote. They are a different thing from the conformance suite, and the distance between the two is what the rest of this article is about.
What it can draw
For the postmortem we rebuilt the captured source fresh and drove it with small programs of our own. This is the result of the first one.
It is a plain red triangle, and every stage behind it is real. A vertex shader and a fragment shader were compiled and linked by the agents' compiler. The corners came from an attribute buffer. The draw call ran through the rasterizer, and the pixels were read back out. Nothing was drawn by a shortcut.
Three more probes ran the same way and returned the right pixels.
- A shader that multiplies a matrix by a vector and reads the first component gives the expected red with full alpha.
- A shader that divides by negative zero and handles very small numbers gives the expected color with no graphics error.
- A query for the list of compressed texture formats returns an empty typed array, which is what a caller needs when there are none.
One probe shows the edge of what works. Reading that same matrix product by index, where the first probe read it by field name, still corrupts the other color channels. The triangle and the probes show small WebGL1 paths working. They are far from showing a finished renderer.
What worked well
Review that checked meaning. In Sprint 8 a code reviewer found that the renderer told callers to step 16 bytes to the next matrix in an array, whatever the matrix size. For larger matrices the values would overlap. The reviewer called it a "data-corruption risk, not a style nit". An old test had accepted the wrong value, and the fix corrected the test along with the code. The repaired spacing is 32, 48 and 64 bytes for the three matrix sizes, and a fresh probe still returns those numbers.
Fixing the test when the test was wrong. In Sprint 12 a reviewer flagged that the compressed-format query returned an ordinary empty array where the API calls for a typed one. The code agent changed it and ran into a test that still demanded the old form. The agents corrected the test to match the API contract. A weaker loop would have bent the code back to keep the test green.
Tracing a wrong pixel to its cause. A code agent chased one bad color channel down to a collision in how values were stored. A four-part vector and a small matrix were both arrays of length four, so multiplication picked the wrong case and fell through to zero. The agent moved the size checks ahead of the ambiguous step.
Small numeric care. Another code agent found that division lost the sign of a zero divisor, and repaired it so that one divided by negative zero gives negative infinity as the shading language requires.
A platform repair with proof. In the last session a platform agent found two folder links that failed to open while the folders they pointed at were fine. It rebuilt them as relative links and reran the same check to show the fix held. A new read-only check still passes through them.
Review that caught missing work. Sprint review compared a sprint's claims with what existed and found that central fixes had never landed. It sent them back, and they were then done.
An honest last report. When the workflow ran out of legal moves, the orchestrator did not call the unfinished tasks finished. Its last record reads "Blocked by a workflow-capacity deadlock", with a brief for resuming.
Most of these compared the work with something outside it, such as the API contract, the shading language rules or the sprint plan.
Check 1. A test that skips the front door
The renderer had a WebGL2 class, and the WebGL2 test runner built that class directly. A real caller never does that. A caller asks a factory for a context and names the version it wants.
That factory returned nothing for an explicit WebGL2 request. The interception code in front of it dropped the requested type and handed back WebGL1. So the class could pass its tests all day while the only route a user would take was broken, and the entry tests that existed asserted the refusal as correct.
Check yours. For each feature, find one test that enters the way a caller enters. If every test constructs the inner object by hand, the wiring is untested, and agents are very good at making the inner object pass.
Check 2. A green wrapper around a red number
The WebGL2 conformance wrapper checked that the run did not crash, that results were recorded, and that two runs agreed. It did not require the conformance number to reach anything. A green wrapper sat on top of 20 of 2,598.
This is the easiest one to write by accident. "The suite ran cleanly" and "the suite passed" look the same in a list of green ticks.
Check yours. Open the test that guards your headline number and look for the line that fails when the number is low. If there is none, the test is checking that the measuring worked.
Check 3. An approval nobody gave
About 110 hours in, the orchestrator, the agent that coordinates the others, raised a concern that the conformance goal looked out of reach. The concern said it needed the human operator. An automated reply came back from another agent. The orchestrator's next record called it an "Operator decision".
That reply proposed interim targets of 80% and 50%. The decision record and the next sprint plan carried those targets forward as authorized. No person had answered it.
An agent cannot tell a human answer from an automated one unless the answer says where it came from, and here the provenance was in the record and dropped one step later.
Check yours. When a step asks for a person, record who answered and keep that with the decision. A timeout default is fine as a default. It must not be able to satisfy "ask the operator".
Check 4. A measurement of fixes that were not there
The same request for direction said a wave of fixes had flipped no additional tests, which made the goal look hopeless. Three hours later, sprint review found that central fixes from that wave had never landed. Commits and regression files the plan claimed did not exist.
The test counts were accurate. They described the program as it stood. They said nothing about the fixes, because the fixes were not in it. The plan had already changed on the strength of an experiment that had not been run.
Check yours. Before anyone reads a before and after, make the report show three things. The change exists, as a commit or a file. The baseline was taken before it. The second measurement was taken after it.
Check 5. A zero that was not a score
The final outside evaluation reported zero tests executed and zero passed. A build tool could not load its Linux dependency, so the test runner stopped before it reached the renderer. The workspace had been built on Windows and mounted into a Linux container with its Windows dependencies inside.
Zero of zero is a setup failure, and it tells you nothing about the code. Our own run made the opposite mistake at the same moment. The orchestrator's last record said it was blocked, and the run's outer status still read completed and successful.
Check yours. Give "did not run" its own outcome, separate from "ran and failed" and from "passed", and build dependencies on the platform that will do the judging.
The short list
- One test per feature goes in through the public entry.
- The test on your headline number fails when the number is low.
- A decision keeps the name of who answered.
- A before and after shows the change exists.
- "Did not run" is its own result.
The full postmortem follows the whole run in time order with the evidence for each claim, and it is here. The code the agents wrote is in the WasmRender-Challenge repository. Favur itself is closed-source and invite-only, and the Favur site has a real run you can drive.
We are running the challenge again in a few days. Follow Favur here on Dev.to or @favurdev on X to see how the second attempt goes.





Top comments (0)