DEV Community

Billie M
Billie M

Posted on Originally published at billiem.uk

What GPT-6.1 Sol chose to build

A cutting layout can look convincing while leaving its cut order ambiguous. A document comparison can work while putting its navigation offscreen. A traffic solver can return the right numbers without making them understandable.

Those were concrete review problems in three autonomous GPT-6.1 Sol experiments. The subjects were freely chosen: sheet cutting, the editing of an official record, and congestion. What changed during review was how someone could inspect and use the result.

I wanted to continue the Choice exercise: let a new model choose what to build, then examine its outputs and style. This batch ran at Extra High reasoning. The coordinating agent assigned compact, substantial and flagship ambition levels. One fresh builder chose all three topics and implemented the first two; a second fresh Sol builder took the flagship before public code existed.

A packing diagram needs a cutting sequence

Try Offcut with sheet dimensions, part quantities, grain constraints and blade width. It produces a layout and a sequence of full cuts through the remaining rectangles. Playback shows when a piece is actually released, rather than merely showing where it ends up.

The supplied job places nine parts on a 1,220 × 610 mm sheet. The blade's lost material changes feasibility: two 50 × 100 mm parts fit a 100 × 100 mm sheet with zero blade width, but only one fits with a 3 mm blade.

A 1,220 by 610 mm sheet contains three labelled shelves, two sides, smaller pieces and hatched remainder areas, with grain running horizontally.

Offcut's supplied nine-part layout. The diagram accounts for blade loss and grain; no physical cutting was tested.

The review problem was in the CSV cut order. Two physical copies of identically named stock were indistinguishable. Stable physical sheet IDs now distinguish them. Mobile feedback also changed because solving could leave the result offscreen.

The outputs include a dimensioned SVG, CSV instructions, editable JSON and remaining-stock inventory. Their usefulness still depends on the model's bounds: axis-aligned rectangular parts and full guillotine cuts. Damaged edges, trimming allowance, clamping and tool clearance are excluded, and no physical cutting was tested. An unplaced part is not proof that every possible arrangement fails; this is a bounded search.

A text comparison needs a way back to its source

Explore The Minutes by selecting sentences across six authored versions of a fictional exhibition incident. It exposes earlier wording and lets you recover facts omitted from the current record.

“The issue was managed promptly” traces back to a request to isolate the power, refused because the display lighting had to stay on. The final reassurance is not established by that first account. The sequence also contains a useful correction: an uncertain pump-stoppage time becomes a time supported by the controller log. Editing is not uniformly treated as concealment.

The earlier isolation request and refusal are struck through above the introduced reassurance, ‘The issue was managed promptly.’

The final record compared with the first account in The Minutes. This ancestry belongs to an explicitly authored fictional incident.

The initial navigation began below the visible page. Review moved it above the document; on mobile, selecting a sentence brings its provenance into view and supplies a return to the sentence. The explanation distinguishes removed facts from introduced assurances. The corpus and its ancestry are authored, so this is an inspectable argument, not a dishonesty detector for uploaded documents.

A solver needs more than its raw output

Change the network in The Shortcut Tax. At the default west demand of six flow units, average travel is 80 minutes without the shortcut and 100 with it under individual route choice. No driver can save time by switching routes alone.

The open shortcut carries four flow units, and average travel is 100 minutes beside an 80-minute outer-roads comparison. Individual route choice is selected.

The default six-unit demand with the shortcut open. These are results of the teaching network, not measured traffic times.

Coordinated routing restores the 80-minute average. Pricing each driver's imposed delay can reproduce it too, with tolls expressed in equivalent minutes. Changing demand can make the same shortcut helpful, harmful or unused.

The prototype initially exposed calculations as raw JSON. Review asked for a legible network, road and route costs, a demand sweep and a toll ledger. A second origin makes the shared bottleneck's effects inspectable by neighbourhood. These are teaching results under fixed demand, instantaneous equilibrium and simple congestion costs, not measured traffic or a forecast with queues.

The published surfaces still resemble September's Astra batch in their warm paper, serif headings and muted colours. The working conditions differ: Astra used three separately assigned builders at Max, with different subjects and a different brief. These portfolios cannot establish a model's taste or superiority.

The concrete comparison is between the prototype and what a visitor can now do. Follow a sentence back, distinguish physical sheets in a kept cut order, or inspect why the network's travel time changes. Review altered those routes into the underlying work, not just its finish.


Want to talk about something I’ve written or built? Get in touch.

This article was adapted with AI assistance from an original article on billiem.uk. The original article was reviewed before publication.

Top comments (0)