Originally published on AI Tech Connect.
What changed this week A new benchmark and a new pipeline. "ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks" (arXiv 2609.18805, September 2026) comes from KAIST, Microsoft Research Montreal and Microsoft AI. A dataset is published at huggingface.co/datasets/microsoft/ProgramDistill. Scale without labellers. The pipeline discovered 1,975 replay-verified behaviours across 26 applications and constructed 4,063 tasks without human intervention. Three difficulty bands. Atomic repair, cumulative repair, and full-application reconstruction — a built-in ladder rather than one undifferentiated pile of problems. Nine frontier coding agents were evaluated. On cumulative workflows within the full-application-reconstruction band, GPT-6 Astra reached 49.2% success and…
Top comments (0)