DEV Community

gentic news
gentic news

Posted on • Originally published at gentic.news

InAgent Hits 90.2% on OSWorld, First Agent Past 90%

InAgent scored 90.2% on OSWorld, first above 90%, with 100% on system-level tasks, surpassing OpenAI, Google, and Anthropic records. Harness engineering, not raw model power, drove the result.

InAgent scored 90.2% on OSWorld in July 2026, the first computer-use agent above 90%. The Chinese agent's result surpasses public records from OpenAI, Google, and Anthropic on the same benchmark.

Key facts

  • 90.2% OSWorld task success, first above 90%
  • 100% success on system-level tasks
  • Surpasses OpenAI, Google, Anthropic records
  • Single run; seed variance not disclosed
  • Harness engineering drives the result

InAgent scored 90.2% task success on OSWorld in July 2026, the first computer-use agent above 90%, according to Pandaily. The agent achieved 100% on system-level tasks, a category that has historically dragged down frontier models. OSWorld measures an agent's ability to complete real-world computer tasks — file operations, web browsing, application control — through screenshots and keyboard/mouse actions.

Key Takeaways

  • InAgent scored 90.2% on OSWorld, first above 90%, with 100% on system-level tasks, surpassing OpenAI, Google, and Anthropic records.
  • Harness engineering, not raw model power, drove the result.

Why the harness matters

The headline number obscures the structural story: the gap between frontier models has narrowed to the point where the scaffold around the model now determines benchmark placement. Harness engineering — the code that plans, verifies, and recovers from errors — has become the new AI competition frontier. InAgent's 90.2% is not primarily a model win; it is a systems win.

This mirrors what Supabase's evals benchmark showed in August 2026: real-world agent performance tracks tooling and orchestration more than raw model capability. Claude Code's strong showing there came from its scaffolding, not just Claude's weights. InAgent confirms the pattern on a harder, more standardized benchmark.

The numbers behind the record

What Screen Agent’s #1 OSWorld ranking means for UI automation i…

The 90.2% figure represents a single run. The source does not disclose variance across seeds, a meaningful omission for a benchmark where stochastic sampling can swing results by several points. OSWorld's system-level tasks — which require multi-step operations like installing software or configuring settings — are where most agents fail; InAgent's 100% there is the more impressive number.

Prior public records on OSWorld sat below 90%. OpenAI, Google, and Anthropic have each published computer-use agents, but none crossed the threshold. InAgent's result places the Chinese lab ahead on this specific metric, though the benchmark is narrow: OSWorld covers desktop tasks, not the full range of enterprise workflows where those companies compete.

What to watch

Watch whether InAgent publishes seed-variance data or a follow-up paper detailing its harness architecture. Also track whether OpenAI, Google, or Anthropic respond with OSWorld scores above 90% in the next two quarters — a response would confirm harness engineering as the new arms race.


Source: pandaily.com


Originally published on gentic.news

Top comments (0)