DEV Community

Cover image for AI coding agents and Theo's Rust TypeScript compiler: the caveats

AI coding agents and Theo's Rust TypeScript compiler: the caveats

Theo Browne announced an experimental Rust TypeScript compiler port on October 7. The project, called ts-rust or tsc-rs, reports substantial speed improvements on some workloads and an implementation produced with AI coding agents. For developers waiting on type checks, that is a good reason to investigate. Its documented compatibility issues and benchmark conditions also give us a fairly concrete review checklist.

TL;DR

  • tsc-rs ports Microsoft's existing Go compiler implementation and inherits its design and tests. That provenance is central to understanding what the agents accomplished.
  • The project's 12.5× headline compares T3 Code with Effect diagnostics against TypeScript 6 and an older Effect plugin. Other comparisons produce different winners.
  • Theo's dollar figures value token usage at API prices. His reported Opus usage ran through Claude subscriptions.
  • Match compiler revisions, compare diagnostics and exercise builds before changing a release gate. The README and upstream tooling manifest currently point to different defaults.
  • My verdict is NEEDS REVIEW. A compiler needs an owner for the next upstream change, alongside evidence that today's code passes its tests.

What AI coding agents actually ported

Theo's announcement leads to a repository with an unusually explicit authorship boundary. Above “The Slop Line” is the author's account of the work. Below it is agent-written documentation. Tables below that line remain project claims, even when their formatting makes them look like the final answer.

The source implementation matters. The README describes a direct Rust port of Microsoft's native Go TypeScript compiler, preserving algorithms and behavior. An existing compiler supplies a detailed reference for the implementation. Its test corpus also supplies a large body of examples against which a port can be checked.

That makes the result interesting for a specific reason: the agents worked on a problem with an unusually strong behavioral reference. TypeScript semantics, diagnostics and the upstream implementation already existed. The difficult task was carrying that behavior into another language while retaining enough fidelity to work on real projects.

“Started from scratch,” in the author's account of Opus, refers to replacing the earlier Rust attempt. Microsoft's compiler work remains part of the story. That distinction helps readers evaluate how much of the achievement they could expect to reproduce on a new project with ambiguous requirements and a small test suite.

The repository reports 181,711 ported Go tests passing, identical diagnostics on TanStack Query core and Hono, and evaluations across 120 repositories. Those are useful claims to inspect. Our episode did no independent compiler run, so I cannot extend them into a universal compatibility guarantee.

Token spend, subscriptions and the cost of iteration

Theo's attached chart is headed “Token spend at API prices.” It assigns more than $400,000 in token-equivalent usage to earlier GPT attempts and roughly $24,047 to Opus. The distinction between a token valuation and a cash bill belongs next to every retelling of those figures.

Theo's original chart labels the figures as token spend at API prices

The author says the Opus work used Claude accounts. He reports an initial version in ten hours, followed by two weeks of work, and usage that heavily exceeded the normal weekly plan limits. These are author reports; we have no independently audited usage logs.

The practical implication is about the shape of the work. A first implementation can arrive quickly while validation and repeated correction continue for much longer. A compiler port benefits from a reference implementation that can keep answering the question “does this behave the same way?” Every correction still consumes resources, whether the billing system presents them as subscription capacity or per-token charges.

For a team evaluating AI coding agents, separate time to first runnable version from time to an acceptable release. Also separate the price paid from the resources consumed. Otherwise the spreadsheet ends up benchmarking the billing plan.

TypeScript compiler benchmarks: read the comparison modes

The project benchmark harness provides enough detail to understand the headline. It pins a T3 Code commit, checks five project paths, and defines separate modes with and without Effect diagnostics. We inspected that methodology rather than reproducing its timings.

With Effect enabled, the table gives tsc-rs 11.13 seconds and TypeScript 6.0.3 with its older Effect language-service plugin 138.63 seconds. That yields the quoted 12.5× result. The Go-based TypeScript 7 plus Effect comparison takes 21.07 seconds, making the Rust result about 1.89× faster in that comparison.

The older plugin uses a different rule set. The reported server diagnostics differ, so the largest multiplier needs that qualification. Diagnostic work is part of the measured workload, and changing that work changes the meaning of the stopwatch.

Without Effect, the same project's T3 Code table gives Bun 4.07 seconds, tsc-rs 7.25, TypeScript 7 16.10 and TypeScript 6 62.63. Bun wins that mode. A separate Effect pass changes Bun's total workload again.

T3 Code mode What the project result supports
Effect enabled, compared with TS6 A large advantage in this comparison, with different older plugin rules
Effect enabled, compared with Go TS7 A smaller reported advantage for tsc-rs
Effect removed Bun is the fastest entry in the project's table

The measurements use an M4 Pro Mac, medians of five runs and one warmup. Incremental checking is disabled. The Node checker gets a specified heap allocation, and the harness uses native checker binaries without the npm launcher.

Build configuration is another important detail. The repository says preview CI artifacts lack the PGO and BOLT optimizations used in its measured build. A downloaded preview can therefore differ from the binary behind the published chart. Keep the build and workload attached to the number when comparing your own results.

Compatibility starts with the upstream revision

The README names a September 29 upstream pin, 673a5f17d713, and recommends a matching TypeScript development release. The current upstream tooling manifest instead identifies June's dc37b5249ab6 as its current default and describes September's pin as a future bump target.

This is a documented mismatch worth checking. Development tooling metadata alone cannot establish which revision an installed npm binary contains. We did no shipped-binary inspection, and the episode explicitly leaves that question open.

For an adoption test, establish the actual revision first. Comparing two compilers built around different TypeScript behavior can produce diagnostic differences that have several possible causes. Version drift makes attribution harder before anyone has found a porting bug.

The platform matrix currently lists Linux x64 and macOS arm64. Windows and Linux arm64 are unavailable. The documented issues include monorepo rootDir diagnostics and extra emit, stale or missing dependency outputs in some build cases, and editor memory growth during repeated edits. A successful check of one small application exercises only a fraction of that surface.

My review sequence would be to keep the existing checker in place, run the port alongside it at an equivalent revision, compare diagnostics, then exercise project-reference builds and long editor sessions. Inspect the emitted artifacts as well as exit codes. Record any divergence with a small reproduction that a maintainer can actually investigate.

Rust and WebAssembly offer another deployment option

The WASM documentation and browser example source describe a browser-capable build. We inspected those sources; we did no independent hosted-browser demonstration.

The documented module is about 4.2 MB, or 1.5 MB with Brotli. It is single-threaded, with a Web Worker recommended for browser use. Its interface lacks watch mode and the native LSP capability, among other limitations. Native compiler features should therefore be checked individually before promising them in a browser deployment.

Moving a type check closer to a user could change a product's server workload. Any VM savings depend on what currently runs server-side, how often the check executes and how much work remains there. This episode contains no infrastructure cost study. The useful next step is a measured deployment experiment on the actual product.

Also in this episode: Haiku pricing and maintainer concentration

Anthropic's Haiku 5.5 launch offers low token prices with an important tier boundary. Input and output pricing rises fivefold above 100,000 prompt tokens. Artificial Analysis also reports heavy output-token usage at maximum effort and warns that its provisional task costs exclude that tier step. Compare successful task cost and latency using your prompt lengths.

The infrastructure maintainer census counts regular contributors across 23 projects and finds eleven with one or two. Its commit threshold can miss review, support and other maintenance work. Even with that limitation, it prompts a useful question for every new port: who owns the ongoing updates?

Verdict: NEEDS REVIEW

I stamped this NEEDS REVIEW because the project's speed claims deserve equal-version testing and a visible maintenance path. Microsoft's compiler will keep changing. A downstream port needs someone accountable for following those changes, investigating regressions and reviewing the next release.

That is the adoption gate I would use before replacing a working checker. Which would block your switch first: platform support, diagnostic fidelity or maintenance ownership?

FAQ

Did Theo spend $420,000 in cash on the compiler? The attached chart values tokens at API prices. His reported Opus work ran through Claude subscriptions. Treat those figures as usage equivalents.

Is tsc-rs always 12.5× faster? That multiplier belongs to the project's T3 Code plus Effect comparison against TS6 and an older plugin. Its other modes give different results.

Can I use the Rust compiler on Windows? The documented platform list currently excludes Windows. Check the repository for changes before planning an evaluation.

Does the manifest prove the shipped compiler uses June's upstream? The manifest describes tooling state. Determining a shipped binary's actual revision requires checking that binary; we have no such result here.

Sources


This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.

Video transcript and primary sources

Top comments (0)