OpenAI has released GPT-6 Astra Ultrafast, a faster inference mode of its Astra model running on NVIDIA Blackwell GPUs. It's available now through the OpenAI API and to eligible ChatGPT Work and Codex users.
The headline figure: Ultrafast delivers up to 8x faster token generation than Astra's Standard mode. OpenAI frames the gain around agentic workflows — coding agents running edit-test-debug cycles, and any application where a model writes code, calls a tool, checks the result, and decides what to do next. Shortening each generation step in that loop shortens the whole cycle, since the delay repeats every time the agent takes an action.
Philippe Tillet, OpenAI's inference lead, said NVIDIA's tooling and documentation let OpenAI's models become "exceptionally good at programming Blackwell and Rubin GPUs," turning that into high-performance kernels that improve latency, throughput and cost on NVIDIA hardware. Uday Ruddarraju, OpenAI's chief technology officer of compute, said the company used its own models to optimize the inference software running on NVIDIA GPUs, with NVIDIA's programmable platform enabling the acceleration behind Ultrafast.
OpenAI also notes that performance work continues after deployment: it's using its own models to keep refining inference software on the same GPU platform, and that a programmable NVIDIA stack lets teams reuse infrastructure across training, inference and reinforcement learning rather than provisioning separately for each.
Developers can access GPT-6 Astra Ultrafast through the API now; OpenAI points to its Ultrafast guide for access, pricing and implementation details.
Top comments (0)