
OpenAI has launched Astra Ultrafast, a new model mode running on NVIDIA Blackwell GPUs designed for high-speed inference. It provides up to eight times faster token generation compared to the standard mode. The release is currently available through the OpenAI API and to specific ChatGPT Work and Codex users.
Architecture and inference acceleration
The Ultrafast mode utilizes NVIDIA Blackwell architecture to increase the speed of token generation. OpenAI researchers used internal models to optimize inference software specifically for these GPUs, creating high-performance kernels that reduce latency. This hardware-software integration allows for faster model responses during complex sequences of tasks.
The system is designed to accelerate cycles where an agent must write code, execute a tool, and evaluate the result repeatedly. By reducing the time between these steps, the model aims to make autonomous agent loops and interactive applications more responsive for developers working with time-sensitive tasks.
Infrastructure and availability
OpenAI is currently serving the model to users of the OpenAI API, alongside ChatGPT Work and Codex subscribers. The development team continues to use their own models to refine how inference software interacts with GPU programmability, suggesting that further productivity gains may be implemented over time without hardware changes.
The source does not provide specific benchmarks regarding the model's reasoning accuracy or cost per token compared to standard versions. While the 8x speed increase refers to token generation, the total end-to-end latency for complex multi-tool workflows remains dependent on the specific tools and external APIs an agent might call.
Original source
This report summarises the source below. Analysis is labelled separately; product and research claims remain attributed to their source.
Read the original at NVIDIA