Skip to content
Brian Sithu
Go back

Three hundred tokens a second

Cover image for Three hundred tokens a second

OpenAI now sells speed itself with Ultrafast. I guess at where 300 tokens a second comes from, then argue latency is the next price dial.

Brian Sithu

Written by

Brian Sithu

View profile

On this page

02 sections

At DevDay, OpenAI launched a plan called Pro 500. Mashable reports it costs $500 a month. The headline feature is Ultrafast. OpenAI calls it “our premium speed tier for workloads where speed matters most.”

The numbers, all OpenAI’s: up to 8x faster token generation at 300 tokens per second in Codex, and up to 6x in the API. The plan itself carries 25 times the ChatGPT Plus allowance. Astra Ultrafast is available today in the API and in ChatGPT Work and Codex on Pro 500 and Enterprise plans. A Sol version is coming soon.

Read that again. Not a bigger model. Not a cheaper token. A faster token. Speed is now a product with a price.

Speed has to come from somewhere

Here is why a speed tier is strange. Generating tokens is memory-bound. Every output token drags the whole weight matrix through GPU memory. Same weights on the same chips run at the same speed. Nobody repeals that for $500 a month.

So 300 tokens a second must come from somewhere else. OpenAI has not said where. Everything below is my guess, and I am labeling it as such because the mechanism is the interesting part and I do not actually know it.

My first guess was wrong, and ruling it out narrows things down. I assumed the agents just got smarter: fewer round trips, more done per call. That improves task time but it cannot move tokens per second. OpenAI quotes tokens per second. So the mechanism sits in serving, not in the harness.

What serving tricks buy real multiples? Speculative decoding is the obvious one. A small cheap model drafts a few tokens, the big model checks them all in one pass, and the accepted ones cost a fraction of the memory traffic. It is a known technique and my guess is that it does a lot of the work here.

Then the boring answer, which is probably the biggest one. Dedicated capacity. Pinned GPUs, warm caches, no shared queue, no noisy neighbor spiking your tail latency at 3pm. You are not buying faster math. You are buying an empty highway.

Prefetch rounds out my guesses. Coding agents reread the same repos and replay the same long contexts. Keep that state hot and the first token stops paying the cold-start tax.

My guess at the Ultrafast serving path: a small model drafts, Astra verifies in one pass, pinned GPUs keep the queue empty

None of this is confirmed. It is the shape I would expect any lab to use, because these are the only levers that exist when the weights are fixed.

Latency is the next price dial

Now the opinion. I like this pricing and I want more of it.

Per-token pricing always hid the queue. Two users paid the same rate while one waited behind the other. Ultrafast makes the queue visible and lets you buy your way out of it. That is honest. Cloud vendors learned this a decade ago with provisioned throughput. The labs just got there.

It also prices the thing I actually feel. My agent loops are latency-bound. A coding run that streams at 300 tokens a second changes what I attempt. I give it bigger jobs. I stop batching my questions like it is 2022. Tokens were never my budget. Minutes were.

The uncomfortable side is plain too. A paid fast lane means the free lane is now explicitly the slow lane. Shared-tier users wait so Pro 500 users do not. That tradeoff was always there, buried in capacity planning. Now it has a price tag, and everyone can see it.

My prediction: every lab ships a speed tier within a year, and benchmarks shift from cost per task to cost per minute. The model is no longer just the weights. It is the weights plus the highway.



Previous Post
The smallest model in your agent stack
Next Post
Rate limits weren't built for an agent army