3 min read

Astra and Sol: Where Token Costs Actually Rise

AstraSolцены API

Published pricing shows Astra at $10 per million input tokens and $50 output, while Sol costs $5 and $30 in standard mode. Fast pricing generally doubles the applicable rates. The claimed 6.25x spend multiplier is not supported by official prices, though long context and output volume can raise bills.

What the prices and limits show

I would not put the claimed 6.25x spending increase into a budget: published prices do not support it. As of September 6, 2026, OpenAI documentation and the official pricing page list Astra’s standard rates as $10 per million input tokens, $1 per million cached input tokens, $12.50 per million cache writes, and $50 per million output tokens.

Astra’s Fast mode doubles the applicable rates. For requests exceeding 272,000 input tokens, the entire request is repriced: input rises to $20 per million and output to $75. This is not a marginal surcharge applied only to the excess tokens; it is a new price for the full request, so crossing the threshold accidentally can skew the bill substantially.

The compiled pricing materials for Sol list $5 per million input tokens and $30 per million output tokens in standard mode. Fast or Priority costs $10 and $60 respectively, which is a 2x multiplier rather than 2.5x. These rates also do not produce a 6.25x spending increase.

Limits are a separate matter. Astra documentation shows growth from 500 requests per minute and 500,000 tokens per minute at Tier 1 to 15,000 requests and 40 million tokens per minute at Tier 5. At high request frequency, tokens per minute usually become the bottleneck before request count does.

Users also report an available-limit reduction of roughly 20% over nine hours and 13% over one hour when working in two threads. Those observations cannot be converted directly into a pricing multiplier without data on context length, output, caching, and the selected mode. For Fable, the collected sources contain no official pricing and limits table, so an honest numerical comparison is not yet possible.

How this changes budget planning

The main risk is not the base price but the combination of expensive output, Fast mode, and long context. I would calculate input and output tokens separately, then double the result for Fast mode and separately model requests that may cross Astra’s 272,000-token threshold.

Repeated prompts materially change Astra’s economics: cached input costs $1 per million versus $10 for standard input. But caching does not solve expensive output, and high throughput does not eliminate TPM constraints. That is why an average cost per request is nearly useless without a length distribution.

In short, official rates indicate a twofold Fast premium, not a 6.25x increase. Until the mechanism behind user limits and Fable’s data are disclosed, the gap between the pricing table and observed charges remains the main unknown.

We previously examined how context costs and Claude Opus 4.6 settings affect total model spending. That analysis helps compare Astra’s token pricing and surcharges with other models.