What the gateway actually measured
Vercel’s AI Gateway Production Index for September 2026 reports anonymized traffic through August. Open-weight models ran 56% of gateway tokens that month, up from 7% in December 2025. Those same models accounted for 14% of estimated spend. Anthropic still took 64% of spend.
The price of a token is falling with that mix. Vercel reports that the average price per token dropped 23.2% in August, the third monthly decline in a row, and that the average token costs less than half what it did five months earlier. Among teams that ran more than ten million tokens in both July and August, the median team paid 7.6% less per token.
Those figures are external market data, not a StratEdge measurement. Spending in the index is estimated from labs’ published list prices, so an invoice can differ. Vercel also says its open-weight classification is broader than in earlier reports, and that earlier months can be revised.
What the split does not prove
A majority of tokens is not a majority of value. Closed models still collect most of the money because some work is priced for judgment. Open-weight models are not universally better than closed models. The data says they are being used for a large share of volume, at a much smaller share of spend. It does not say they win every task.
Vercel’s own line on the same report is the one worth keeping: teams can get more inference from the same budget and reserve frontier models for the tasks that justify the premium.
The future may not be one giant model doing everything. It may be specialized intelligence doing exactly what each workflow requires.
How StratEdge routes the work
That routing idea is StratEdge’s operating approach. Read it as interpretation, not as another statistic from the index.
- Use an open-weight model where the job is narrow.
- Train or optimize it for that specific operational task.
- Reserve a frontier model for work that genuinely needs frontier reasoning.
- Choose on cost, control, latency, and task-specific performance.
StratEdge has not published a measured savings percentage for this approach. The point is the rule, not a claim about our bill. A document to classify, a field to extract, a limit to check, or a routine step to draft can sit inside workflow control without sending every step to the most expensive model on the market. When the decision is ambiguous, or the cost of being wrong is high, the frontier model is the right spend.
A test before the next model call
Name the job before you name the model. If the job is narrow and the standard is clear, test a specialized open-weight model against that standard. If the job needs frontier reasoning, pay for a frontier model and stop treating the price as a failure.
The cost curve is breaking because buyers can separate throughput from judgment. It is not breaking because one kind of model replaced the other. More of this kind of note lives on StratEdge Research.