Nesyona Research // Press Kit

Press Kit: AI API Token Price Decay 2022-2026

Published 2026-05-09 · Embargo: none, available for immediate publication · Read the full study

Dataset summary

Five quotable findings

"The cost of a standard 1,000-input, 500-output AI task fell roughly 530 times between March 2023 and August 2024. GPT-4 32K cost twelve cents per task; Gemini 1.5 Flash, after its August 2024 price cut, costs about two-hundredths of a cent. The decay rate is fastest at the bottom of the price ladder, not the top."

Attribute to: Vincent, Nesyona Research, in Nesyona's 2026 AI API Token Price Decay study

"Claude Opus held its price for 20 months. Anthropic priced Opus 3 at fifteen dollars input and seventy-five dollars output per million tokens in March 2024 and shipped Claude 4 Opus at the identical numbers in May 2025, choosing a faster model at the same price rather than a cut. That ended in November 2025, when Claude Opus 4.5 launched at five dollars input and twenty-five output, a 67 percent cut."

Attribute to: Vincent, Nesyona Research

"Gemini 1.5 Flash holds the price floor, and it got there through a price cut rather than at launch. Google cut it from $0.35 to $0.075 per million input tokens in August 2024, and no later model in the dataset, including Gemini 2.0 Flash, has undercut it. Google owns its TPU stack and can amortize silicon across consumer Search workloads, which lets Flash list price approach marginal compute cost rather than carrying a sales-cycle margin."

Attribute to: Vincent, Nesyona Research

"Open-weight hosted pricing leads each closed-weight cost cut by roughly one quarter. DeepSeek V2 priced at fourteen cents per million input tokens in May 2024. Two months later GPT-4o mini matched at fifteen cents. Together AI and Fireworks publish open-weight pricing the week the weights become public, which establishes a marginal-cost reference rate that closed-weight providers then chase at the next product launch."

Attribute to: Vincent, Nesyona Research

"Output is now eight times more expensive than input on the GPT-5 tier. GPT-3.5 was one-to-one. The widening reflects compute economics: prefill is parallelizable across the prompt while generation is sequential. Builders running output-heavy workloads (writing, reasoning, code generation) face a meaningfully smaller cost cut than the headline 530x decay implies."

Attribute to: Vincent, Nesyona Research

Suggested coverage angles

This study lands at the intersection of AI infrastructure economics, vendor strategy, and developer tooling. Angles the data supports:

Downloads

Embed the headline chart

Free to embed under CC-BY 4.0 with link attribution to the study page:

<iframe src="https://nesyona.com/research/ai-token-price-decay-2026/embed/decay-curve.html"
  width="100%" height="400" frameborder="0" loading="lazy"
  title="AI API token price decay 2022 to 2026, Nesyona Research">
</iframe>

Media contact

Vincent, Nesyona Research

Email: vincent@deepsynthesis.org

Available for: written quotes (24 hour turnaround), data drill-down requests, podcast or video interview, on-record commentary on AI API pricing trends, vendor strategy reads, and developer-cost implications.

Response time: typically same-day during ET business hours.

About Nesyona

Nesyona is an independent AI tool economy research site at nesyona.com. We track real pricing, free-tier longevity, capability boundaries, and cost-per-outcome across the AI product stack. Nesyona is part of the DeepSynthesis Lattice, a constellation of niche research sites covering ecommerce, finance, education, health, and AI verticals.