al-Qantara
Observatory

AI adoption, in numbers.

Each dimension is a short data essay from a public source — with the chart, the story behind it, and the open math for how the number was reached. Pick a dimension to explore.

Pick a dimension…

Capabilities

From text to work: what AI can already do

0%cheaper

To match GPT-4 on GPQA Diamond, the output price per token fell from US$60 to US$2.19 per million — same test, nearly 30× cheaper, in under two years.

In little over three years, models jumped from 'completing sentences' to solving graduate-level problems, closing code issues on their own and sustaining tens-of-billions companies. This dimension measures that jump through three lenses — what they know, what they can execute, and how much money it already moves — and ends on the exam built to be unbeatable.

Cognitive capability — by year and by price

Each dot is a frontier model: when it shipped (horizontal), how much it scored on GPQA Diamond — 'Google-proof' graduate questions (vertical) — and which price bracket it plays in (color). The ceiling rises while the lighter dots show the same intelligence getting cheap: DeepSeek-R1 delivered 71% at US$2 per million tokens.

  • 2023-03GPT-435.7%US$60/M
  • 2024-03Claude 3 Opus50.4%US$75/M
  • 2024-05GPT-4o53.6%US$10/M
  • 2024-06Claude 3.5 Sonnet59.4%US$15/M
  • 2024-12o178%US$60/M
  • 2025-02Claude 3.7 Sonnet78.2%US$15/M
  • 2025-03Gemini 2.5 Pro84%US$10/M
  • 2025-08GPT-588.4%US$10/M
  • 2025-11Gemini 3 Pro91.9%US$12/M
  • 2026-02Gemini 3.1 Pro94.1%US$12/M

What this chart concludes

The vertical axis is GPQA Diamond accuracy (100% = perfect). In 2023 GPT-4 scored 36% at US$60/M tokens; DeepSeek-R1 soon hit 71% — double — at US$2.19. Hence the headline: beating GPT-4 got 96% cheaper. Capability up, price down.

Method
GPQA Diamond (198 questões de pós-graduação, 'à prova de Google'), score oficial pass@1 na data de lançamento. Preço = tabela de saída do provedor (US$/M tokens); cor = faixa de preço, linha = recorde acumulado.
Source
Epoch AI — GPQA Diamond
Snapshot from
July 05, 2026

Agentic capability — solving, not just answering

Answering well is one thing; executing a task end-to-end is another. SWE-bench Verified measures exactly that: 500 real GitHub issues the model must resolve by editing the repo until tests pass. In under two years, the frontier went from a third to almost everything.

  1. 2024-08GPT-4o
    33%
  2. 2024-103.5 Sonnet
    49%
  3. 2025-023.7 Sonnet
    62.3%
  4. 2025-05Opus 4
    72.5%
  5. 2025-08GPT-5
    74.9%
  6. 2025-11Gemini 3
    76.2%
  7. 2025-11Opus 4.5
    80.9%
  8. 2026-04Opus 4.7
    87.6%
  9. 2026-06Fable 5 👑
    95%

Claude Fable 5 across three agentic suites today

  • SWE-bench Verifiedcódigo de ponta a ponta
    95
  • Terminal-Bench 2.1tarefas no terminal
    88
  • τ-benchatendimento com ferramentas
    89.2
Method
SWE-bench Verified: 500 issues reais do GitHub resolvidas de ponta a ponta (editar o repo até os testes passarem) — o padrão de capacidade agêntica de código. Escada = recorde por data; barras = líder atual em três suítes. Nota: em 2026 um estudo de Berkeley mostrou que essas suítes podem ser 'gameadas' — leia como tendência.
Source
SWE-bench (leaderboard oficial)
Snapshot from
July 05, 2026

The economics behind it — who earns, and how much

Capability turned into revenue at an unprecedented pace. OpenAI opened the decade ahead, but in 2026 Anthropic crossed the curve — pulled by enterprise coding usage. Figures are annualized run-rate (the month times twelve), reported by the companies and the press.

US$0 biUS$7,5 biUS$15 biUS$22,5 biUS$30 bi202420252026
OpenAIAnthropic

Valuation, latest round

US$852 biOpenAI · 2026
US$965 biAnthropic · mai/2026

The end of the ruler

Humanity's Last Exam — the test built to be unbeatable

Once models started acing tests like MMLU, benchmarks lost their point: you could no longer tell who was ahead. The answer, launched in January 2025 by the Center for AI Safety with Scale AI, was to assemble the hardest exam humanity could write.

How it's made

  1. 1About a thousand experts from 500+ institutions submitted over 70,000 questions in their fields — from topology to virology.
  2. 2A question only makes it in if the best models of the day get it wrong: easy ones are dropped on the spot, an adversarial filter against the frontier itself.
  3. 3Survivors go through two rounds of peer review until 3,000 remain — 2,500 public and 500 secret, to catch models that 'memorized' the test.
  4. 4They're closed-ended (76% exact short answers, 24% multiple choice; ~14% need an image), so they can be graded automatically and compared fairly.

Best score over time (%)

  • 2025-01o1
    8%
  • 2025-04o3
    17%
  • 2025-08GPT-5
    25.3%
  • 2025-11Gemini 3 Pro
    37.5%
  • 2026-02Gemini 3.1 Pro
    45%
  • 2026-06Fable 5
    53.3%

The first models got 3% to 8% right. Eighteen months later, the frontier passed the halfway mark — on an exam designed not to be beaten any time soon. That's the pace this dimension tries to capture.

Method
Center for AI Safety (CAIS) + Scale AI. 3,000 questões (2,500 públicas + 500 secretas) em 100+ disciplinas, filtradas contra a fronteira. Lançado em 24/01/2025.
Source
Humanity's Last Exam (CAIS + Scale AI)
Snapshot from
July 05, 2026

Explore the dimensions