Back to Langfuse

Model Price Audit Memory

.agents/skills/add-model-price/references/model-audit-memory.md

3.224.114.7 KB
Original Source

Model Price Audit Memory

This file is an optional snapshot of the latest automated audit whose per-model results add useful context for a future run. It is orientation only; reconfirm every price and tier against official provider sources before making a change or reporting a row as confirmed.

The audit agent may replace the snapshot below with its complete current table. Keep only one snapshot, never append an unbounded run history, never persist a partial set of checked models, and do not update this file only to refresh the date.

Latest useful snapshot

Audit date: 2026-07-23

All prices listed as $X / MTok (per million tokens). Per-token JSON values: divide by 1,000,000.

ProviderModel / pricing entryPricing checkedPrice confirmedTiering checkedTiering correctChangeOfficial source(s)Comments
Anthropicclaude-fable-5Input $10/MTok (10e-6), Output $50/MTok (50e-6), 5m cache write $12.5/MTok (12.5e-6), 1h cache write $20/MTok (20e-6), cache read $1/MTok (1e-6)YesFlat 1M context at standard pricing — no large-context tierYesNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed. In flat long-context list.
Anthropicclaude-mythos-5Input $10/MTok (10e-6), Output $50/MTok (50e-6), 5m $12.5/MTok, 1h $20/MTok, read $1/MTokYesFlat 1M context — no large-context tierYesNonehttps://platform.claude.com/docs/en/about-claude/pricingLimited availability (Project Glasswing). Same prices as Fable 5.
Anthropicclaude-opus-4-8Input $5/MTok (5e-6), Output $25/MTok (25e-6), 5m $6.25/MTok (6.25e-6), 1h $10/MTok (10e-6), read $0.50/MTok (0.5e-6)YesFlat 1M context — no large-context tierYesNonehttps://platform.claude.com/docs/en/about-claude/pricingFast mode at $10/$50 per input/output MTok (additional SKU).
Anthropicclaude-opus-4-7Input $5/MTok (5e-6), Output $25/MTok (25e-6), 5m $6.25/MTok, 1h $10/MTok, read $0.50/MTokYesFlat 1M context — no large-context tierYesNonehttps://platform.claude.com/docs/en/about-claude/pricingFast mode deprecated July 24 2026 per official page.
Anthropicclaude-opus-4-6Input $5/MTok (5e-6), Output $25/MTok (25e-6), 5m $6.25/MTok, 1h $10/MTok, read $0.50/MTokYesFlat 1M context — no large-context tierYesNonehttps://platform.claude.com/docs/en/about-claude/pricinginference_geo: "us" adds 1.1× multiplier for this and later models.
Anthropicclaude-opus-4-5-20251101Input $5/MTok (5e-6), Output $25/MTok (25e-6), 5m $6.25/MTok, 1h $10/MTok, read $0.50/MTokYesNOT on flat long-context list; no explicit large-context tier pricing publishedNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricing"Claude Opus 4.5" is now listed on the official pricing page (confirmed July 2026). matchPattern covers both claude-opus-4-5 and claude-opus-4-5-20251101. Not on flat long-context list (flat list: Fable 5, Mythos 5, Mythos Preview, Opus 4.8/4.7/4.6, Sonnet 5, Sonnet 4.6).
Anthropicclaude-opus-4-1-20250805Input $15/MTok (15e-6), Output $75/MTok (75e-6), 5m $18.75/MTok, 1h $30/MTok, read $1.50/MTokYesDeprecated model — no tieringNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingDeprecated. Retained for backward pricing compatibility.
Anthropicclaude-opus-4-20250514Input $15/MTok (15e-6), Output $75/MTok (75e-6), 5m $18.75/MTok, 1h $30/MTok, read $1.50/MTokYesRetired modelNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingRetired except on Google Cloud.
Anthropicclaude-sonnet-5Input $2/MTok (2e-6), Output $10/MTok (10e-6), 5m $2.50/MTok, 1h $4/MTok, read $0.20/MTokYesFlat 1M context; introductory pricing through August 31 2026YesNonehttps://platform.claude.com/docs/en/about-claude/pricingIntroductory pricing confirmed through Aug 31 2026. Standard from Sep 1 2026: $3/$15, cache $3.75/$6/$0.30. File must be updated after Aug 31 2026.
Anthropicclaude-sonnet-4-6Input $3/MTok (3e-6), Output $15/MTok (15e-6), 5m $3.75/MTok, 1h $6/MTok, read $0.30/MTokYesFlat 1M context — no large-context tierYesNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed.
Anthropicclaude-sonnet-4-5-20250929Input $3/MTok (3e-6), Output $15/MTok (15e-6), 5m $3.75/MTok, 1h $6/MTok, read $0.30/MTokYesLarge Context >200K: 2× input / 1.5× output — tier in file, not explicitly published by AnthropicNoNonehttps://platform.claude.com/docs/en/about-claude/pricingSonnet 4.5 NOT on flat list. Anthropic page does not publish per-tier pricing for this model. The Large Context tier was set when model was first added. Unresolved.
Anthropicclaude-sonnet-4-20250514Input $3/MTok (3e-6), Output $15/MTok (15e-6), 5m $3.75/MTok, 1h $6/MTok, read $0.30/MTokYesRetired modelNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingRetired except on Bedrock and Google Cloud.
Anthropicclaude-haiku-4-5-20251001Input $1/MTok (1e-6), Output $5/MTok (5e-6), 5m $1.25/MTok (1.25e-6), 1h $2/MTok (2e-6), read $0.10/MTok (1e-7)YesNo large-context tierNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed. Haiku 4.5 NOT on flat long-context list.
Anthropicclaude-3-5-haiku-20241022Input $0.80/MTok (8e-7), Output $4/MTok (4e-6), 5m $1/MTok, 1h $1.60/MTok, read $0.08/MTokYesRetired modelNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingRetired except on Bedrock and Google Cloud.
Anthropicclaude-3.7-sonnet-20250219Input $3/MTok, Output $15/MTok, cache $3.75/$6/$0.30 per MTokNoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingNot listed on current official pricing page. Legacy prices retained.
Anthropicclaude-3.5-sonnet-20241022Input $3/MTok, Output $15/MTok, cache $3.75/$6/$0.30 per MTokNoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingNot listed on current official pricing page. Legacy prices retained.
Anthropicclaude-3-5-sonnet-20240620Input $3/MTok, Output $15/MTok, cache $3.75/$6/$0.30 per MTokNoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingNot listed on current official pricing page. Legacy prices retained.
Anthropicclaude-3-opus-20240229Input $15/MTok, Output $75/MTokNoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingLegacy. Not on current pricing page.
Anthropicclaude-3-sonnet-20240229Input $3/MTok, Output $15/MTok, cache setNoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingLegacy. Not on current pricing page.
Anthropicclaude-3-haiku-20240307Input $0.25/MTok, Output $1.25/MTokNoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingLegacy. Not on current pricing page.
OpenAIgpt-5.6-solInput $5/MTok (5e-6), Cached $0.50/MTok (0.5e-6), Output $30/MTok (30e-6)YesLarge Context (>272K): input $10/MTok, cached $1/MTok, output $45/MTokYesNonehttps://developers.openai.com/api/docs/pricingReasoning model. 272K threshold unique to gpt-5.6 family.
OpenAIgpt-5.6-terraInput $2.50/MTok (2.5e-6), Cached $0.25/MTok (0.25e-6), Output $15/MTok (15e-6)YesLarge Context (>272K): input $5/MTok, cached $0.50/MTok, output $22.50/MTokYesNonehttps://developers.openai.com/api/docs/pricingReasoning model.
OpenAIgpt-5.6-lunaInput $1/MTok (1e-6), Cached $0.10/MTok (0.1e-6), Output $6/MTok (6e-6)YesLarge Context (>272K): input $2/MTok, cached $0.20/MTok, output $9/MTokYesNonehttps://developers.openai.com/api/docs/pricingReasoning model.
OpenAIgpt-5.5-2026-04-23 (also matches gpt-5.5)Input $5/MTok (5e-6), Cached $0.50/MTok (0.5e-6), Output $30/MTok (30e-6)YesNo large-context tier in file; large context mentioned by OpenAI but threshold not confirmedNoNonehttps://developers.openai.com/api/docs/pricingReasoning model. matchPattern covers plain gpt-5.5 (dated suffix is optional). Large-context threshold unresolved.
OpenAIgpt-5.5-pro-2026-04-23 (also matches gpt-5.5-pro)Input $30/MTok (30e-6), Output $180/MTok (180e-6); no cacheYesNo tiering in fileNot applicableNonehttps://developers.openai.com/api/docs/pricingReasoning model; no cached input pricing. matchPattern covers plain gpt-5.5-pro.
OpenAIgpt-5.4Input $2.50/MTok (2.5e-6), Cached $0.25/MTok (0.25e-6), Output $15/MTok (15e-6)YesNo large-context tier confirmedNoNonehttps://developers.openai.com/api/docs/pricingReasoning model. Large-context threshold unresolved.
OpenAIgpt-5.4-2026-03-05Input $2.50/MTok, Cached $0.25/MTok, Output $15/MTokYesNo tieringNot applicableNonehttps://developers.openai.com/api/docs/pricingDated snapshot for gpt-5.4.
OpenAIgpt-5.4-proInput $30/MTok (30e-6), Output $180/MTok (180e-6); no cacheYesNo tieringNot applicableNonehttps://developers.openai.com/api/docs/pricingReasoning model; no cached input pricing.
OpenAIgpt-5.4-miniInput $0.75/MTok (0.75e-6), Cached $0.075/MTok (0.075e-6), Output $4.50/MTok (4.5e-6)YesNo tieringNot applicableNonehttps://developers.openai.com/api/docs/pricingNon-reasoning model. Confirmed.
OpenAIgpt-5.4-nanoInput $0.20/MTok (0.2e-6), Cached $0.02/MTok (0.02e-6), Output $1.25/MTok (1.25e-6)YesNo tieringNot applicableNonehttps://developers.openai.com/api/docs/pricingNon-reasoning model. Confirmed.
OpenAIgpt-5-chat-latestInput $1.25/MTok (1.25e-6), Cached $0.125/MTok (1.25e-7), Output $10/MTok (10e-6)YesNo provider tieringNot applicableNonehttps://developers.openai.com/api/docs/models/gpt-5-chat-latestConfirmed via specific model page in July 2026 audit.
Googlegemini-2.5-flashInput $0.30/MTok (3e-7), Audio $1/MTok (1e-6), Output $2.50/MTok (2.5e-6), Cache read $0.03/MTok (3e-8)YesNo large-context tier; single tier confirmedYesNonehttps://ai.google.dev/pricingCache read = 10% of standard input. Confirmed.
Googlegemini-2.5-flash-liteInput $0.10/MTok (1e-7), Audio $0.30/MTok (3e-7), Output $0.40/MTok (4e-7), Cache read $0.01/MTok (1e-8)YesNo large-context tierYesNonehttps://ai.google.dev/pricingCache read = 10% of input. Confirmed.
Googlegemini-2.5-proInput $1.25/MTok (1.25e-6) ≤200K, $2.50/MTok >200K; Output $10/MTok ≤200K, $15/MTok >200K; Cache read $0.125/MTok ≤200KYesLarge Context (>200K) confirmedYesNonehttps://ai.google.dev/pricingTwo tiers confirmed. Cache = 10% of input at each tier.
Googlegemini-3.5-flashInput $1.50/MTok (1.5e-6), Output $9.00/MTok (9e-6), Cache read $0.15/MTok (1.5e-7)YesNo large-context tierYesNonehttps://ai.google.dev/pricingCache = 10% of input. Confirmed.
Googlegemini-3.1-flash-liteInput $0.25/MTok (2.5e-7), Audio $0.50/MTok (5e-7), Output $1.50/MTok (1.5e-6), Cache read $0.025/MTok (2.5e-8)YesNo large-context tierYesNonehttps://ai.google.dev/pricingCache = 10% of text input. Confirmed on AI Studio page.
Googlegemini-3.1-flash-lite-previewInput $0.25/MTok, Output $1.50/MTok (same as GA)NoNo large-context tierNot applicableNonehttps://ai.google.dev/pricingPreview variant; same prices as GA version in file; not separately listed on official page.
Googlegemini-3.1-pro-previewInput $2/MTok (2e-6) ≤200K, $4/MTok >200K; Output $12/MTok ≤200K, $18/MTok >200KYesLarge Context (>200K) confirmedYesNonehttps://ai.google.dev/pricingTwo tiers confirmed on AI Studio page.
Googlegemini-3-flash-previewInput $0.50/MTok (5e-7), Audio $1/MTok (1e-6), Output $3/MTok (3e-6), Cache read $0.05/MTok (5e-8)YesNo large-context tierYesNonehttps://ai.google.dev/pricingConfirmed on AI Studio page.
Googlegemini-3-pro-previewInput $2/MTok (2e-6) ≤200K, $4/MTok >200K; Output $12/MTok ≤200K, $18/MTok >200KNoLarge Context (>200K) set in fileNot applicableNonehttps://ai.google.dev/pricingNot listed on current AI Studio pricing page. Existing prices retained.
Googlegemini-2.0-flashInput $0.10/MTok (1e-7), Output $0.40/MTok (4e-7)NoDeprecated (shut down June 1, 2026)Not applicableNonehttps://ai.google.dev/pricingOfficially shut down June 1 2026. Existing prices retained for backward compatibility.
Googlegemini-2.0-flash-001Same as gemini-2.0-flashNoDeprecated (shut down June 1, 2026)Not applicableNonehttps://ai.google.dev/pricingOfficially shut down. Existing prices retained for backward compatibility.
Googlegemini-3.6-flashInput $1.50/MTok (1.5e-6), Output $7.50/MTok (7.5e-6), Cache read $0.15/MTok (1.5e-7)YesNo large-context tier on official pageYesNonehttps://ai.google.dev/pricingConfirmed on official AI Studio pricing page. Output ($7.50) is lower than gemini-3.5-flash ($9.00) — correct per official page.
Googlegemini-3.5-flash-liteInput $0.30/MTok (3e-7), Output $2.50/MTok (2.5e-6)YesNo large-context tierYesUpdatedhttps://ai.google.dev/pricingContext caching is "Not available" per official page (all tiers: Standard, Batch, Flex, Priority). Cache keys removed from pricing entry on 2026-07-23. Do NOT add cache pricing for this model without explicit official evidence.

Unresolved findings (updated 2026-07-23)

  1. gpt-5.3-codex — Confirmed in OpenAI models list, described as "most agentic coding model". Not in selectable models or pricing file. Not adding without explicit official pricing page confirmation.

  2. gpt-5.5 / gpt-5.4 large-context tiers — OpenAI page mentions extended context with doubled rates for these models, but the exact tier threshold (possibly 200K) is not confirmed. No Large Context tiers in the file for these models. Future audits should verify.

  3. claude-sonnet-4-5-20250929 Large Context tier — The file has a Large Context (>200K) tier for this model. The official Anthropic page does not explicitly publish per-tier pricing for this model separately. Future audits should verify.

  4. claude-sonnet-5 introductory pricing — Introductory pricing ($2/$10/MTok) expires August 31, 2026. Standard pricing ($3/$15/MTok, cache $3.75/$6/$0.30) takes effect September 1, 2026. The pricing file must be updated before or on September 1, 2026.