Back to Langfuse

Model Price Audit Memory

.agents/skills/add-model-price/references/model-audit-memory.md

4.6.020.5 KB
Original Source

Model Price Audit Memory

This file is an optional snapshot of the latest automated audit whose per-model results add useful context for a future run. It is orientation only; reconfirm every price and tier against official provider sources before making a change or reporting a row as confirmed.

The audit agent may replace the snapshot below with its complete current table. Keep only one snapshot, never append an unbounded run history, never persist a partial set of checked models, and do not update this file only to refresh the audit date.

Latest useful snapshot

Audit date: 2026-08-04

All prices listed as $X / MTok (per million tokens). Per-token JSON values: divide by 1,000,000.

ProviderModel / pricing entryPricing checkedPrice confirmedTiering checkedTiering correctChangeOfficial source(s)Comments
Anthropicclaude-fable-5Input $10/MTok, Output $50/MTok, 5m $12.50/MTok, 1h $20/MTok, read $1/MTokYesFlat 1M context at standard pricingYesNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed via full pricing table fetch.
Anthropicclaude-mythos-5Same as Fable 5YesFlat 1M contextYesNonehttps://platform.claude.com/docs/en/about-claude/pricingLimited availability (Project Glasswing). Confirmed.
Anthropicclaude-opus-5Input $5/MTok, Output $25/MTok, 5m $6.25/MTok, 1h $10/MTok, read $0.50/MTokYesFlat 1M contextYesNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed. Fast mode $10/$50 MTok.
Anthropicclaude-opus-4-8Same as Opus 5YesFlat 1M contextYesNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed.
Anthropicclaude-opus-4-7Same as Opus 5YesFlat 1M contextYesNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed.
Anthropicclaude-opus-4-6Same as Opus 5YesFlat 1M contextYesNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed. inference_geo: "us" adds 1.1x.
Anthropicclaude-opus-4-5-20251101Same as Opus 5YesFlat 1M contextYesNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed.
Anthropicclaude-opus-4-1-20250805Input $15/MTok, Output $75/MTok, 5m $18.75/MTok, 1h $30/MTok, read $1.50/MTokYesDeprecated — no tieringNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingDeprecated, retires August 5, 2026 (tomorrow as of this audit). Still listed on official page; entry retained.
Anthropicclaude-opus-4-20250514Input $15/MTok, Output $75/MTok, 5m $18.75/MTok, 1h $30/MTok, read $1.50/MTokYesRetired except Google Cloud — no tieringNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed present on current page's main table this run (row: "Claude Opus 4 (retired, except on Google Cloud)"), correcting a prior audit's belief it was absent from the page.
Anthropicclaude-sonnet-5Input $2/MTok, Output $10/MTok (through Aug 31, 2026); 5m $2.50/MTok, 1h $4/MTok, read $0.20/MTokYesFlat 1M context; introductory pricing through Aug 31 2026YesNonehttps://platform.claude.com/docs/en/about-claude/pricingStandard pricing $3/$15 (cache $3.75/$6/$0.30) takes effect Sep 1, 2026 — update the file then.
Anthropicclaude-sonnet-4-6Input $3/MTok, Output $15/MTok, 5m $3.75/MTok, 1h $6/MTok, read $0.30/MTokYesFlat 1M contextYesNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed.
Anthropicclaude-sonnet-4-5-20250929Input $3/MTok, Output $15/MTok, 5m $3.75/MTok, 1h $6/MTok, read $0.30/MTokYesNo large-context tier (200k hard context-window cap)YesUpdatedhttps://platform.claude.com/docs/en/about-claude/pricing https://platform.claude.com/docs/en/build-with-claude/context-windowsRemoved an incorrect "Large Context (>200K)" tier. The context-windows page confirms Sonnet 4.5 has a hard 200k-token context window (not on the 1M list) and that exceeding it returns a 400 error rather than being billed at a premium, so the tier's condition could never legitimately fire. Resolves a finding left unresolved since at least the June 2026 audit.
Anthropicclaude-sonnet-4-20250514Input $3/MTok, Output $15/MTok, 5m $3.75/MTok, 1h $6/MTok, read $0.30/MTokYesRetired except Bedrock/Google Cloud — no tieringNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed present on current page's main table this run.
Anthropicclaude-haiku-4-5-20251001Input $1/MTok, Output $5/MTok, 5m $1.25/MTok, 1h $2/MTok, read $0.10/MTokYesNo large-context tier (200k context window, not on flat 1M list)Not applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed.
Anthropicclaude-3-5-haiku-20241022Input $0.80/MTok, Output $4/MTok, 5m $1/MTok, 1h $1.60/MTok, read $0.08/MTokYesRetired except Bedrock/Google CloudNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingConfirmed present on current page's main table this run ("Claude Haiku 3.5").
Anthropicclaude-3.7-sonnet-20250219Input $3/MTok, Output $15/MTok, cache $3.75/$6/$0.30NoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingNot on current page this run either. Legacy prices retained, not re-verified.
Anthropicclaude-3.5-sonnet-20241022Input $3/MTok, Output $15/MTok, cache $3.75/$6/$0.30NoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingNot re-verified this run. Legacy prices retained. Bedrock has a separate, higher "Public Extended Access" SKU — see comments on Bedrock below; not representable in the schema.
Anthropicclaude-3-5-sonnet-20240620Input $3/MTok, Output $15/MTok, cache $3.75/$6/$0.30NoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingNot re-verified this run. Legacy prices retained. Same Bedrock Public Extended Access caveat as v2 above.
Anthropicclaude-3-opus-20240229Input $15/MTok, Output $75/MTokNoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingNot re-verified this run. Legacy.
Anthropicclaude-3-sonnet-20240229Input $3/MTok, Output $15/MTokNoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingNot re-verified this run. Legacy.
Anthropicclaude-3-haiku-20240307Input $0.25/MTok, Output $1.25/MTokNoNot on current pageNot applicableNonehttps://platform.claude.com/docs/en/about-claude/pricingNot re-verified this run. Legacy.
AWS Bedrockclaude-3-5-sonnet-20240620 / claude-3.5-sonnet-20241022 (Public Extended Access SKU)$6.00/MTok input, $30.00/MTok output, $7.50/MTok cache write, $0.60/MTok cache readYes (SKU confirmed real)Distinct dated SKU, not a context-length tierNot applicableUnresolvedhttps://aws.amazon.com/bedrock/pricing/Confirmed via targeted verbatim fetch that this is a real, distinct "Public Extended Access, Effective 1 Dec 2025" Bedrock SKU at double the standard $3/$15 rate for the same model IDs. Not actionable: Langfuse's schema matches by model-ID string only and cannot distinguish which Bedrock billing tier a given request used, so no file change was made. Documented as a permanent, confirmed limitation — see provider-sources-and-price-keys.md.
OpenAIgpt-5.6-solInput $5/MTok, Cached $0.50/MTok, Cache write $6.25/MTok, Output $30/MTokYesLarge Context (>272K): $10/$1.00/$12.50/$45YesNonehttps://developers.openai.com/api/docs/pricing https://developers.openai.com/api/docs/models/gpt-5.6-solConfirmed cache-write pricing (1.25x input) already correctly present in file.
OpenAIgpt-5.6-terraInput $2/MTok, Cached $0.20/MTok, Cache write $2.50/MTok, Output $12/MTokYesLarge Context (>272K): $4/$0.40/$5.00/$18YesNonehttps://developers.openai.com/api/docs/pricing https://developers.openai.com/api/docs/models/gpt-5.6-terraConfirmed stable since July 31 2026 price cut.
OpenAIgpt-5.6-lunaInput $0.20/MTok, Cached $0.02/MTok, Cache write $0.25/MTok, Output $1.20/MTokYesLarge Context (>272K): $0.40/$0.04/$0.50/$1.80YesNonehttps://developers.openai.com/api/docs/pricing https://developers.openai.com/api/docs/models/gpt-5.6-lunaConfirmed stable since July 31 2026 price cut.
OpenAIgpt-5.5-2026-04-23 (alias gpt-5.5)Input $5/MTok, Cached $0.50/MTok, Output $30/MTokYesLarge Context (>272K): $10/$1.00/$45YesUpdatedhttps://developers.openai.com/api/docs/pricing https://developers.openai.com/api/docs/models/gpt-5.5Added Large Context tier (272K threshold now confirmed, resolving a long-standing unresolved finding) plus missing cache_read_input_tokens/reasoning_tokens aliases. No cache-write pricing for this model (confirmed).
OpenAIgpt-5.5-pro-2026-04-23 (alias gpt-5.5-pro)Input $30/MTok, Output $180/MTok; no cacheYesLarge Context (>272K): $60/$270YesUpdatedhttps://developers.openai.com/api/docs/pricing https://developers.openai.com/api/docs/models/gpt-5.5-proAdded Large Context tier plus missing reasoning_tokens alias.
OpenAIgpt-5.4Input $2.50/MTok, Cached $0.25/MTok, Output $15/MTokYesLarge Context (>272K): $5.00/$0.50/$22.50YesUpdatedhttps://developers.openai.com/api/docs/pricing https://developers.openai.com/api/docs/models/gpt-5.4Added Large Context tier plus missing cache_read_input_tokens/reasoning_tokens aliases.
OpenAIgpt-5.4-2026-03-05Same as gpt-5.4YesLarge Context (>272K): $5.00/$0.50/$22.50YesUpdatedhttps://developers.openai.com/api/docs/pricing https://developers.openai.com/api/docs/models/gpt-5.4Dated snapshot sibling of gpt-5.4; same fix applied.
OpenAIgpt-5.4-proInput $30/MTok, Output $180/MTok; no cacheYesLarge Context (>272K): $60/$270YesUpdatedhttps://developers.openai.com/api/docs/pricing https://developers.openai.com/api/docs/models/gpt-5.4Added Large Context tier plus missing reasoning_tokens alias.
OpenAIgpt-5.4-pro-2026-03-05Same as gpt-5.4-proYesLarge Context (>272K): $60/$270YesUpdatedhttps://developers.openai.com/api/docs/pricing https://developers.openai.com/api/docs/models/gpt-5.4Dated snapshot sibling; same fix applied.
OpenAIgpt-5.4-miniInput $0.75/MTok, Cached $0.075/MTok, Output $4.50/MTokYesNo large-context tier (dashes confirmed in official table)YesNonehttps://developers.openai.com/api/docs/pricingConfirmed no large-context tier applies; left unchanged.
OpenAIgpt-5.4-mini-2026-03-17Same as gpt-5.4-miniYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingDated snapshot sibling; confirmed unchanged.
OpenAIgpt-5.4-nanoInput $0.20/MTok, Cached $0.02/MTok, Output $1.25/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-5.4-nano-2026-03-17Same as gpt-5.4-nanoYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingDated snapshot sibling; confirmed unchanged.
OpenAIgpt-5.3-codexInput $1.75/MTok, Cached $0.175/MTok, Output $14.00/MTokYesNo large-context tier (400k context window, single tier)YesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-5.2-2025-12-11Input $1.75/MTok, Cached $0.175/MTok, Output $14.00/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-5.1-2025-11-13Input $1.25/MTok, Cached $0.125/MTok, Output $10/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-5-2025-08-07Input $1.25/MTok, Cached $0.125/MTok, Output $10/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-5-mini-2025-08-07Input $0.25/MTok, Cached $0.025/MTok, Output $2/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-5-nano-2025-08-07Input $0.05/MTok, Cached $0.005/MTok, Output $0.40/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-5-pro-2025-10-06Input $15/MTok, Output $120/MTok; no cacheYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/models/gpt-5-proConfirmed unchanged via dedicated model page.
OpenAIgpt-5.2-pro-2025-12-11Input $21/MTok, Output $168/MTok; no cacheYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/models/gpt-5.2-proConfirmed unchanged via dedicated model page.
OpenAIgpt-4.1-2025-04-14Input $2/MTok, Cached $0.50/MTok, Output $8/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-4.1-mini-2025-04-14Input $0.40/MTok, Cached $0.10/MTok, Output $1.60/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-4.1-nano-2025-04-14Input $0.10/MTok, Cached $0.025/MTok, Output $0.40/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-4o-2024-08-06Input $2.50/MTok, Cached $1.25/MTok, Output $10/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-4o-mini-2024-07-18Input $0.15/MTok, Cached $0.075/MTok, Output $0.60/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIo3-2025-04-16Input $2/MTok, Cached $0.50/MTok, Output $8/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIo3-mini-2025-01-31Input $1.10/MTok, Cached $0.55/MTok, Output $4.40/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIo4-mini-2025-04-16Input $1.10/MTok, Cached $0.275/MTok, Output $4.40/MTokYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-3.5-turboInput $0.50/MTok, Output $1.50/MTok; no cacheYesNo large-context tierYesNonehttps://developers.openai.com/api/docs/pricingConfirmed unchanged.
OpenAIgpt-5-chat-latestInput $1.25/MTok, Cached $0.125/MTok, Output $10/MTokNoNo provider tieringNot applicableNonehttps://developers.openai.com/api/docs/models/gpt-5-chat-latestNot re-verified this run; retained from July 2026 audit.
Googlegemini-2.5-flashInput $0.30/MTok, Audio $1/MTok, Output $2.50/MTok, Cache read $0.03/MTok (audio $0.10/MTok)YesNo large-context tierYesNonehttps://ai.google.dev/pricingConfirmed unchanged.
Googlegemini-2.5-flash-liteInput $0.10/MTok, Audio $0.30/MTok, Output $0.40/MTok, Cache read $0.01/MTok (audio $0.03/MTok)YesNo large-context tierYesNonehttps://ai.google.dev/pricingConfirmed unchanged.
Googlegemini-2.5-proInput $1.25/$2.50 MTok (≤200K/>200K), Output $10/$15, Cache read $0.125/$0.25YesLarge Context (>200K) confirmedYesNonehttps://ai.google.dev/pricingConfirmed unchanged.
Googlegemini-3.5-flashInput $1.50/MTok, Output $9.00/MTok, Cache read $0.15/MTokYesNo large-context tierYesNonehttps://ai.google.dev/pricingConfirmed unchanged.
Googlegemini-3.5-flash-liteInput $0.30/MTok, Output $2.50/MTok, Cache read $0.03/MTokYesNo large-context tierYesUpdatedhttps://ai.google.dev/pricing https://ai.google.dev/gemini-api/docs/pricingRe-added cache-read pricing after two independent verbatim, column-explicit fetches confirmed the Paid tier has context caching at $0.03/MTok (10% of input, matching Google's universal ratio); only the Free tier says "Not available". Reverses the July 22-31 2026 audits' conclusion that caching was unavailable on all tiers — see provider-sources-and-price-keys.md for the full flip-flop history and the lesson learned.
Googlegemini-3.1-flash-liteInput $0.25/$0.50 (text/audio), Output $1.50, Cache read $0.025/$0.05YesNo large-context tierYesNonehttps://ai.google.dev/pricingConfirmed unchanged.
Googlegemini-3.1-flash-lite-previewSame as GA gemini-3.1-flash-liteNoNo large-context tierNot applicableNonehttps://ai.google.dev/pricingNot separately listed on official page this run either; not re-verified.
Googlegemini-3.1-pro-previewInput $2/$4 MTok (≤200K/>200K), Output $12/$18YesLarge Context (>200K) confirmedYesNonehttps://ai.google.dev/pricingConfirmed unchanged.
Googlegemini-3-flash-previewInput $0.50/$1.00 (text/audio), Output $3.00, Cache read $0.05/$0.10YesNo large-context tierYesNonehttps://ai.google.dev/pricingConfirmed unchanged.
Googlegemini-3-pro-previewInput $2/$4 MTok (≤200K/>200K), Output $12/$18NoLarge Context (>200K) set in fileNot applicableNonehttps://ai.google.dev/pricingStill not listed on official AI Studio page this run; existing prices retained, not re-verified.
Googlegemini-3.6-flashInput $1.50/MTok, Output $7.50/MTok, Cache read $0.15/MTokYesNo large-context tierYesNonehttps://ai.google.dev/pricingConfirmed unchanged.
Googlegemini-2.0-flashInput $0.10/MTok, Output $0.40/MTokNoDeprecated (shut down June 1, 2026)Not applicableNonehttps://ai.google.dev/pricingNot re-verified this run; retained for backward compatibility.
Googlegemini-2.0-flash-001Same as gemini-2.0-flashNoDeprecated (shut down June 1, 2026)Not applicableNonehttps://ai.google.dev/pricingNot re-verified this run; retained for backward compatibility.

Unresolved findings (updated 2026-08-04)

  1. claude-sonnet-5 introductory pricing — Introductory pricing ($2/$10/MTok) expires August 31, 2026. Standard pricing ($3/$15/MTok, cache $3.75/$6/$0.30) takes effect September 1, 2026. The pricing file must be updated before or on that date.

  2. claude-opus-4-1-20250805 retirement — Deprecated, retiring August 5, 2026 (i.e. the day after this audit). Still listed on the official pricing page as of this run. File entry retained for backward pricing compatibility; no action required now, but check whether the model is fully removed from the official page in the next audit.

  3. AWS Bedrock "Claude 3.5 Sonnet (Public Extended Access)" pricing — Confirmed real this run (see the dedicated row above and provider-sources-and-price-keys.md), but not representable in Langfuse's schema because it matches by model-ID string only. No file change should be made for this; treat it as a permanent, documented limitation rather than something to re-investigate each run.

  4. Legacy Claude 3.x / 3.5 / 3.7 models not on the current pricing page — Not re-verified this run (claude-3.7-sonnet-20250219, claude-3.5-sonnet-20241022, claude-3-5-sonnet-20240620, claude-3-opus-20240229, claude-3-sonnet-20240229, claude-3-haiku-20240307). Existing prices retained. Low priority since these are retired/legacy, but a future audit could check the model-deprecations page directly if there is reason to believe pricing changed.

  5. gemini-3.1-flash-lite-preview and gemini-3-pro-preview — Still not separately listed on the official AI Studio pricing page. Existing prices retained without fresh confirmation. Re-verify if these move from preview to GA or gain their own pricing row.

  6. OpenAI "cache writes" may expand beyond the gpt-5.6 family — As of this audit, the distinct 1.25x-of-input cache-write billing dimension applies only to gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna. Future audits should re-check the OpenAI pricing table's "cache writes" column whenever a new reasoning model is added, since this could spread to other model families.