.agents/skills/add-model-price/references/model-audit-memory.md
This file is an optional snapshot of the latest automated audit whose per-model results add useful context for a future run. It is orientation only; reconfirm every price and tier against official provider sources before making a change or reporting a row as confirmed.
The audit agent may replace the snapshot below with its complete current table. Keep only one snapshot, never append an unbounded run history, never persist a partial set of checked models, and do not update this file only to refresh the audit date.
Audit date: 2026-07-27
All prices listed as $X / MTok (per million tokens). Per-token JSON values: divide by 1,000,000.
| Provider | Model / pricing entry | Pricing checked | Price confirmed | Tiering checked | Tiering correct | Change | Official source(s) | Comments |
|---|---|---|---|---|---|---|---|---|
| Anthropic | claude-fable-5 | Input $10/MTok (10e-6), Output $50/MTok (50e-6), 5m cache write $12.5/MTok (12.5e-6), 1h cache write $20/MTok (20e-6), cache read $1/MTok (1e-6) | Yes | Flat 1M context at standard pricing — no large-context tier | Yes | None | https://platform.claude.com/docs/en/about-claude/pricing | Confirmed. In flat long-context list. |
| Anthropic | claude-mythos-5 | Input $10/MTok (10e-6), Output $50/MTok (50e-6), 5m $12.5/MTok, 1h $20/MTok, read $1/MTok | Yes | Flat 1M context — no large-context tier | Yes | None | https://platform.claude.com/docs/en/about-claude/pricing | Limited availability (Project Glasswing). Same prices as Fable 5. |
| Anthropic | claude-opus-5 | Input $5/MTok (5e-6), Output $25/MTok (25e-6), 5m $6.25/MTok (6.25e-6), 1h $10/MTok (10e-6), read $0.50/MTok (0.5e-6) | Yes | Flat 1M context at standard pricing — no large-context tier | Yes | None | https://platform.claude.com/docs/en/about-claude/pricing https://platform.claude.com/docs/en/about-claude/models/overview | Added July 25 2026. API ID: claude-opus-5. Bedrock ID: anthropic.claude-opus-5. Google Cloud ID: claude-opus-5. In flat long-context list. Fast mode at $10/$50 MTok (shared with Opus 4.8). |
| Anthropic | claude-opus-4-8 | Input $5/MTok (5e-6), Output $25/MTok (25e-6), 5m $6.25/MTok (6.25e-6), 1h $10/MTok (10e-6), read $0.50/MTok (0.5e-6) | Yes | Flat 1M context — no large-context tier | Yes | None | https://platform.claude.com/docs/en/about-claude/pricing | Fast mode at $10/$50 per input/output MTok (additional SKU). |
| Anthropic | claude-opus-4-7 | Input $5/MTok (5e-6), Output $25/MTok (25e-6), 5m $6.25/MTok, 1h $10/MTok, read $0.50/MTok | Yes | Flat 1M context — no large-context tier | Yes | None | https://platform.claude.com/docs/en/about-claude/pricing | Confirmed. |
| Anthropic | claude-opus-4-6 | Input $5/MTok (5e-6), Output $25/MTok (25e-6), 5m $6.25/MTok, 1h $10/MTok, read $0.50/MTok | Yes | Flat 1M context — no large-context tier | Yes | None | https://platform.claude.com/docs/en/about-claude/pricing | inference_geo: "us" adds 1.1× multiplier for this and later models. |
| Anthropic | claude-opus-4-5-20251101 | Input $5/MTok (5e-6), Output $25/MTok (25e-6), 5m $6.25/MTok, 1h $10/MTok, read $0.50/MTok | Yes | Flat 1M context at standard pricing per official page | Yes | None | https://platform.claude.com/docs/en/about-claude/pricing https://platform.claude.com/docs/en/about-claude/models/overview | "Claude Opus 4.5" now explicitly listed on official pricing page (confirmed July 27 2026). API ID: claude-opus-4-5-20251101, alias claude-opus-4-5. matchPattern covers both. |
| Anthropic | claude-opus-4-1-20250805 | Input $15/MTok (15e-6), Output $75/MTok (75e-6), 5m $18.75/MTok, 1h $30/MTok, read $1.50/MTok | Yes | Deprecated model — no tiering | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Deprecated. Retires August 5, 2026. Retained for backward pricing compatibility. |
| Anthropic | claude-opus-4-20250514 | Input $15/MTok (15e-6), Output $75/MTok (75e-6), 5m $18.75/MTok, 1h $30/MTok, read $1.50/MTok | No | Retired model | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Retired except on Google Cloud. Not on current pricing page main table. |
| Anthropic | claude-sonnet-5 | Input $2/MTok (2e-6), Output $10/MTok (10e-6), 5m $2.50/MTok, 1h $4/MTok, read $0.20/MTok | Yes | Flat 1M context; introductory pricing through August 31 2026 | Yes | None | https://platform.claude.com/docs/en/about-claude/pricing | Introductory pricing confirmed through Aug 31 2026. Standard from Sep 1 2026: $3/$15, cache $3.75/$6/$0.30. File must be updated after Aug 31 2026. |
| Anthropic | claude-sonnet-4-6 | Input $3/MTok (3e-6), Output $15/MTok (15e-6), 5m $3.75/MTok, 1h $6/MTok, read $0.30/MTok | Yes | Flat 1M context — no large-context tier | Yes | None | https://platform.claude.com/docs/en/about-claude/pricing | Confirmed. |
| Anthropic | claude-sonnet-4-5-20250929 | Input $3/MTok (3e-6), Output $15/MTok (15e-6), 5m $3.75/MTok, 1h $6/MTok, read $0.30/MTok | Yes | Large Context >200K: 2× input / 1.5× output — tier in file, not explicitly published by Anthropic | No | None | https://platform.claude.com/docs/en/about-claude/pricing | Sonnet 4.5 NOT on flat long-context list. Anthropic page does not publish per-tier pricing for this model. The Large Context tier was set when model was first added. Unresolved. |
| Anthropic | claude-sonnet-4-20250514 | Input $3/MTok (3e-6), Output $15/MTok (15e-6), 5m $3.75/MTok, 1h $6/MTok, read $0.30/MTok | Yes | Retired model | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Retired except on Bedrock and Google Cloud. |
| Anthropic | claude-haiku-4-5-20251001 | Input $1/MTok (1e-6), Output $5/MTok (5e-6), 5m $1.25/MTok (1.25e-6), 1h $2/MTok (2e-6), read $0.10/MTok (1e-7) | Yes | No large-context tier | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Confirmed. Haiku 4.5 NOT on flat long-context list. |
| Anthropic | claude-3-5-haiku-20241022 | Input $0.80/MTok (8e-7), Output $4/MTok (4e-6), 5m $1/MTok, 1h $1.60/MTok, read $0.08/MTok | Yes | Retired model | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Retired except on Bedrock and Google Cloud. |
| Anthropic | claude-3.7-sonnet-20250219 | Input $3/MTok, Output $15/MTok, cache $3.75/$6/$0.30 per MTok | No | Not on current page | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Not listed on current official pricing page. Legacy prices retained. |
| Anthropic | claude-3.5-sonnet-20241022 | Input $3/MTok, Output $15/MTok, cache $3.75/$6/$0.30 per MTok | No | Not on current page | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Not listed on current official pricing page. Legacy prices retained. |
| Anthropic | claude-3-5-sonnet-20240620 | Input $3/MTok, Output $15/MTok, cache $3.75/$6/$0.30 per MTok | No | Not on current page | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Not listed on current official pricing page. Legacy prices retained. |
| Anthropic | claude-3-opus-20240229 | Input $15/MTok, Output $75/MTok | No | Not on current page | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Legacy. Not on current pricing page. |
| Anthropic | claude-3-sonnet-20240229 | Input $3/MTok, Output $15/MTok, cache set | No | Not on current page | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Legacy. Not on current pricing page. |
| Anthropic | claude-3-haiku-20240307 | Input $0.25/MTok, Output $1.25/MTok | No | Not on current page | Not applicable | None | https://platform.claude.com/docs/en/about-claude/pricing | Legacy. Not on current pricing page. |
| OpenAI | gpt-5.3-codex | Input $1.75/MTok (1.75e-6), Cached $0.175/MTok (0.175e-6), Output $14.00/MTok (14e-6) | Yes | No large-context tier; 400k context window (single tier) | Yes | Added | https://developers.openai.com/api/docs/pricing https://developers.openai.com/api/docs/models/gpt-5.3-codex | "Most capable agentic coding model." 400k context window, 128k max output. No date-stamped snapshot at launch. Added to pricing file and openAIModels in July 27 2026 audit. |
| OpenAI | gpt-5.6-sol | Input $5/MTok (5e-6), Cached $0.50/MTok (0.5e-6), Output $30/MTok (30e-6) | Yes | Large Context (>272K): input $10/MTok, cached $1/MTok, output $45/MTok | Yes | None | https://developers.openai.com/api/docs/pricing | Reasoning model. 272K threshold unique to gpt-5.6 family. |
| OpenAI | gpt-5.6-terra | Input $2.50/MTok (2.5e-6), Cached $0.25/MTok (0.25e-6), Output $15/MTok (15e-6) | Yes | Large Context (>272K): input $5/MTok, cached $0.50/MTok, output $22.50/MTok | Yes | None | https://developers.openai.com/api/docs/pricing | Reasoning model. |
| OpenAI | gpt-5.6-luna | Input $1/MTok (1e-6), Cached $0.10/MTok (0.1e-6), Output $6/MTok (6e-6) | Yes | Large Context (>272K): input $2/MTok, cached $0.20/MTok, output $9/MTok | Yes | None | https://developers.openai.com/api/docs/pricing | Reasoning model. |
| OpenAI | gpt-5.5-2026-04-23 (also matches gpt-5.5) | Input $5/MTok (5e-6), Cached $0.50/MTok (0.5e-6), Output $30/MTok (30e-6) | Yes | No large-context tier in file; large context mentioned by OpenAI but threshold not confirmed | No | None | https://developers.openai.com/api/docs/pricing | Reasoning model. matchPattern covers plain gpt-5.5 (dated suffix is optional). Large-context threshold unresolved. |
| OpenAI | gpt-5.5-pro-2026-04-23 (also matches gpt-5.5-pro) | Input $30/MTok (30e-6), Output $180/MTok (180e-6); no cache | Yes | No tiering in file | Not applicable | None | https://developers.openai.com/api/docs/pricing | Reasoning model; no cached input pricing. matchPattern covers plain gpt-5.5-pro. |
| OpenAI | gpt-5.4 | Input $2.50/MTok (2.5e-6), Cached $0.25/MTok (0.25e-6), Output $15/MTok (15e-6) | Yes | No large-context tier confirmed | No | None | https://developers.openai.com/api/docs/pricing | Reasoning model. Large-context threshold unresolved. |
| OpenAI | gpt-5.4-2026-03-05 | Input $2.50/MTok, Cached $0.25/MTok, Output $15/MTok | Yes | No tiering | Not applicable | None | https://developers.openai.com/api/docs/pricing | Dated snapshot for gpt-5.4. |
| OpenAI | gpt-5.4-pro | Input $30/MTok (30e-6), Output $180/MTok (180e-6); no cache | Yes | No tiering | Not applicable | None | https://developers.openai.com/api/docs/pricing | Reasoning model; no cached input pricing. |
| OpenAI | gpt-5.4-mini | Input $0.75/MTok (0.75e-6), Cached $0.075/MTok (0.075e-6), Output $4.50/MTok (4.5e-6) | Yes | No tiering | Not applicable | None | https://developers.openai.com/api/docs/pricing | Non-reasoning model. Confirmed. |
| OpenAI | gpt-5.4-nano | Input $0.20/MTok (0.2e-6), Cached $0.02/MTok (0.02e-6), Output $1.25/MTok (1.25e-6) | Yes | No tiering | Not applicable | None | https://developers.openai.com/api/docs/pricing | Non-reasoning model. Confirmed. |
| OpenAI | gpt-5-chat-latest | Input $1.25/MTok (1.25e-6), Cached $0.125/MTok (1.25e-7), Output $10/MTok (10e-6) | No | No provider tiering | Not applicable | None | https://developers.openai.com/api/docs/pricing | Prices match the file. Not re-verified directly this run (page did not list it in the summary). Retained from July 25 audit. |
| gemini-2.5-flash | Input $0.30/MTok (3e-7), Audio $1/MTok (1e-6), Output $2.50/MTok (2.5e-6), Cache read $0.03/MTok (3e-8) | Yes | No large-context tier; single tier confirmed | Yes | None | https://ai.google.dev/pricing | Cache read = 10% of standard input. Confirmed. | |
| gemini-2.5-flash-lite | Input $0.10/MTok (1e-7), Audio $0.30/MTok (3e-7), Output $0.40/MTok (4e-7), Cache read $0.01/MTok (1e-8) | Yes | No large-context tier | Yes | None | https://ai.google.dev/pricing | Cache read = 10% of input. Confirmed. | |
| gemini-2.5-pro | Input $1.25/MTok (1.25e-6) ≤200K, $2.50/MTok >200K; Output $10/MTok ≤200K, $15/MTok >200K; Cache read $0.125/MTok ≤200K | Yes | Large Context (>200K) confirmed | Yes | None | https://ai.google.dev/pricing | Two tiers confirmed. Cache = 10% of input at each tier. | |
| gemini-3.5-flash | Input $1.50/MTok (1.5e-6), Output $9.00/MTok (9e-6), Cache read $0.15/MTok (1.5e-7) | Yes | No large-context tier | Yes | None | https://ai.google.dev/pricing | Cache = 10% of input. Confirmed. | |
| gemini-3.1-flash-lite | Input $0.25/MTok (2.5e-7), Audio $0.50/MTok (5e-7), Output $1.50/MTok (1.5e-6), Cache read $0.025/MTok (2.5e-8) | No | No large-context tier | No | None | https://ai.google.dev/pricing | Contradictory fetch results: July 23 2026 explicit page fetch confirmed caching; July 25 and July 27 2026 summary fetches say "Not available." Per skill guidance, do not remove cache pricing based on a summary artifact alone. Cache pricing retained. Future audits should re-verify with an explicit page fetch. | |
| gemini-3.1-flash-lite-preview | Input $0.25/MTok, Output $1.50/MTok (same as GA) | No | No large-context tier | Not applicable | None | https://ai.google.dev/pricing | Preview variant; same prices as GA version in file; not separately listed on official page. | |
| gemini-3.1-pro-preview | Input $2/MTok (2e-6) ≤200K, $4/MTok >200K; Output $12/MTok ≤200K, $18/MTok >200K | Yes | Large Context (>200K) confirmed | Yes | None | https://ai.google.dev/pricing | Two tiers confirmed on AI Studio page. | |
| gemini-3-flash-preview | Input $0.50/MTok (5e-7), Audio $1/MTok (1e-6), Output $3/MTok (3e-6), Cache read $0.05/MTok (5e-8) | Yes | No large-context tier | Yes | None | https://ai.google.dev/pricing | Confirmed on AI Studio page. | |
| gemini-3-pro-preview | Input $2/MTok (2e-6) ≤200K, $4/MTok >200K; Output $12/MTok ≤200K, $18/MTok >200K | No | Large Context (>200K) set in file | Not applicable | None | https://ai.google.dev/pricing | Not listed on current AI Studio pricing page. Existing prices retained. | |
| gemini-2.0-flash | Input $0.10/MTok (1e-7), Output $0.40/MTok (4e-7) | No | Deprecated (shut down June 1, 2026) | Not applicable | None | https://ai.google.dev/pricing | Officially shut down June 1 2026. Existing prices retained for backward compatibility. | |
| gemini-2.0-flash-001 | Same as gemini-2.0-flash | No | Deprecated (shut down June 1, 2026) | Not applicable | None | https://ai.google.dev/pricing | Officially shut down. Existing prices retained for backward compatibility. | |
| gemini-3.6-flash | Input $1.50/MTok (1.5e-6), Output $7.50/MTok (7.5e-6), Cache read $0.15/MTok (1.5e-7) | Yes | No large-context tier on official page | Yes | None | https://ai.google.dev/pricing | Confirmed on official AI Studio pricing page. Output ($7.50) is lower than gemini-3.5-flash ($9.00) — correct per official page. | |
| gemini-3.5-flash-lite | Input $0.30/MTok (3e-7), Output $2.50/MTok (2.5e-6) | Yes | No large-context tier | Yes | None | https://ai.google.dev/pricing | Context caching is "Not available" per official page (all tiers: Standard, Batch, Flex, Priority). Confirmed by both July 25 audit and July 27 2026 fetch. Cache keys absent from pricing entry. Do NOT add cache pricing for this model without explicit official evidence. |
gpt-5.5 / gpt-5.4 large-context tiers — OpenAI page mentions extended context with doubled rates for these models, but the exact tier threshold (possibly 200K) is not confirmed. No Large Context tiers in the file for these models. Future audits should verify.
claude-sonnet-4-5-20250929 Large Context tier — The file has a Large Context (>200K) tier for this model. The official Anthropic page does not explicitly publish per-tier pricing for this model separately. Future audits should verify.
claude-sonnet-5 introductory pricing — Introductory pricing ($2/$10/MTok) expires August 31, 2026. Standard pricing ($3/$15/MTok, cache $3.75/$6/$0.30) takes effect September 1, 2026. The pricing file must be updated before or on September 1, 2026.
claude-opus-4-1-20250805 retirement — Deprecated and retiring August 5, 2026. File entry retained for backward pricing compatibility. No action required; the entry stays so historical traces can still be priced.
gemini-3.1-flash-lite cache pricing — Contradictory fetch results across multiple audit runs. July 23 2026 explicit page fetch confirmed caching at $0.025/MTok; subsequent summary-level fetches say "Not available". Per skill guidance, cache pricing is retained until a definitive explicit page fetch confirms its removal. Future audits should fetch the model-specific row on the AI Studio pricing page.