LLM · Price Tracker

Sources

The 52 official pricing pages that Model Price Watch cites. The tracker reads each of them once a day and extracts every model's price with a tested recipe.

Last run
Pages healthy
51 of 52
Fetched, and every model extracted
Models extracted
169 of 193
On the last run
Untracked models
23
Retired or without a public price
Differences from MPW
5
5 explained by page evidence

How tracking works

Official pages

Official pricing pages
ProviderPageStatusHTTPPage changedDays OKModelsUntrackedDifferencesDetails
Metaai.developer.meta.com/docs/pricing-rate-limits (opens in new tab)
via dev.meta.ai/docs/pricing-rate-limits.md
OK200no change seen13 of 300Details
Googleai.google.dev/gemini-api/docs/pricing (opens in new tab)
via ai.google.dev/gemini-api/docs/pricing.md.txt
OK200no change seen19 of 900Details
AI21 Labsai21.com/pricing (opens in new tab)OK200no change seen12 of 200Details
Alibabaalibabacloud.com/help/en/model-studio/model-pricing (opens in new tab)Failed200no change seen0 / 1 failed0 of 2400Details
DeepSeekapi-docs.deepseek.com/quick_start/pricing (opens in new tab)OK200no change seen12 of 200Details
Inceptionapi.inceptionlabs.ai/v1/models (opens in new tab)OK200no change seen11 of 100Details
Amazonaws.amazon.com/nova/pricing (opens in new tab)
via b0.p.awsstatic.com/pricing/2.0/meteredUnitMaps/bedrock/USD/current/bedrock.json
OK200no change seen13 of 300Details
Tencentcloud.tencent.com/document/product/1823/130055 (opens in new tab)OK200no change seen14 of 401Details
Coherecohere.com/pricing (opens in new tab)OK200no change seen11 of 110Details
Groqconsole.groq.com/docs/models (opens in new tab)
via console.groq.com/docs/models.md
OK200no change seen12 of 240Details
Sakana AIconsole.sakana.ai/pricing (opens in new tab)OK200no change seen11 of 100Details
DeepInfradeepinfra.com/pricing (opens in new tab)OK200no change seen15 of 511Details
IBMdevelopers.cloudflare.com/workers-ai/models/granite-4.0-h-micro (opens in new tab)
via developers.cloudflare.com/workers-ai/models/granite-4.0-h-micro/index.md
OK200no change seen11 of 100Details
OpenAIdevelopers.openai.com/api/docs/changelog (opens in new tab)
via developers.openai.com/api/docs/changelog.md
OK200no change seen11 of 100Details
OpenAIdevelopers.openai.com/api/docs/models/gpt-5.3-codex (opens in new tab)
via developers.openai.com/api/docs/models/gpt-5.3-codex.md
OK200no change seen11 of 100Details
OpenAIdevelopers.openai.com/api/docs/pricing.md (opens in new tab)OK200no change seen113 of 1300Details
OpenAIdevelopers.openai.com/api/docs/pricing (opens in new tab)
via developers.openai.com/api/docs/pricing.md
OK200no change seen13 of 300Details
Coheredocs.cohere.com/docs/command-a (opens in new tab)OK200no change seen11 of 100Details
Coheredocs.cohere.com/docs/command-r (opens in new tab)OK200no change seen11 of 100Details
Fireworksdocs.fireworks.ai/serverless/pricing (opens in new tab)
via docs.fireworks.ai/serverless/pricing.md
OK200no change seen14 of 4100Details
Inceptiondocs.inceptionlabs.ai/get-started/models (opens in new tab)
via docs.inceptionlabs.ai/get-started/models.md
OK200no change seen12 of 200Details
Perplexitydocs.perplexity.ai/docs/agent-api/models (opens in new tab)
via docs.perplexity.ai/docs/agent-api/models.md
OK200no change seen11 of 100Details
Rekadocs.reka.ai/pricing (opens in new tab)
via developer.reka.ai/models
OK200no change seen11 of 120Details
TypeSafe AIdocs.typesafe.ai/models (opens in new tab)
via docs.typesafe.ai/models.md
OK200no change seen11 of 100Details
Voyage AIdocs.voyageai.com/docs/pricing (opens in new tab)
via docs.voyageai.com/docs/pricing.md
OK200no change seen18 of 800Details
xAIdocs.x.ai/developers/models/grok-4.20 (opens in new tab)OK200no change seen11 of 100Details
xAIdocs.x.ai/developers/models/grok-4.3 (opens in new tab)OK200no change seen11 of 100Details
xAIdocs.x.ai/developers/models/grok-4.7 (opens in new tab)OK200no change seen11 of 100Details
xAIdocs.x.ai/developers/models/grok-build-0.1 (opens in new tab)OK200no change seen11 of 100Details
xAIdocs.x.ai/developers/models (opens in new tab)OK200no change seen12 of 200Details
Z.AIdocs.z.ai/guides/overview/pricing (opens in new tab)
via docs.z.ai/guides/overview/pricing.md
OK200no change seen121 of 2120Details
Fireworksfireworks.ai/models/fireworks/ember-1 (opens in new tab)OK200no change seen11 of 100Details
Groqgroq.com/pricing (opens in new tab)OK200no change seen10 of 010Details
IBMibm.com/products/watsonx-ai/pricing (opens in new tab)OK200no change seen12 of 200Details
Inference.netinference.net/models (opens in new tab)OK200no change seen12 of 200Details
Meituanlongcat.ai/platform/docs/Pricing/LongCat-2.0.html (opens in new tab)
via longcat.ai/platform/docs/pricing/longcat-2.0
OK200no change seen11 of 101Details
Mistralmistral.ai/pricing/api (opens in new tab)OK200no change seen16 of 610Details
Baichuannovita.ai/models (opens in new tab)OK200no change seen11 of 100Details
Apodexplatform.apodex.ai/docs/pricing (opens in new tab)OK200no change seen12 of 200Details
Anthropicplatform.claude.com/docs/en/about-claude/pricing.md (opens in new tab)OK200no change seen11 of 100Details
Anthropicplatform.claude.com/docs/en/about-claude/pricing (opens in new tab)
via platform.claude.com/docs/en/about-claude/pricing.md
OK200no change seen112 of 1200Details
Moonshotplatform.kimi.ai/docs/pricing/chat-k3 (opens in new tab)
via platform.kimi.ai/docs/pricing/chat.md
OK200no change seen11 of 100Details
Moonshotplatform.kimi.ai/docs/pricing/chat (opens in new tab)
via platform.kimi.ai/docs/pricing/chat.md
OK200no change seen13 of 310Details
MiniMaxplatform.minimax.io/docs/guides/pricing-paygo (opens in new tab)
via platform.minimax.io/docs/guides/pricing-paygo.md
OK200no change seen12 of 200Details
OpenAIplatform.openai.com/docs/pricing (opens in new tab)
via developers.openai.com/api/docs/pricing.md
OK200no change seen116 of 1600Details
Relacerelace.ai/pricing (opens in new tab)OK200no change seen12 of 200Details
Sakana AIsakana.ai/fugu (opens in new tab)OK200no change seen12 of 200Details
Kwaipilotstreamlake.com/product/wanqing (opens in new tab)OK200no change seen12 of 200Details
Thinking Machinestinker-docs.thinkingmachines.ai/tinker/models (opens in new tab)
via tinker-docs.thinkingmachines.ai/tinker/serverless.json
OK200no change seen12 of 200Details
Togethertogether.ai/pricing (opens in new tab)OK200no change seen19 of 901Details
Unbiasedunbiased.ai/pricing (opens in new tab)OK200no change seen11 of 100Details
Upstageupstage.ai/pricing/api (opens in new tab)OK200no change seen13 of 301Details

Page details

Meta ai.developer.meta.com/docs/pricing-rate-limits OK

Format markdown

Surface (checked 2026-10-08): page_url returns HTTP 500 (a 1542-byte Facebook error page, as in the hint local_probe) to the tracker's browser User-Agent, also with the .md suffix. With a non-browser User-Agent the same host returns HTTP 200, but it serves an older CMS copy (front matter 'cms: alias: /model-api/docs/pricing-rate-limits', no SAM 3.1 section). https://ai.developer.meta.com/ redirects (302) to https://dev.meta.ai/, Meta's current developer site. The fetch_url is its official Markdown twin (text/markdown, about 9 KB), linked from https://dev.meta.ai/llms.txt ('Meta Model API documentation'). The rendered page https://dev.meta.ai/docs/pricing-rate-limits (HTML, about 220 KB) shows the same Standard tier table. Muse Spark 1.3, 1.2 and 1.1 share one 'Standard tier' table (Usage | Price per 1M tokens: Cached input, Input, Output). Each model has one recipe per field. The section requires the model id (in backticks) on the 'Models:' line directly under the 'Standard tier' heading (an optional {#anchor} suffix is allowed). section_end stops at the next heading or 'Models:' line, so no recipe can read the Contributor table. Each pattern starts at a two-column table header 'Price per 1M tokens' (unit guard), steps only through table rows, and reads the exact label cell with a single price cell, so row order and added rows do not matter. The Contributor tier (muse-spark-1.3-contributor, muse-spark-1.2-contributor: training-eligible, 100 RPM) is intentionally not tracked, per MPW price_note. Web search grounding ($2.50 per 1,000 queries) is billed separately and is not part of the token rate.

Models on ai.developer.meta.com/docs/pricing-rate-limits
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Muse Spark 1.3$1.25$0.15$4.25$1.25$0.15$4.25Agree
Muse Spark 1.2$1.25$0.15$4.25$1.25$0.15$4.25Agree
Muse Spark 1.1$1.25$0.15$4.25$1.25$0.15$4.25Agree
Google ai.google.dev/gemini-api/docs/pricing OK

Format markdown

Fetches the official markdown twin pricing.md.txt (listed in https://ai.google.dev/gemini-api/docs/llms.txt). The twin is English for any accept-language; the HTML page without accept-language serves a translated page. The engine's default browser user-agent gets a 302 to /oauth2authorize?...prompt=none that loops without cookies ('redirect count exceeded'); a non-browser user-agent gets 200, so headers override user-agent. Every recipe starts at the model's exact '## <name>' heading and stops at the next '## ' heading or at the first '### ' heading that is not 'Standard', so it reads only the Standard table, paid column, even if Google reorders the Batch, Flex and Priority tables. Cross-checked on 2026-10-08: the Standard rows of all nine models in the twin match the English HTML page.

Models on ai.google.dev/gemini-api/docs/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Gemini 3.1 Pro$2.00$0.20$12.00$2.00$0.20$12.00Agree
Gemini 3.5 Flash$1.50$0.15$9.00$1.50not listed$9.00Agree
Gemini 3.6 Flash$0.75$0.075$3.75$0.75$0.075$3.75Agree
Gemini 3.7 Flash$0.75$0.075$3.75$0.75$0.075$3.75Agree
Gemini 3.8 Flash$0.75$0.075$3.75$0.75$0.075$3.75Agree
Gemini 3.5 Flash-Lite$0.30$0.03$2.50$0.30$0.03$2.50Agree
Gemini 3 Pro Image$2.00not listed$120.00$2.00not listed$120.00Agree
Gemini 3.1 Flash Image$0.50not listed$60.00$0.50not listed$60.00Agree
Gemini 3.1 Flash Lite Image$0.25not listed$30.00$0.25not listed$30.00Agree
AI21 Labs ai21.com/pricing OK

Format html

Server-rendered WordPress HTML (wp-json/wp/v2/pages/23501: status publish, modified 2026-05-06). The 'Foundation Models' section has one card per Jamba model: an h3 name line, an optional description line, then a footer that the <br> splits into '$X / 1M input tokens' and '$Y / 1M output tokens' lines. These are the AI21 platform pay-as-you-go prices: the Pay As You Go plan says 'See usage pricing for Foundation Models', and the Custom Plan only promises volume discounts without numbers. No cached-input or batch price is published (docs jamba-batch-api.md has no price). Alternatives checked on 2026-10-08: /pricing.md is 404, accept: text/markdown still returns HTML, www.ai21.com/llms.txt has no pricing entry, docs.ai21.com/docs/usage-cost.md defers to this page for platform prices, and the wp-json record only wraps the same rendered HTML (acf is empty). The page carries robots noindex and is not in the site navigation; if it disappears, check docs usage-cost.md for its replacement.

Models on ai21.com/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Jamba Large$2.00not listed$8.00$2.00not listed$8.00Agree
Jamba Mini$0.20not listed$0.40$0.20not listed$0.40Agree
Alibaba alibabacloud.com/help/en/model-studio/model-pricing Failed

Format html. Last error: empty page (HTTP 200)

Fetches the HTML page (format html). Every recipe reads the Singapore tab (rows whose Deployment scope cell is 'International') of the named sub-table, bounded by the sub-heading and the next 'China (Beijing)' tab label. An official Markdown twin exists at page_url + '.md' (MDX with JSX <table> markup and escaped '\$' and '\<'), but the normalized HTML text is more compact and its cells are simpler to step through, so HTML is used. Prices are the standard real-time list rates at the base (smallest) context tier; batch, cache and limited-time discounts are not tracked, except the qwen3.8-omni-flash cache-hit column and the derived 10% explicit-cache-hit rate for qwen3.7-max. qwen3.8-max gets no derived cache rate: the official Context Cache page lists it as an exception to the 10% rule, and its Model Info page prints a different cache price (see that recipe's notes).

Models on alibabacloud.com/help/en/model-studio/model-pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
QwQ-Plus
fetch empty page (HTTP 200)
not listednot listednot listed$0.80not listed$2.40MPW only
Qwen-Flash
fetch empty page (HTTP 200)
not listednot listednot listed$0.05not listed$0.40MPW only
Qwen-Plus
fetch empty page (HTTP 200)
not listednot listednot listed$0.40not listed$1.20MPW only
Qwen-Turbo
fetch empty page (HTTP 200)
not listednot listednot listed$0.05not listed$0.20MPW only
Qwen3-Max
fetch empty page (HTTP 200)
not listednot listednot listed$1.20not listed$6.00MPW only
Qwen3.6-Flash
fetch empty page (HTTP 200)
not listednot listednot listed$0.25not listed$1.50MPW only
Qwen3.7-Flash
fetch empty page (HTTP 200)
not listednot listednot listed$0.03not listed$0.13MPW only
Qwen3.7-Max
fetch empty page (HTTP 200)
not listednot listednot listed$2.50$0.25$7.50MPW only
Qwen3.7-Plus
fetch empty page (HTTP 200)
not listednot listednot listed$0.40not listed$1.60MPW only
Qwen3.8-Max
fetch empty page (HTTP 200)
not listednot listednot listed$2.00$0.20$6.00MPW only
Qwen3.8-2.4T-A95B
fetch empty page (HTTP 200)
not listednot listednot listed$2.00not listed$6.00MPW only
Qwen3.8-27B
fetch empty page (HTTP 200)
not listednot listednot listed$0.50not listed$3.00MPW only
Qwen3.8-Flash
fetch empty page (HTTP 200)
not listednot listednot listed$0.15not listed$0.47MPW only
Qwen3.8-Omni-Flash
fetch empty page (HTTP 200)
not listednot listednot listed$0.15$0.016$0.47MPW only
Qwen3.6-Plus
fetch empty page (HTTP 200)
not listednot listednot listed$0.50not listed$3.00MPW only
Qwen3.6-27B
fetch empty page (HTTP 200)
not listednot listednot listed$0.60not listed$3.60MPW only
Qwen3.5-397B-A17B
fetch empty page (HTTP 200)
not listednot listednot listed$0.60not listed$3.60MPW only
Qwen3.5-Plus
fetch empty page (HTTP 200)
not listednot listednot listed$0.40not listed$2.40MPW only
Qwen3.5-Flash
fetch empty page (HTTP 200)
not listednot listednot listed$0.10not listed$0.40MPW only
Qwen3.5-122B-A10B
fetch empty page (HTTP 200)
not listednot listednot listed$0.40not listed$3.20MPW only
Qwen3.5-27B
fetch empty page (HTTP 200)
not listednot listednot listed$0.30not listed$2.40MPW only
Qwen3.5-35B-A3B
fetch empty page (HTTP 200)
not listednot listednot listed$0.25not listed$2.00MPW only
Qwen3-Coder-Next
fetch empty page (HTTP 200)
not listednot listednot listed$0.30not listed$1.50MPW only
Qwen3-32B
fetch empty page (HTTP 200)
not listednot listednot listed$0.16not listed$0.64MPW only
DeepSeek api-docs.deepseek.com/quick_start/pricing OK

Format html

Static HTML (Docusaurus) renders the full price table server-side. No machine-readable twin exists: pricing.md, llms.txt and llms-full.txt fall back to the docs root page, and accept: text/markdown returns the same HTML. One table, two model columns: column 1 = deepseek-flash (MODEL VERSION DeepSeek-V4.1-Flash), column 2 = deepseek-v4-pro (MODEL VERSION DeepSeek-V4-Pro-0813). The MODEL and MODEL VERSION label cells span 3 columns and the PRICING rows use rowspans for the first two cells, so in each PEAK row the first price cell is column 1 and the second is column 2. Each price row has OFF-PEAK and PEAK sub-rows; recipes read the PEAK sub-row, which is what MPW tracks (off-peak is half of peak). Sections are anchored on the MODEL / MODEL VERSION header rows, so a column swap or a column inserted before a model breaks the section match instead of silently swapping models; a column appended after Pro keeps both recipes correct. The Flash anchor needs another cell after DeepSeek-V4.1-Flash, so it also fails if Flash becomes the only column; then drop the trailing space after the final pipe. Cross-check: the zh-cn page (https://api-docs.deepseek.com/zh-cn/quick_start/pricing) has the same table and column order in CNY.

Models on api-docs.deepseek.com/quick_start/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
DeepSeek V4.1 Flash$0.30$0.006$1.20$0.30$0.006$1.20Agree
DeepSeek V4 Pro$1.32$0.044$3.96$1.32$0.044$3.96Agree
Inception api.inceptionlabs.ai/v1/models OK

Format json

Inception's own unauthenticated model list (chat-scoped). Prices are per token as strings; scale 1e6 gives USD per 1M tokens. input_cache_writes is 0 and is not tracked. The list also carries mercury-2.5 (promotional rate, tracked from the docs rate card at list price) and mercury-decide; neither is in this source's hint file.

Models on api.inceptionlabs.ai/v1/models
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Mercury 2$0.25$0.025$0.75$0.25$0.025$0.75Agree
Amazon aws.amazon.com/nova/pricing OK

Format json

The HTML page renders prices client-side: each table cell is a token {priceOf!bedrock/bedrock!<key>!*!1000} that the AWS pricing widget resolves against the official AWS metered-unit-map JSON (data-pricing-endpoint=https://b0.p.awsstatic.com/pricing/2.0/meteredUnitMaps, data-token-paths bedrock/bedrock). We fetch that JSON (gzip, ~2.5 MB) and read regions["US East (N. Virginia)"][<key>].price with raw regexes scoped to that region object (section_end stops at the }}," that closes it, just before "South America (Sao Paulo)"). Prices are USD per 1K tokens, scaled x1000. Each record holds only rateCode, price and RegionlessRateCode (= the key); the first part of rateCode is the AWS Price List SKU, so every key can be checked or found again in the official Price List file https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/AmazonBedrock/current/us-east-1/index.json by the usagetype named in each recipe's notes. Input and output = the Standard Tier table on the 'Geo Cross-region inference and in-region' tab (the 'Global cross-Region inference' tab has no Nova Micro/Lite/Pro rows); the 'Nova Pro (w/ latency optimized inference)' row and the Priority, Flex and Batch tables use other keys and are not read. Cached = the base-model on-demand prompt-cache-read SKU (usagetype USE1-Nova<Model>-cache-read-input-token-count, feature On-demand Inference). The page shows no base cache-read column for these three models; its only cache-read column is in the custom-model 'On Demand Inference' table, whose cache-read keys are the -custom-model SKUs (feature Model Customization, us-east-1 only), so those keys are not used. The page says on-demand inference for custom Nova models is priced the same as base, and both SKUs show the same price today. MPW has no cached value. If AWS re-keys a SKU the recipe fails with 'pattern matched 0 time(s)': re-read input/output keys from data-pricing-markup on page_url, and the cache-read keys from the Price List usagetype. Raw mode is used because the JSON path parser splits on the dot in "N. Virginia". Patterns use character classes ([(], [.], [{], [^}], [0-9.]) instead of backslash escapes; [^}]* keeps each match inside one flat record and does not depend on attribute order. Oregon and Ohio show the same prices for all nine SKUs.

Models on aws.amazon.com/nova/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Nova Micro$0.035$0.00875$0.14$0.035not listed$0.14Agree
Nova Lite$0.06$0.015$0.24$0.06not listed$0.24Agree
Nova Pro$0.80$0.20$3.20$0.80not listed$3.20Agree
Tencent cloud.tencent.com/document/product/1823/130055 OK

Format html, CNY converted at 0.1471 USD

Surface: the official HTML page (server-rendered; the price tables are in the static HTML). The language-model table under "按 Token 计费(后付费)" has region tabs 广州 (Guangzhou, domestic list) then 新加坡 (Singapore). Recipes read only the first (Guangzhou) table. The section requires 广州 to be the first tab label after "按 Token 计费(后付费)", allows up to four more tab labels before the first "模型名称" header (so a new region tab, such as the Silicon Valley tab on the international page, does not break it), and ends at the next "模型名称" header line, which starts the second tab's table. Prices are CNY per 1M tokens (元/百万 tokens), so no scale. usd_rate 0.1471 = 1/6.80 CNY per USD, close to the rates MPW used for Hy3 (6.8031) and Hy-MT2 (6.78); the ECB reference rate in late Aug 2026 was about 6.72, within 1.2%. Hy4 preview is an accepted mismatch: MPW's figure is Tencent's USD list price on the international rate card (https://www.tencentcloud.com/document/product/1300/78937), which is the same ¥6 / ¥18 / ¥0.3 at about 7.197 CNY/USD, and one usd_rate cannot serve both. That international card is a different rate card, not a twin of this page, so it is not the fetch_url: it lists Hy3 at $0.132 / $0.528 / $0.033, while MPW's Hy3 is this CNY card at about 6.80, so switching would only move the mismatch to Hy3. Tier names map to model sizes per Tencent's 混元调用指南 (https://cloud.tencent.com/document/product/1823/132252): hy-mt2-pro = 30B-A3B, hy-mt2-plus = 7B, hy-mt2-lite = 1.8B.

Models on cloud.tencent.com/document/product/1823/130055
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Hy4 preview$0.8826$0.04413$2.65$0.834$0.042$2.50Mismatch
Hunyuan Hy3$0.1471$0.03678$0.5884$0.147$0.037$0.588Agree
Hy-MT2 30B-A3B$0.07355not listed$0.2942$0.074not listed$0.295Agree
Hy-MT2 1.8B$0.04413not listed$0.1765$0.044not listed$0.177Agree

Differences from Model Price Watch

  • Hy4 preview: Input page 0.8826 vs MPW 0.834; Output page 2.6478 vs MPW 2.501; Cached input page 0.04413 vs MPW 0.042
    Same SKU and tier; only the currency surface differs. This page's Guangzhou table row is "Hy4 preview | - | - | 6 | 18 | 0.3" under 推理输入 / 推理输出 / 缓存命中 (元/百万 tokens); the 新加坡 tab row is identical, and MPW's price_note quotes the same ¥6 / ¥18 / ¥0.3. MPW's $0.834 / $2.501 / $0.042 is Tencent's own USD list price on its international rate card (https://www.tencentcloud.com/document/product/1300/78937, Singapore and Guangzhou tabs, same values on 2026-10-08), which is ¥6 / ¥18 / ¥0.3 at about 7.197 CNY/USD. This source converts every row at one usd_rate (1/6.80) because Hy3 and Hy-MT2 need it, so Hy4 reads about $0.883 / $2.648 / $0.044, about 5.8% above MPW. No single rate passes all four models within 3%: Hy4 needs at least 6.988 CNY/USD, while Hy3 and Hy-MT2-Pro need at most 6.966.
Cohere cohere.com/pricing OK

Format html

Visible HTML text. The FAQ legacy block is server-rendered. Tab content (Generative / Advanced retrieval models) is in the Next.js flight JSON in the raw body, not in the visible text.

Models on cohere.com/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Command R+ 08-2024$2.50not listed$10.00$2.50not listed$10.00Agree

Untracked models

  • Embed 4: cohere.com/pricing (checked 2026-10-08) no longer lists Embed 4 anywhere: not in the visible text, not in the Model Vault instance table (Embed 5 Fast/Pro, Rerank 4 Fast/Pro, Rerank 3.5, Parse 5), and not in the embedded Next.js flight data, whose Advanced retrieval models group holds only Embed 5 Pro (Cost 0.12 per 1M tokens), Embed 5 Fast (0.08), Rerank 4 Fast/Pro (per 1K searches) and Parse 5 (per 1K pages), with no image cost entry. The $0.12 that MPW shows now belongs to Embed 5 Pro, a different SKU: docs.cohere.com/docs/cohere-embed lists embed-v5.0-pro and embed-v4.0 as separate models and prints no price for either. Other official surfaces checked with no Embed 4 price: cohere.com/pricing.md (404), accept: text/markdown (same HTML), cohere.com/llms.txt (no prices), cohere.com/{ja,ko,fr,de}/pricing (Embed 5 only), docs.cohere.com/docs/cohere-embed as HTML and docs.cohere.com/docs/how-does-cohere-pricing-work (no prices).
Groq console.groq.com/docs/models OK

Format markdown

Fetches the official markdown twin of the models page (console.groq.com/docs/models.md, text/markdown). Its Production Models rows match the HTML page_url (checked 2026-10-08). Prices are in the 'PRICE PER 1M TOKENS' column as '$X input$Y output'. Rows are anchored on the API model id cell (after the closing ')' of the docs link) so openai/gpt-oss-safeguard-20b cannot match gpt-oss-20b. Tier check: Groq's public model API https://api.groq.com/public/v1/models/dev (no key; the console docs JS loads it) gives per-token prices under metadata.model_price.<tier>. For openai/gpt-oss-120b it has on_demand 0.15/0.60, batch 0.075/0.30 and performance 0.60/2.40 per 1M, so this table shows the on_demand tier. That API is a fallback JSON surface if the .md layout changes (select data[id=openai/gpt-oss-120b], fields metadata.model_price.on_demand.PriceInTokens / PriceOutTokens / PriceInCachedTokens, scale 1000000). The page shows no cached-input price; MPW has none either. As of 2026-10-08 Llama 3.1 8B and Llama 3.3 70B show 'ContactSales' (Enterprise only), and Qwen3.6-27B and Llama 4 Scout are gone from the page (see console.groq.com/docs/deprecations.md). None of the four is in the public API, and their docs/model/<id>.md pages have no PRICING block. If Groq lists public prices for them again, add recipes. The page also lists qwen/qwen3.8-27b ($0.80/$4.00), which has no MPW entry.

Models on console.groq.com/docs/models
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
GPT-OSS 120B$0.15not listed$0.60$0.15not listed$0.60Agree
GPT-OSS 20B$0.075not listed$0.30$0.075not listed$0.30Agree

Untracked models

  • Llama 3.1 8B: Page shows no public price: Production Models row 'Llama 3.1 8B Enterprise llama-3.1-8b-instant | 560 | ContactSales | ContactSales'. Groq deprecated llama-3.1-8b-instant for free and developer tiers on 08/16/26 (docs/deprecations.md); it is now Enterprise-only with contract pricing, so MPW's $0.05/$0.08 is no longer published here. Also checked: docs/model/llama-3.1-8b-instant.md shows an 'Enterprise' badge and no PRICING block, and the public API https://api.groq.com/public/v1/models/dev and /free does not list the model.
  • Llama 3.3 70B: Page shows no public price: Production Models row 'Llama 3.3 70B Enterprise llama-3.3-70b-versatile | 280 | ContactSales | ContactSales'. Groq deprecated llama-3.3-70b-versatile for free and developer tiers on 08/16/26 (docs/deprecations.md); it is now Enterprise-only with contract pricing, so MPW's $0.59/$0.79 is no longer published here. Also checked: docs/model/llama-3.3-70b-versatile.md shows an 'Enterprise' badge and no PRICING block, and the public API https://api.groq.com/public/v1/models/dev and /free does not list the model.
  • Qwen3.6-27B: qwen/qwen3.6-27b is no longer listed on the models page. docs/deprecations.md: deprecated 09/14/26 in favor of qwen/qwen3.8-27b. The page now lists only 'Qwen/Qwen3.8-27B qwen/qwen3.8-27b | 450 | $0.80 input$4.00 output', which is a different SKU, so it must not be used for Qwen3.6-27B. Also checked: docs/model/qwen/qwen3.6-27b.md has no PRICING block ('Loading model information...'), and the public API https://api.groq.com/public/v1/models/dev and /free lists only qwen/qwen3.8-27b.
  • Llama 4 Scout: meta-llama/llama-4-scout-17b-16e-instruct is no longer listed on the models page (no row in Production or Preview tables). docs/deprecations.md: deprecated 07/17/26 for free and developer tiers. No official public price remains. Also checked: docs/model/meta-llama/llama-4-scout-17b-16e-instruct.md has no PRICING block ('Loading model information...'), and the public API https://api.groq.com/public/v1/models/dev and /free does not list the model.
Sakana AI console.sakana.ai/pricing OK

Format html

Server-rendered HTML; prices are in real tables (cells joined by ' | '). The page also lists Fugu Ultra and Fugu Max tables, but MPW cites those from sakana.ai/fugu/ (source sakana-ai-fugu), so only Namazu is tracked here. The Namazu recipe is limited to the sakana-namazu-v* token table, which ends at 'Tool usage', so the web-search and code-execution fees cannot match; without the section the pattern would read the Fugu Ultra table first. No machine-readable twin exists (checked 2026-10-08): console.sakana.ai/llms.txt, /llms-full.txt and /pricing.md return 404, 'accept: text/markdown' still returns HTML, and api.sakana.ai/v1/models needs an API key (401). The RSC payload in the raw HTML repeats the same table, so the visible HTML table is the surface. Sakana's product page sakana.ai/namazu/ shows the same Namazu figures (Input $0.95, Output $4.00, Cached input $0.15 per 1M tokens).

Models on console.sakana.ai/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Sakana Namazu$0.95$0.15$4.00$0.95$0.15$4.00Agree
DeepInfra deepinfra.com/pricing OK

Format html

Surface: the server-rendered HTML of deepinfra.com/pricing. Each family has a 'Model | Context | $ per 1M input tokens | $ per 1M output tokens' table, rendered twice in the page; recipes take the first match. Rows normalize to 'Model-Name' on one line, then '| ctx | $in / $cached cached | $out |' (or '| ctx | $in | $out |' when there is no cache price). Recipes anchor the full model name as a whole line, so dated snapshots (-0731), -Flash vs -Pro, -Vision-Exp and Guard variants cannot match. The table shows the standard (1x base) tier; Priority (1.5x) and Flex (0.8x) are only multipliers in the Service Tiers table. 'cached' is the cache-read price (cents_per_input_token x rate_per_input_token_cached); cache writes are not in the table. Promotions: the page script (pricing chunk, components dD/nd) adds an 'N% off' badge to the name cell and a struck-through list price before the promo price in each price cell, so a promotion makes these patterns fail instead of silently reading a promo price. The official JSON API https://api.deepinfra.com/models/list and the page's own __NEXT_DATA__ (_pricingPageData) carry the same prices and were used to cross-check; neither is used as the surface because they give cached only as a ratio (rate_per_input_token_cached), and the engine cannot multiply two fields. Checked 2026-10-08: /pricing.md and deepinfra.com/llms.txt return 404, and 'accept: text/markdown' returns the same HTML.

Models on deepinfra.com/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
DeepSeek V4 Flash$0.06$0.015$0.18$0.06$0.015$0.18Agree
DeepSeek V4 Pro$1.30$0.10$2.60$1.30$0.10$2.60Agree
Qwen3-32B$0.08not listed$0.28$0.08not listed$0.28Agree
Llama 4 Scout$0.10not listed$0.30$0.10not listed$0.30Agree
NVIDIA Nemotron 3.5 Lightning$0.06$0.03$0.16$0.08$0.04$0.20Mismatch

Differences from Model Price Watch

  • NVIDIA Nemotron 3.5 Lightning: Input page 0.06 vs MPW 0.08; Output page 0.16 vs MPW 0.2; Cached input page 0.03 vs MPW 0.04
    Pricing page Nemotron table row (2026-10-08): 'NVIDIA-Nemotron-3.5-Lightning | 256k | $0.06 / $0.03 cached | $0.16 |', linked to /nvidia/NVIDIA-Nemotron-3.5-Lightning. It is the only Nemotron 3.5 Lightning text model on the page and in the API (page siblings: Nemotron-Content-Safety-3.5, NVIDIA-Nemotron-3-Ultra-550B-A55B, NVIDIA-Nemotron-3-Super-120B-A12B). It has only the standard tier (rate_per_service_tier_priority and _flex are null) and no promotion: the page's __NEXT_DATA__, the model page and https://api.deepinfra.com/models/list all give discount null, cents_per_input_token 0.000006, cents_per_output_token 0.000016, rate_per_input_token_cached 0.5 = $0.06 / $0.03 cached / $0.16. MPW's $0.08 / $0.04 / $0.20 is DeepInfra's earlier list price: Wayback snapshots of DeepInfra's own model page (2026-08-22, 2026-09-12 and 2026-09-23) show cents_per_input_token 0.000008, cents_per_output_token 0.00002, rate_per_input_token_cached 0.5, discount null. DeepInfra cut the price after 2026-09-23; MPW should update to $0.06 / $0.03 / $0.16.

Untracked models

  • Llama 4 Maverick: Llama 4 Maverick is no longer on deepinfra.com/pricing: the Llama 4 table lists only Llama-4-Scout-17B-16E and Llama-Guard-4-12B, and neither the visible HTML nor the embedded __NEXT_DATA__ has a 'Maverick' string. Other official surfaces tried on 2026-10-08: /pricing.md and deepinfra.com/llms.txt return 404; 'accept: text/markdown' returns the same HTML; docs.deepinfra.com/llms.txt and llms-full.txt have no model prices; the model pages /meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 and -Turbo return 404. The official API https://api.deepinfra.com/models/list still lists meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 at $0.20 / $0.80 (cents_per_input_token 0.00002, cents_per_output_token 0.00008; matches MPW), but marks it deprecated (1790897569 = 2026-10-01T23:32:49Z) with replaced_by google/gemma-4-31B-it-turbo. DeepInfra's docs (Model deprecation) say requests to a deprecated model are forwarded to the replacement, so $0.20 / $0.80 is no longer a live price. MPW gives no anchor_sku. Retire or repoint the MPW entry.
IBM developers.cloudflare.com/workers-ai/models/granite-4.0-h-micro OK

Format markdown

Fetches the official Markdown twin of the Cloudflare Workers AI model page (index.md, linked from the page as 'View as Markdown'). Cloudflare hosts the model, so its docs are the official price surface for this SKU. The 'Unit Pricing' row of the Model Info table is the only price on the page.

Models on developers.cloudflare.com/workers-ai/models/granite-4.0-h-micro
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Granite 4.0 H Micro$0.017not listed$0.112$0.017not listed$0.112Agree
OpenAI developers.openai.com/api/docs/changelog OK

Format markdown

Fetches the official Markdown twin (changelog.md, text/markdown), which returns Markdown for any Accept header. The bare page URL returns the same Markdown only when the Accept header lists text/markdown (the tracker's default header does); with Accept */* or a browser's text/html it returns the HTML page, which shows the same Sep 22 entry. GPT-6 Sol prices come from the 2026-09-22 release entry, which states standard pricing per 1M tokens for prompts with up to 272K input tokens. A changelog entry is historical: it does not change when the rate card changes, so a later price change for gpt-6-sol will not show here. Live official surfaces to check for a change: the HTML rate card no longer shows a gpt-6-sol row, but on 2026-10-08 the rate card's Markdown twin (https://developers.openai.com/api/docs/pricing.md, '### Standard pricing data', short-context columns) and the model page (https://developers.openai.com/api/docs/models/gpt-6-sol.md, '### Text tokens') still listed gpt-6-sol with the same input, cached input and output values as this entry.

Models on developers.openai.com/api/docs/changelog
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
GPT-6 Sol$2.00$0.20$10.00$2.00$0.20$10.00Agree
OpenAI developers.openai.com/api/docs/models/gpt-5.3-codex OK

Format markdown

Fetches the official Markdown twin (gpt-5.3-codex.md, text/markdown), which returns Markdown for any Accept header. The bare page URL returns the same bytes only when the Accept header lists text/markdown (the tracker's default header does); with Accept */* or a browser's text/html it returns the HTML page, which shows the same Text tokens values. Prices come from the first 'Text tokens' table under '## Pricing'. That table is the Standard tier: on 2026-10-08 the rate card (pricing.md, Standard 'Grouped Pricing Table data', Codex row) had the same three values, and its Fast table had a higher row. The Markdown twin drops the HTML tier tabs: on pages with several tiers (for example o3.md) each tier is an unlabeled '### Text tokens' table, Standard first. The 'Quick comparison' table repeats the row and also lists GPT-5.2-Codex; the recipe does not read it. The model is deprecated: the HTML page shows a Deprecated badge, and the deprecations page lists shutdown on 2027-04-01 with gpt-6-sol as the replacement. Retired model pages such as gpt-4.5-preview.md keep their Text tokens table, so after the shutdown this source will probably keep reporting the last listed price.

Models on developers.openai.com/api/docs/models/gpt-5.3-codex
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
GPT-5.3 Codex$1.75$0.175$14.00$1.75$0.175$14.00Agree
OpenAI developers.openai.com/api/docs/pricing.md OK

Format markdown

Fetches the page itself, the official markdown twin https://developers.openai.com/api/docs/pricing.md (text/markdown, 200, no auth). The same fetch_url is shared by platform-openai-com-docs-pricing, developers-openai-com-api-docs-pricing and developers-openai-com-api-docs-pricing-md so the tracker fetches it once. Text-model recipes read the Flagship models "### Standard pricing data" table, short-context columns; realtime and audio recipes read the "Realtime and audio generation models" table, Text modality rows. Each section pins its heading and its column header ("Model | Short context input | Short context cached input | Short context cache writes | Short context output" and "Model | Modality | Input | Cached input | Output / cost"), so a column reorder or insertion fails loudly instead of shifting fields. section_end "^[^|\n]" ends the section at the first line that is not a table row, so each scope is exactly one table: Batch, Flex, Fast and Ultrafast stay out. Anchors end the model name at a " | " cell boundary so -mini, -nano, -pro, dated snapshots and the 2.1 realtime models cannot match their siblings.

Models on developers.openai.com/api/docs/pricing.md
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
GPT-Realtime-2$4.00$0.40$24.00$4.00$0.40$24.00Agree
GPT-Realtime-2.1$4.00$0.40$24.00$4.00$0.40$24.00Agree
GPT-Realtime-2.1 mini$0.60$0.06$2.40$0.60$0.06$2.40Agree
GPT-Audio-1.5$2.50not listed$10.00$2.50not listed$10.00Agree
GPT-4o$2.50$1.25$10.00$2.50$1.25$10.00Agree
GPT-4o mini$0.15$0.075$0.60$0.15$0.075$0.60Agree
GPT-5.1$1.25$0.125$10.00$1.25$0.125$10.00Agree
GPT-5$1.25$0.125$10.00$1.25$0.125$10.00Agree
GPT-5 mini$0.25$0.025$2.00$0.25$0.025$2.00Agree
GPT-5 nano$0.05$0.005$0.40$0.05$0.005$0.40Agree
GPT-5 Pro$15.00not listed$120.00$15.00not listed$120.00Agree
o3$2.00$0.50$8.00$2.00$0.50$8.00Agree
o3-pro$20.00not listed$80.00$20.00not listed$80.00Agree
OpenAI developers.openai.com/api/docs/pricing OK

Format markdown

Fetches the official markdown twin https://developers.openai.com/api/docs/pricing.md (text/markdown, 200, no auth). The HTML page_url shows only the first 3 rows of each table (the rest sit behind an "All models" expander), while the .md twin holds every row and equals the page's own text/markdown response. The same fetch_url is shared by platform-openai-com-docs-pricing, developers-openai-com-api-docs-pricing and developers-openai-com-api-docs-pricing-md so the tracker fetches it once. Recipes read the Flagship models "### Standard pricing data" table, short-context columns. The section pins that heading and the column header "Model | Short context input | Short context cached input | Short context cache writes | Short context output", so a column reorder or insertion fails loudly instead of shifting fields (these rows have a numeric cache-writes cell that a shifted pattern would read as output). section_end "^[^|\n]" ends the section at the first line that is not a table row, so Batch, Flex, Fast and Ultrafast stay out. Anchors end the model name at a " | " cell boundary, so gpt-6-astra and gpt-6.1-sol cannot match gpt-6-sol.

Models on developers.openai.com/api/docs/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
GPT-6 Astra$10.00$1.00$50.00$10.00$1.00$50.00Agree
GPT-6.1 Sol$2.00$0.10$10.00$2.00$0.10$10.00Agree
GPT-6 Luna$0.10$0.01$0.50$0.10$0.01$0.50Agree
Cohere docs.cohere.com/docs/command-a OK

Format html

Fetch the HTML model card with accept: text/html. Fern serves the .md twin (finalUrl command-a.md) whenever Accept lists text/markdown or text/plain, at any q-value (checked 2026-10-08), and the engine default Accept lists both. The .md twin omits the ModelShowcase card (Pricing, Specifications, API Endpoints). The raw HTML also embeds the page MDX in the Copy page markdown prop, where <ModelShowcase> holds id: 'command-a-03-2025' and pricing: { input: 2.50, output: 10.0 }. That agrees with the card, but it is HTML- and JSON-escaped and states no unit, so the recipe reads the visible card, which prints '/ 1M tokens'. Do not add text/markdown or text/plain to the header.

Models on docs.cohere.com/docs/command-a
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Command A$2.50not listed$10.00$2.50not listed$10.00Agree
Cohere docs.cohere.com/docs/command-r OK

Format html

Fetch the HTML model card with accept: text/html. Fern serves the .md twin (finalUrl command-r.md) whenever Accept lists text/markdown or text/plain, at any q-value (checked 2026-10-08), and the engine default Accept lists both. The .md twin omits the ModelShowcase card (Pricing, Specifications, API Endpoints). The raw HTML also embeds the page MDX in the Copy page markdown prop, where <ModelShowcase> holds id: 'command-r-08-2024' and pricing: { input: 0.15, output: 0.60 }. That agrees with the card, but it is HTML- and JSON-escaped and states no unit, so the recipe reads the visible card, which prints '/ 1M tokens'. Do not add text/markdown or text/plain to the header.

Models on docs.cohere.com/docs/command-r
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Command R 08-2024$0.15not listed$0.60$0.15not listed$0.60Agree
Fireworks docs.fireworks.ai/serverless/pricing OK

Format markdown

Fetches the official Mintlify markdown twin (pricing.md). Recipes read the 'Text and vision models' table. The section regex requires the intro sentence that states the cell order 'input / cached input / output' directly above a header whose first price column is 'Standard', so a reordered cell or column layout fails loudly instead of being misread. The Standard cell (first cell after the name) is used; Priority is ignored. The closing ']' of each link text ends the name, so Flash, Fast and (US) rows cannot match. Ten hint models are not on the current rate card: the official changelog (docs.fireworks.ai/updates/changelog.md, entries 2026-08-27 and 2026-09-26) removed eight of them from serverless, and the 'Models by tier' table on docs.fireworks.ai/serverless/rate-limits.md, which lists every serverless model, omits all ten. If Fireworks re-lists any of them, add a recipe.

Models on docs.fireworks.ai/serverless/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
GLM-5.3$1.40$0.26$4.40$1.40$0.26$4.40Agree
GPT-OSS 120B$0.15$0.015$0.60$0.15$0.015$0.60Agree
MiniMax-M3$0.30$0.06$1.20$0.30$0.06$1.20Agree
NVIDIA Nemotron 3 Ultra$0.60$0.12$2.40$0.60$0.12$2.40Agree

Untracked models

  • DeepSeek V4 Flash: Removed from Fireworks serverless on 2026-09-25. Changelog 2026-09-26 ('Serverless deprecation: DeepSeek V4 Pro (0813), DeepSeek V4 Flash (0731), and related models') says they are no longer on public serverless; migrate to DeepSeek V4.1 Flash. The rate card lists only DeepSeek V4.1 Flash, a different SKU; the rate-limits 'Models by tier' list omits V4 Flash; fireworks.ai/models/fireworks/deepseek-v4-flash-0731 returns 404.
  • DeepSeek V4 Pro: Removed from Fireworks serverless on 2026-09-25 (changelog 2026-09-26: DeepSeek V4 Pro (0813), migrate to DeepSeek V4.1 Flash; the undated DeepSeek V4 Pro left serverless on 2026-08-27). No DeepSeek V4 Pro row on the rate card or in the rate-limits 'Models by tier' list; fireworks.ai/models/fireworks/deepseek-v4-pro-0813 returns 404 and .../deepseek-v4-pro shows 'Serverless: Not supported'.
  • GLM-5.1: Not on the serverless rate card or in the rate-limits 'Models by tier' list (MPW also had no anchor on this page). The official model card fireworks.ai/models/fireworks/glm-5p1 shows 'Serverless: Not supported' and offers only Fine-tuning and On-demand Deployment, with no per-token price.
  • GLM-5.2: Removed from Fireworks serverless on 2026-09-25 (changelog 2026-09-26: GLM 5.2, including GLM 5.2 Fast, Fast US and US, migrate to GLM 5.3). No GLM 5.2 row on the rate card or in the rate-limits 'Models by tier' list. The model card fireworks.ai/models/fireworks/glm-5p2 still prints a stale 'Available Serverless' block with MPW's old figures, but its own FAQ says the serverless deprecation started on September 25, 2026, so that block is not a live serverless price.
  • Muse Glimmer 30B: Removed from Fireworks serverless on 2026-09-25 (changelog 2026-09-26: migrate to NVIDIA Nemotron 3.5 Lightning 30B A3B). Not on the rate card or in the rate-limits 'Models by tier' list; the model card shows 'Serverless: Not supported'. fireworks.ai/pricing lists it only under the Serverless Training API (training prices, not inference).
  • GPT-OSS 20B: Deprecated from Fireworks serverless effective 2026-08-27 (changelog 2026-08-27 'Serverless deprecation: MiniMax M2.7, GPT OSS 20B, ...'; migrate to GPT OSS 120B or Qwen3 8B). Not on the rate card or in the rate-limits 'Models by tier' list; the model card shows 'Serverless: Not supported'.
  • Kimi K2.6: Removed from Fireworks serverless on 2026-09-25 (changelog 2026-09-26: Kimi K2.6, migrate to GLM 5.3 or Kimi K3). Not on the rate card or in the rate-limits 'Models by tier' list; the model card shows 'Serverless: Not supported'.
  • Kimi K2.7 Code: Removed from Fireworks serverless on 2026-09-25 (changelog 2026-09-26: Kimi K2.7 Code, migrate to GLM 5.3 or Kimi K3). Not on the rate card or in the rate-limits 'Models by tier' list; the model card shows 'Serverless: Not supported'.
  • MiniMax-M2.7: Deprecated from Fireworks serverless effective 2026-08-27 (changelog 2026-08-27: MiniMax M2.7, migrate to MiniMax M3). Not on the rate card or in the rate-limits 'Models by tier' list; the model card shows 'Serverless: Not supported'.
  • Qwen3.7-Plus: Not on the serverless rate card (the Qwen row is now Qwen 3.8 Max, a different SKU) or in the rate-limits 'Models by tier' list. The official model card fireworks.ai/models/fireworks/qwen3p7-plus shows 'Serverless: Not supported' and no per-token price.
Inception docs.inceptionlabs.ai/get-started/models OK

Format markdown

Fetches the Mintlify markdown twin (page_url + .md), published by Inception's docs site. The HTML page shows the same rate card (<del>$0.20</del> → <strong>$0.04</strong>), but normalized HTML text drops the <del> mark, while the markdown keeps the list price as ~~strikethrough~~ ahead of the promotional price. `section` is the rate card's header row, so the recipes run only while the columns are Model | Input | Cached Input | Output, each per 1M tokens; `section_end` stops at the first line after the table. A price cell must be '~~list~~ → **promo**' (the struck list price is read, as price_note asks) or one plain or bold price. Any other shape (a second price in the cell, promo first, an inserted column, a unit change) fails loudly instead of reading a wrong number. The table also lists Mercury Voice, Mercury Decide and Mercury 2, which are not in this source's hint file.

Models on docs.inceptionlabs.ai/get-started/models
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Mercury 2.5$0.20$0.02$0.75$0.20$0.02$0.75Agree
Mercury Edit 2$0.25$0.025$0.75$0.25$0.025$0.75Agree
Perplexity docs.perplexity.ai/docs/agent-api/models OK

Format markdown

Fetches the official Mintlify markdown twin (page URL + .md, text/markdown, ~140KB) instead of the ~1MB HTML; the HTML page_url shows the same perplexity/sonar row under the same header. The Available Models section has one <Tab> per provider family, each with one table: Model | Input ($/1M) | Output ($/1M) | Cache read ($/1M) | Service tiers | Docs. Listed prices are default processing; flex is 0.5x and priority 2x of the listed prices (Service tiers section), not tracked. The .md also embeds a PRICING JS object for the calculator widget; it holds the same agent rate and, under a separate sonar key, the different Sonar API product ($1/$1). Recipes read the human table, not that object.

Models on docs.perplexity.ai/docs/agent-api/models
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Sonar (Agent API)$0.25$0.0625$2.50$0.25$0.0625$2.50Agree
Reka docs.reka.ai/pricing OK

Format html

docs.reka.ai/pricing (and its .md twin) lists no prices. It says 'The model catalog on developer.reka.ai lists the current input, output and cached-input price for every model' and points to GET /v1/models. api.reka.ai/v1/models and inference.api.reka.ai/v1/models return 401 without an API key, and developer.reka.ai has no llms.txt, .md or JSON twin (404). developer.reka.ai/models is a Next.js page whose model list (with per-Mtok 'rates') is embedded only in the RSC flight data (self.__next_f.push), so recipes use mode raw. The visible text of /models has no prices, so the daily page hash does not change when only a price changes; the recipe still reads the rates. The catalog lists reka-edge-2603 and reka-flash-3 as the only Reka models; legacy Reka Core and Reka Flash (v1/v2) are absent. Fallback surfaces if the flight embedding changes: the same records come as plain JSON with the request header 'RSC: 1' (content type text/x-component; the reka-edge pattern already reads that body), and the developer.reka.ai home page has a server-rendered table 'Text model prices per million tokens, from /v1/models' (Model | Context | Input $/M | Output $/M | Cached input $/M).

Models on docs.reka.ai/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Reka Edge$0.10not listed$0.10$0.10not listed$0.10Agree

Untracked models

  • Reka Core: Reka Core has no price on any official Reka surface reachable without login. docs.reka.ai/pricing and its .md twin list no prices. The developer.reka.ai/models catalog, the developer.reka.ai home price table and docs.reka.ai/chat/models list no Reka Core model, and none of the docs .md pages indexed in docs.reka.ai/llms.txt names it. The reka.ai JSON-LD names Reka Core as a product with no price, and the Reka Core launch post links 'Reka Model Pricing' to docs.reka.ai/pricing. api.reka.ai/v1/models and inference.api.reka.ai/v1/models need an API key (401). MPW's $2/$6 cannot be verified.
  • Reka Flash: MPW's $0.80/$2.00 is the legacy Reka Flash SKU, which no official surface lists now. The only Flash SKU on developer.reka.ai/models, on the developer.reka.ai home price table and in docs.reka.ai/chat/models is reka-flash-3 ('Reka Flash 3', a 21B text-only reasoning model with 64k context, $0.10 in / $0.20 out per Mtok). That is a different model generation: MPW's seed metadata for Reka Flash (128k context, image, video and audio input) describes the legacy multimodal Flash, so it is not the same SKU. The only other official 'Reka Flash' price is the Vision API quick-tagging endpoint (per input video minute plus $2 per 1M output tokens, the same rate as every Vision QA endpoint), which is not the chat SKU. Map reka-flash-3 here only if MPW confirms it means Flash 3.
TypeSafe AI docs.typesafe.ai/models OK

Format markdown

Fetches the official Mintlify markdown twin of the page (models.md, text/markdown). The Jev table prints one price row 'Price (per Btok / per Mtok) | $42 / $0.042'; both numbers are the input rate in two units, and the recipe reads the per-Mtok number. Output is captured from the note under the table, 'Charged per input token. Output tokens are free.' (parseNum maps 'free' to 0). The markdown escapes dollar signs with a backslash.

Models on docs.typesafe.ai/models
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Jev 1.13$0.042not listed$0.00$0.042not listed$0.00Agree
Voyage AI docs.voyageai.com/docs/pricing OK

Format markdown

Fetches the official markdown twin (pricing.md). The HTML page says "Append .md to any documentation page URL to get its markdown version" and points to https://docs.voyageai.com/llms.txt. The twin and the HTML page showed the same prices on 2026-10-08. The HTML flattens grouped cells into separate lines; the markdown keeps each group in one cell joined by <br />. Each section regex starts at the table heading and requires the first table under that heading to have the expected column headers, so a column insert, drop or reorder fails loudly instead of shifting the captured price to another column (for example a batch column). section_end "^[^|]" ends the scope at the first line after the header that is not a table row, so each recipe reads only that one table. Patterns accept a model name anywhere inside a grouped cell and end the name at its closing backtick, so rerank-2.5 cannot match rerank-2.5-lite and voyage-4 cannot match voyage-4-large. voyage-code-3, voyage-context-3, rerank-2.5 and rerank-2.5-lite are read from the Older models table, because the current tables now list voyage-code-4, voyage-context-4, rerank-3 and rerank-3-lite. Standard rates only; the Batch API (33% discount) is not used. All models bill input only, so output is null.

Models on docs.voyageai.com/docs/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
voyage-4-large$0.12not listednot listed$0.12not listednot listedAgree
voyage-4$0.06not listednot listed$0.06not listednot listedAgree
voyage-4-lite$0.02not listednot listed$0.02not listednot listedAgree
voyage-code-3$0.18not listednot listed$0.18not listednot listedAgree
voyage-context-3$0.18not listednot listed$0.18not listednot listedAgree
voyage-multimodal-3.5$0.12not listednot listed$0.12not listednot listedAgree
rerank-2.5$0.05not listednot listed$0.05not listednot listedAgree
rerank-2.5-lite$0.02not listednot listed$0.02not listednot listedAgree
xAI docs.x.ai/developers/models/grok-4.20 OK

Format markdown

Surface: page_url with an explicit 'Accept: text/markdown' header, the method docs.x.ai documents in https://docs.x.ai/llms.txt. With Accept text/html or */* the same URL returns a ~400 KB Next.js HTML page; without this header the fetch only got markdown because the engine's default Accept lists text/markdown;q=0.9 and the server ignores q-values. Do not switch to the .md twin: it returns 404 whenever Accept contains text/markdown, which the engine default does. Recipe reads the first price column ('< 200k prompt tokens (per 1M tokens)', standard tier) of the Pricing table on the Grok 4.20 card (model name grok-4.20-0309-reasoning); the ≥ 200k column is long-context pricing and Batch API is a 20% discount, both ignored. Cross-checked 2026-10-08: the HTML card shows the same $1.25 / $0.20 / $2.50 per 1M tokens, and https://docs.x.ai/developers/pricing shows the same in the grok-4.20-0309-reasoning (< 200k) row.

Models on docs.x.ai/developers/models/grok-4.20
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Grok 4.20$1.25$0.20$2.50$1.25$0.20$2.50Agree
xAI docs.x.ai/developers/models/grok-4.3 OK

Format markdown

Surface: page_url with an explicit 'Accept: text/markdown' header, the method docs.x.ai documents in https://docs.x.ai/llms.txt. With Accept text/html or */* the same URL returns a ~380 KB Next.js HTML page; without this header the fetch only got markdown because the engine's default Accept lists text/markdown;q=0.9 and the server ignores q-values. Do not switch to the .md twin: it returns 404 whenever Accept contains text/markdown, which the engine default does. Recipe reads the first price column ('< 200k prompt tokens (per 1M tokens)', standard tier) of the Pricing table on the Grok 4.3 card (model name grok-4.3); the ≥ 200k column is long-context pricing and Batch API is a 20% discount, both ignored. Cross-checked 2026-10-08: the HTML card shows the same $1.25 / $0.20 / $2.50 per 1M tokens, and https://docs.x.ai/developers/pricing shows the same in the grok-4.3 (< 200k) row; the eu-west-1 cluster in the HTML data has the same price.

Models on docs.x.ai/developers/models/grok-4.3
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Grok 4.3$1.25$0.20$2.50$1.25$0.20$2.50Agree
xAI docs.x.ai/developers/models/grok-4.7 OK

Format markdown

Surface: page_url with an explicit 'Accept: text/markdown' header, the method docs.x.ai documents in https://docs.x.ai/llms.txt. With Accept text/html or */* the same URL returns a ~380 KB Next.js HTML page; without this header the fetch only got markdown because the engine's default Accept lists text/markdown;q=0.9 and the server ignores q-values. Do not switch to the .md twin: it returns 404 whenever Accept contains text/markdown, which the engine default does. Recipe reads the first price column ('< 200k prompt tokens (per 1M tokens)', standard tier) of the Pricing table on the Grok 4.7 card (model name grok-4.7); the ≥ 200k column is long-context pricing and is ignored. Not used: the US regional endpoint (1.1x, $2.20 / $0.55 / $6.60; the us-central-1 cluster in the HTML data), Priority (2x) and 'Grok 4.7 Fast' ($4.00 / $1.00 / $12.00, Cursor and Grok Build plans only, not on the public API), all on https://docs.x.ai/developers/pricing. Cross-checked 2026-10-08: the HTML card shows the same $2.00 / $0.50 / $6.00 per 1M tokens.

Models on docs.x.ai/developers/models/grok-4.7
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Grok 4.7$2.00$0.50$6.00$2.00$0.50$6.00Agree
xAI docs.x.ai/developers/models/grok-build-0.1 OK

Format markdown

Surface: page_url with an explicit 'Accept: text/markdown' header, the method docs.x.ai documents in https://docs.x.ai/llms.txt. With Accept text/html or */* the same URL returns a ~380 KB Next.js HTML page; without this header the fetch only got markdown because the engine's default Accept lists text/markdown;q=0.9 and the server ignores q-values. Do not switch to the .md twin: it returns 404 whenever Accept contains text/markdown, which the engine default does. Recipe reads the first price column ('< 200k prompt tokens (per 1M tokens)', standard tier) of the Pricing table on the Grok Build 0.1 card (model name grok-build-0.1; aliases grok-code-fast-1, grok-code-fast, grok-code-fast-1-0825); the ≥ 200k column is long-context pricing and is ignored. This is the API model, not the Grok Build coding product or its 'Grok 4.7 Fast' plan pricing. Cross-checked 2026-10-08: the HTML card shows the same $1.00 / $0.20 / $2.00 per 1M tokens, and https://docs.x.ai/developers/pricing shows the same in the grok-build-0.1 (< 200k) row.

Models on docs.x.ai/developers/models/grok-build-0.1
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Grok Build 0.1$1.00$0.20$2.00$1.00$0.20$2.00Agree
xAI docs.x.ai/developers/models OK

Format markdown

Surface: page_url with an explicit 'Accept: text/markdown' header, the method docs.x.ai documents in https://docs.x.ai/llms.txt. With Accept text/html or */* the same URL returns a ~400 KB Next.js HTML page whose visible text has no Text API Pricing table; without this header the fetch only got markdown because the engine's default Accept lists text/markdown;q=0.9 and the server ignores q-values. Do not switch to the .md twin (https://docs.x.ai/developers/models.md): it returns 404 whenever Accept contains text/markdown, which the engine default does. The section starts at the table header row, so the column order Model | Context | Input / 1M tokens | Cached input / 1M tokens | Output / 1M tokens is checked, not assumed. Row '<sku> (< 200k prompt tokens)' is the standard tier; the '(≥ 200k prompt tokens)' rows are long-context pricing and are ignored. Batch (20% off for some models), Priority (2x) and the US regional endpoint (1.1x) are not in this table. Cross-checked 2026-10-08: https://docs.x.ai/developers/pricing shows the same rows, and the us-east-1/us-west-2 records embedded in the HTML page match. The table also lists grok-4.7, grok-4.3, grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning, grok-build-0.1 and grok-4.20-multi-agent-0309; MPW cites this page only for grok-4.6 and grok-4.5 (grok-4.7, grok-4.3, grok-4.20 and grok-build-0.1 are tracked from their own model pages; the non-reasoning and multi-agent SKUs have no MPW id).

Models on docs.x.ai/developers/models
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Grok 4.6$2.00$0.50$6.00$2.00$0.50$6.00Agree
Grok 4.5$2.00$0.30$6.00$2.00not listed$6.00Agree
Z.AI docs.z.ai/guides/overview/pricing OK

Format markdown

Fetches the official Mintlify Markdown twin (page_url + .md); it carries the same tables and cell values as the HTML page (checked 2026-10-08, no struck-through cells). The z.ai/pricing page's config API (api.z.ai/api/biz/operation/query?ids=1163,1164) shows the same prices for GLM-5.3, GLM-5.3-Flash, GLM-5.2, GLM-4.6V and GLM-OCR. Columns of the three per-token tables: Model | Input | Cached Input | Cached Input Storage | Output. Cached Input Storage ('Limited-time Free') is not a published rate and is skipped. Every recipe searches the block from the first of '### Latest Models', '### Text Models' or '### Vision Models' up to the next heading that is not one of these three, so a row that moves between these tables at a model launch is still read, and other tables (tools, image, video, audio, agents, or a new batch table) are out of scope. Each pattern anchors at the line start on the exact '| <name> |' cell, so sibling rows (-Flash, -FlashX, -Air, -AirX, -X, V) cannot match. Markdown escapes the dollar sign ('\$0.15'); patterns accept '\$', '$' or no sign. Free-tier rows print 'Free', which the engine parses as 0; patterns accept a number or 'Free', so a switch to paid pricing is still read. Mapping is by column position: on paid rows a new column before Output makes the match fail, because the 'Limited-time Free' storage cell cannot parse as a price. GLM-5-Turbo and GLM-5V-Turbo are absent from this page and from every other official surface checked (see untracked).

Models on docs.z.ai/guides/overview/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
GLM-5.3-Flash$0.15$0.03$0.50$0.15$0.03$0.50Agree
GLM-5.3-FlashX$0.37$0.075$1.25$0.37$0.075$1.25Agree
GLM-5.3$1.40$0.26$4.40$1.40$0.26$4.40Agree
GLM-5.2$1.40$0.26$4.40$1.40$0.26$4.40Agree
GLM-5.1$1.40$0.26$4.40$1.40$0.26$4.40Agree
GLM-5$1.00$0.20$3.20$1.00$0.20$3.20Agree
GLM-4.7$0.60$0.11$2.20$0.60$0.11$2.20Agree
GLM-4.7-FlashX$0.07$0.01$0.40$0.07$0.01$0.40Agree
GLM-4.6$0.60$0.11$2.20$0.60$0.11$2.20Agree
GLM-4.5$0.60$0.11$2.20$0.60$0.11$2.20Agree
GLM-4.5-X$2.20$0.45$8.90$2.20$0.45$8.90Agree
GLM-4.5-Air$0.20$0.03$1.10$0.20$0.03$1.10Agree
GLM-4.5-AirX$1.10$0.22$4.50$1.10$0.22$4.50Agree
GLM-4-32B-0414$0.10not listed$0.10$0.10not listed$0.10Agree
GLM-4.7-Flash$0.00$0.00$0.00$0.00$0.00$0.00Agree
GLM-4.5-Flash$0.00$0.00$0.00$0.00$0.00$0.00Agree
GLM-4.6V$0.30$0.05$0.90$0.30$0.05$0.90Agree
GLM-OCR$0.03not listed$0.03$0.03not listed$0.03Agree
GLM-4.6V-FlashX$0.04$0.004$0.40$0.04$0.004$0.40Agree
GLM-4.5V$0.60$0.11$1.80$0.60$0.11$1.80Agree
GLM-4.6V-Flash$0.00$0.00$0.00$0.00$0.00$0.00Agree

Untracked models

  • GLM-5-Turbo: GLM-5-Turbo has no row on docs.z.ai/guides/overview/pricing (HTML page or .md twin) as of 2026-10-08: the Latest, Text and Vision Models tables do not list it. No other official surface checked on 2026-10-08 gives a USD per-token price: its unlisted model page docs.z.ai/guides/llm/glm-5-turbo (and .md) gives only context 200K and max output 128K and links back to the pricing page; docs.z.ai/llms.txt, llms-full.txt, sitemap.xml, openapi.json and release-notes/new-released do not mention it; the z.ai/pricing config API (api.z.ai/api/biz/operation/query?ids=1163,1164,1166) has no GLM-5-Turbo card; api.z.ai/api/paas/v4/models needs an API key. The configs that name it (ids 1119, 1128, 1129) are bigmodel.cn China-platform news, product and carousel lists with no price. A maintainer must find Z.AI's current USD price surface if MPW keeps citing this page.
  • GLM-5V-Turbo: GLM-5V-Turbo has no row on docs.z.ai/guides/overview/pricing (HTML page or .md twin) as of 2026-10-08: the Latest, Text and Vision Models tables do not list it. No other official surface checked on 2026-10-08 gives a USD per-token price: its unlisted model page docs.z.ai/guides/vlm/glm-5v-turbo (and .md) gives only context 200K and max output 128K and links back to the pricing page; docs.z.ai/llms.txt, llms-full.txt, sitemap.xml, openapi.json and release-notes/new-released do not mention it; the z.ai/pricing config API (api.z.ai/api/biz/operation/query?ids=1163,1164,1166) has no GLM-5V-Turbo card; api.z.ai/api/paas/v4/models needs an API key. Configs 1119 and 1128 name it in bigmodel.cn news and product lists with no price. The only priced mention (config id 1162, isOverseas false) is the bigmodel.cn China FAQ in CNY with 32K context tiers (input 5/7, output 22/26 CNY per 1M), a different price list from Z.AI's flat USD card, so it cannot back MPW's 1.2/4.0/0.24 USD. A maintainer must find Z.AI's current USD price surface if MPW keeps citing this page.
Fireworks fireworks.ai/models/fireworks/ember-1 OK

Format html

Server-rendered model card HTML (the page_url itself); the 'Available Serverless' card prints input / cached input / output per 1M tokens. No machine-readable twin exists: ember-1.md and fireworks.ai/llms.txt return 404, 'accept: text/markdown' still returns HTML, and the embedded RSC payload holds only the rendered price strings, not structured price fields. The same Standard price also appears on docs.fireworks.ai/serverless/pricing.md (Ember-1 row) if this page layout changes.

Models on fireworks.ai/models/fireworks/ember-1
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Ember-1$3.00$0.30$15.00$3.00$0.30$15.00Agree
Groq groq.com/pricing OK

Format html

As of 2026-10-08 https://groq.com/pricing returns HTTP 308 to https://groq.com/ (the homepage), which has no price table; groq.com/pricing/ also 308s back to /pricing. The redirect does not depend on user agent or accept header (text/markdown too). It is part of a site redesign: groq.com/groqcloud 308s to /platform, groq.com/sitemap.xml has no pricing page, and neither the homepage nor /platform has prices. The only remaining official price lists are console.groq.com/docs/models (source console-groq-com-docs-models), the docs/model/<id>.md cards and Groq's public model API https://api.groq.com/public/v1/models/dev; none of them lists a price for Qwen3 32B. fetch_url stays on the pricing URL so the daily page hash shows if Groq restores the page; then add a recipe. Until then the hash tracks the homepage text, so homepage edits also show as page changes.

Untracked models

  • Qwen3-32B: groq.com/pricing now 308-redirects to the groq.com homepage, which has no prices. qwen/qwen3-32b is also absent from console.groq.com/docs/models(.md); docs/deprecations.md says it was deprecated 07/17/26 for free and developer tiers in favor of openai/gpt-oss-120b. No official surface publishes a price for 'Qwen3 32B 131k' any more. Also checked 2026-10-08: docs/model/qwen/qwen3-32b.md has no PRICING block ('Loading model information...') and its HTML embeds no price data; the public API https://api.groq.com/public/v1/models/dev and /free does not list qwen/qwen3-32b; console.groq.com/llms.txt and llms-full.txt have no token prices for it; groq.com/pricing.md, console.groq.com/pricing and console.groq.com/docs/pricing(.md) are 404 pages.
IBM ibm.com/products/watsonx-ai/pricing OK

Format html

Server-rendered HTML; the model tables are in the visible text. Granite 4 H Small is read from the 'IBM Foundation Models' table (pay-as-you-go, per million tokens). Granite Embedding 278M Multilingual is read from the 'Embedding model library' table. The hosting/deploy-on-demand column is per-hour GPU pricing and is not tracked.

Models on ibm.com/products/watsonx-ai/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Granite 4 H Small$0.0636not listed$0.265$0.0636not listed$0.265Agree
Granite Embedding 278M Multilingual$0.106not listednot listed$0.106not listednot listedAgree
Inference.net inference.net/models OK

Format html

Server-rendered HTML model catalogue (redirects to /models/); all cards are in the static HTML. Each card normalizes to: name line, '<ctx> input context... $in / $out' summary line, three price lines in header order, then the API slug line. The header normalizes to the single line 'InputCache readOutput' (tooltips: 'Input price per 1M tokens', 'Cached input price per 1M tokens', 'Output price per 1M tokens'). Recipes use that header as the section, so a reordered or added column fails instead of mapping fields wrongly, and they anchor on the exact name and slug lines so sibling SKUs cannot match. Prices are the standard synchronous serverless rates; the async /v1/slow API has no published rate. The catalogue rounds displayed prices to three decimals (e.g. DeepSeek V4 Flash cached 0.0028 shows as $0.003). Fallback surface (checked 2026-10-08, no auth, matched all 63 cards): the page's own data API https://observability-api.inference.net/modelCatalog.listPublic?input={"cacheVersion":"v10","surface":"catalog"} with result.data[catalogId=inference-net/schematron-v2-turbo].pricing.costInputPerToken / costCachedInputPerToken / costOutputPerToken (per token, scale 1000000). Not used because it is an undocumented internal tRPC route and the engine reads a renamed JSON field as a valid null instead of an error.

Models on inference.net/models
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Schematron V2 Turbo$0.03$0.03$0.15$0.03$0.03$0.15Agree
Schematron V2 Small$0.05$0.05$0.23$0.05$0.05$0.23Agree
Meituan longcat.ai/platform/docs/Pricing/LongCat-2.0.html OK

Format html

page_url is an orphaned file. The static host keeps files from old builds, and page_url still serves the 2026-08-20 docs build (app.BEWjZ5om.js, chunk Pricing_LongCat-2.0.md.Bxi8E3M7). The current site does not link it, so it gets no price updates; /platform/docs/pricing/long-cat-2.0 is the same orphan. fetch_url is the current LongCat-2.0 pricing page that the docs sidebar links to (build app.DMwFWoaD.js, chunk Pricing_longcat-2.0.md.DlMFDZCW, 2026-09-26); /platform/docs/Pricing/longcat-2.0.html serves the same build. The page is server-rendered. No machine-readable twin exists: the .md paths, /llms.txt, /platform/docs/llms.txt and /platform/docs/llms-full.txt each return HTTP 200 with a JSON ResourceNotFoundException body, accept: text/markdown returns the HTML, and GET /openai/v1/models needs an API key. The VitePress page chunk on s3.meituan.net holds the page markdown (frontmatter rawMarkdown), but its file name carries a build hash that changes on each deploy, so it cannot be a fixed fetch_url. The page has a USD table (data-edition-only="ai", shown on longcat.ai) and then a CNY table (data-edition-only="chat", for longcat.chat) with the same row labels. The patterns require '$' before the number, so they read only the USD table. The section starts at the heading that names LongCat-2.0, because LongCat-2.5-Preview has its own page with the same layout and prices. The recipes read the first price column. Today the only column is 'Discounted Price $/1M Tokens (limited-time)'. If a list-price column returns in front of it, as in the 2026-08-20 build, the recipes read the list price, as MPW does.

Models on longcat.ai/platform/docs/Pricing/LongCat-2.0.html
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
LongCat-2.0$0.30$0.006$1.20$0.75$0.015$2.95Mismatch

Differences from Model Price Watch

  • LongCat-2.0: Input page 0.3 vs MPW 0.75; Output page 1.2 vs MPW 2.95; Cached input page 0.006 vs MPW 0.015
    The current official page (fetch_url, the LongCat-2.0 page that the docs sidebar links to, build of 2026-09-26) has one USD price column, 'Discounted Price $/1M Tokens (limited-time)': Uncached Input $0.30, Cached Input $0.006, Output $1.20. The API Pay-As-You-Go guide (/platform/docs/api-pay-as-you-go) shows the same single column for LongCat-2.5-Preview and LongCat-2.0. MPW tracks the list price ($0.75 input, $0.015 cached, $2.95 output). That list price appears only in the orphaned 2026-08-20 build that page_url still serves; no current official surface states a list price. The tracker follows the current page.
Mistral mistral.ai/pricing/api OK

Format html

page_url 301-redirects to https://docs.mistral.ai/inference/pricing, a server-rendered Next.js page; fetch_url keeps page_url so the tracker follows Mistral's own redirect. In the raw HTML (checked 2026-10-08) every price table renders with USD pressed (EUR not), the 'Regional inference' checkbox unchecked and the 'Standard' pricing mode pressed (Batch and Priority not), so the visible numbers are the standard pay-as-you-go global prices; Batch is a 50% discount per the Batch Processing docs and is not tracked. Columns are Model | Input | Cached input | Output in USD per 1M tokens ('Prices /M Tokens'). Cached input is the cache-read price: the docs prompt-caching page says cached prompt tokens are billed at 10% of the standard input price, which matches every row; it is not a cache write. Each recipe's section pins the section heading, the 'Prices /M Tokens' label and that exact header row within a few lines, so a column reorder, an added column or a unit change fails loudly instead of shifting fields. Rows must be name, three '$' cells and the line end; the link arrow ' ↗' is optional. Surfaces checked 2026-10-08: /inference/pricing.md is 404, 'accept: text/markdown' returns the same HTML, docs llms.txt does not list the pricing page and llms-full.txt has no price tables, so the HTML is the surface. Promotions: a sale row renders as 'Name ↗Sale price | Original price: $X Sale price: $Y | ...' (Mistral Large 4 on 2026-10-08); the patterns then fail on purpose instead of silently reading the list or sale price, so decide which price to track before updating them.

Models on mistral.ai/pricing/api
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Codestral$0.30$0.03$0.90$0.30not listed$0.90Agree
Codestral 2508$0.30$0.03$0.90$0.30not listed$0.90Agree
Ministral 3 3B$0.10$0.01$0.10$0.10not listed$0.10Agree
Mistral Large 3$0.50$0.05$1.50$0.50not listed$1.50Agree
Mistral Medium 3.5$1.50$0.15$7.50$1.50not listed$7.50Agree
Mistral Small 4$0.15$0.015$0.60$0.15not listed$0.60Agree

Untracked models

  • Voxtral Small 24B: Not on the pricing page (fetched 2026-10-08). The Specialized models table renders only OCR 4.1, Voxtral Mini Transcribe 2, Voxtral TTS and Mistral Moderation 2, and the page's RSC payload builds that table from the slug list ocr-4-1, voxtral-mini-transcribe-26-02, voxtral-tts-26-03, mistral-moderation-26-03, with no voxtral-small-25-07. 'Voxtral Small' appears in the raw HTML only as a sidebar link. The model is still current (docs.mistral.ai/models lists Voxtral Small v25.07, Apache 2.0, and it is not in the deprecated or retired table), and its official model card https://docs.mistral.ai/models/voxtral-small-25-07 shows $0.004 /Min audio and $0.1 /M Tokens input, and $0.4 /M Tokens output, which matches MPW. That card is a different page, and this source has one fetch_url, which must stay on the pricing page for the other six models. Other surfaces checked: /inference/pricing.md and /models/voxtral-small-25-07.md are 404, 'accept: text/markdown' returns the same HTML, docs llms.txt does not list the pricing page, and llms-full.txt has no Voxtral Small price.
Baichuan novita.ai/models OK

Format markdown

Novita serves an official markdown twin of /models (Accept: text/markdown; the .md URL variant https://novita.ai/models.md returns the same text). Each LLM card is a name line (optionally prefixed by a logo image), then '$X/MtInput', an optional '$Y/MtCache Read' line and '$Z/MtOutput'; on 2026-10-08 all 104 LLM cards used one of these two layouts. Verified official fallback: the models API https://api.novita.ai/openai/v1/models (answered without auth on 2026-10-08, although the docs say Bearer auth) has data[id=baichuan/baichuan-m2-32b] with pricing.prompt.price_per_m_decimal and pricing.completion.price_per_m_decimal (pricing.input_cache_read for models that have cache reads).

Models on novita.ai/models
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Baichuan M2-32B$0.07not listed$0.07$0.07not listed$0.07Agree
Apodex platform.apodex.ai/docs/pricing OK

Format html

Server-rendered HTML; the 'LLM API — Core Models' table is in the visible text (header 'Model | Input / 1M tokens | Cached input | Output / 1M tokens | Access'). No machine-readable twin exists: /docs/pricing.md, /llms.txt, /llms-full.txt and /docs/llms.txt all return 404, and www.apodex.ai/pricing is a consumer credit page with no token rates. The server-rendered rows print the standard list prices, which MPW tracks. The 20%-off campaign ('promoBanner', 'promoBadge', 'promoNote: Struck-through prices are the standard list prices') and a free campaign ('freeBadge', 'freeNote') exist only as i18n strings: in the RSC payload every price cell is [false, <span>$x</span>] and every name cell has false badge slots, so no campaign is rendered server-side. If a campaign turns on server-side, a price cell will probably hold two numbers or the name line a badge; the strict ' |' after each captured number and the name line anchor then make the recipe fail instead of reading the wrong number, so recheck which number is the list price. Both recipes' section requires the heading followed by the header row, so a column reorder or unit change fails instead of mis-mapping. Requests over 200K input tokens bill at 2x; base tier is tracked.

Models on platform.apodex.ai/docs/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Apodex 1.1 Mini$0.10$0.01$1.00$0.10$0.01$1.00Agree
Apodex 1.1$0.30$0.03$3.00$0.30$0.03$3.00Agree
Anthropic platform.claude.com/docs/en/about-claude/pricing.md OK

Format markdown

Official markdown twin of https://platform.claude.com/docs/en/about-claude/pricing. It returns the same text/markdown document for any Accept header; the explicit accept header only states the intent and keeps one shared fetch with platform-claude-com-docs-en-about-claude-pricing, which fetches the same URL with the same header, so a page change affects both sources. The recipe reads only the "## Model pricing" table: section anchors on its header row "| Model | Base input tokens | 5m cache writes | 1h cache writes | Cache hits and refreshes | Output tokens |" (the two cache-write names are wildcards), which proves the column order the pattern relies on; section_end ^(?!\|) stops at the first line that is not a table row. Cached comes from the Cache hits and refreshes column (footnote 1: 0.025x base input). The Claude Mythos 5.1 row has the same prices but is a different model, and MPW does not cite it for this page.

Models on platform.claude.com/docs/en/about-claude/pricing.md
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Claude Fable 5.1$10.00$0.25$50.00$10.00$0.25$50.00Agree
Anthropic platform.claude.com/docs/en/about-claude/pricing OK

Format markdown

fetch_url is the official .md twin of page_url (same document, also fetched by platform-claude-com-docs-en-about-claude-pricing-md). page_url itself returns markdown only when the Accept header lists text/markdown (the server sends Vary: Accept; an HTML-only Accept gets the 936 KB HTML page), so the .md twin, which returns markdown for any Accept, is fetched with an explicit accept header. Recipes read only the "## Model pricing" table: section anchors on its header row "| Model | Base input tokens | 5m cache writes | 1h cache writes | Cache hits and refreshes | Output tokens |" (the two cache-write names are wildcards), which proves the column order the patterns rely on; section_end ^(?!\|) stops at the first line that is not a table row. Each pattern ends the model name with " |" (Opus 5 cannot match Opus 5.5, Fable 5 cannot match Fable 5.1, Sonnet 5 cannot match Sonnet 5.5) and requires exactly six cells, so an added, removed or reordered column fails loudly instead of shifting fields. Page rows that MPW does not cite for this page are not tracked: Mythos 5.1, Mythos 5, Opus 4.1, Opus 4, Sonnet 4, Haiku 5.5 (two prompt-length tiers) and Haiku 3.5. Fable 5.1 is tracked by the -md source.

Models on platform.claude.com/docs/en/about-claude/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Claude Fable 5$10.00$1.00$50.00$10.00$1.00$50.00Agree
Claude Haiku 4.5$1.00$0.10$5.00$1.00$0.10$5.00Agree
Claude Opus 4.5$5.00$0.50$25.00$5.00$0.50$25.00Agree
Claude Opus 4.6$5.00$0.50$25.00$5.00$0.50$25.00Agree
Claude Opus 4.7$5.00$0.50$25.00$5.00$0.50$25.00Agree
Claude Opus 4.8$5.00$0.50$25.00$5.00$0.50$25.00Agree
Claude Opus 5$5.00$0.50$25.00$5.00$0.50$25.00Agree
Claude Opus 5.5$4.00$0.20$20.00$4.00$0.20$20.00Agree
Claude Sonnet 4.5$3.00$0.30$15.00$3.00$0.30$15.00Agree
Claude Sonnet 4.6$3.00$0.30$15.00$3.00$0.30$15.00Agree
Claude Sonnet 5$2.00$0.20$10.00$2.00$0.20$10.00Agree
Claude Sonnet 5.5$2.00$0.10$10.00$2.00$0.10$10.00Agree
Moonshot platform.kimi.ai/docs/pricing/chat-k3 OK

Format markdown

page_url /docs/pricing/chat-k3 now returns 308 to /docs/pricing/chat, and chat-k3.md returns 307 to chat.md (checked 2026-10-08). This source therefore fetches the same official markdown twin as platform-kimi-ai-docs-pricing-chat, https://platform.kimi.ai/docs/pricing/chat.md, because the HTML draws the price rows on the client. The recipe reads the K3 Series Models DocTable row. The section regex pins the complete column list from Model through Output Price.

Models on platform.kimi.ai/docs/pricing/chat-k3
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Kimi K3$3.00$0.30$15.00$3.00$0.30$15.00Agree
Moonshot platform.kimi.ai/docs/pricing/chat OK

Format markdown

Fetches the official Mintlify markdown twin (page_url + .md, listed in https://platform.kimi.ai/docs/llms.txt). The HTML page draws the DocTable rows on the client; its embedded MDX payload has the same numbers as the .md twin (checked 2026-10-08). llms-full.txt strips the DocTable props, so it has no prices. The .md body keeps the DocTable JSX rows such as ["kimi-k2.6", "1M tokens", <>{"$"}0.16</>, ...]. Each section regex pins the complete column list from Model through Output Price, so an inserted, removed or reordered column fails the recipe instead of shifting the field mapping. Price cells match both the JSX form <>{"$"}0.16</> and the string form "$0.16" (used on the official batch page). Inference is pay-as-you-go with one context tier; the separate batch page (/docs/pricing/batch) is a different tier and is not read.

Models on platform.kimi.ai/docs/pricing/chat
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Kimi K2.6$0.95$0.16$4.00$0.95$0.16$4.00Agree
Kimi K2.7 Code$0.95$0.19$4.00$0.95$0.19$4.00Agree
Kimi K2.7 Code HighSpeed$1.90$0.38$8.00$1.90$0.38$8.00Agree

Untracked models

  • Kimi K2.5: kimi-k2.5 is retired, so no official price surface remains. The pricing page K2 table lists only kimi-k2.7-code, kimi-k2.7-code-highspeed and kimi-k2.6 (HTML payload and .md twin agree). The official model list (https://platform.kimi.ai/docs/models.md) says kimi-k2.5 was officially discontinued on August 31, 2026 and marks it Deprecated. The official changelog (August 2026, in https://platform.kimi.ai/docs/llms-full.txt) says kimi-k2.5 was retired across all platforms and calls now return a 404 "model not found" error. The K2.5 quickstart guide now redirects to /docs/overview, the batch page does not list it, llms-full.txt has no kimi-k2.5 price row, and the China platform pricing twin (platform.moonshot.cn redirects to https://platform.kimi.com/docs/pricing/chat.md) has no kimi-k2.5 row either. Checked 2026-10-08.
MiniMax platform.minimax.io/docs/guides/pricing-paygo OK

Format markdown

Fetches the official Mintlify markdown twin of the page (page URL + .md, text/markdown, about 11KB) instead of the 588KB HTML. The .md URL is listed in the official docs index https://platform.minimax.io/docs/llms.txt, and the HTML page shows the same prices. LLM section: MiniMax-M3 sits inside <Tabs> with a Standard tab and a Priority tab (1.5x); MiniMax-M2.7 is in the untabbed table after </Tabs>. Prices are USD per 1M tokens. The MiniMax-M3 rows show a struck-through list price (~~$0.60~~) followed by the current 'Permanent 50% off' price; the recipe skips the struck value and reads the current one.

Models on platform.minimax.io/docs/guides/pricing-paygo
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
MiniMax-M3$0.30$0.06$1.20$0.30$0.06$1.20Agree
MiniMax-M2.7$0.30$0.06$1.20$0.30$0.06$1.20Agree
OpenAI platform.openai.com/docs/pricing OK

Format markdown

Fetches the official markdown twin https://developers.openai.com/api/docs/pricing.md (text/markdown, 200, no auth). page_url redirects to https://developers.openai.com/api/docs/pricing; that HTML shows only the first 3 rows of each table (the rest sit behind an "All models" expander), while the .md twin holds every row and equals the page's own text/markdown response. The same fetch_url is shared by platform-openai-com-docs-pricing, developers-openai-com-api-docs-pricing and developers-openai-com-api-docs-pricing-md so the tracker fetches it once. Text-model recipes read the Flagship models "### Standard pricing data" table, short-context columns. Each section pins its heading and its column header (here "Model | Short context input | Short context cached input | Short context cache writes | Short context output"), so a column reorder or insertion fails loudly instead of shifting fields. section_end "^[^|\n]" ends the section at the first line that is not a table row, so each scope is exactly one table: Batch, Flex, Fast, Ultrafast and the Cyber models table stay out. The image and embedding recipes also pin the "Standard" tab label. Anchors end the model name at a " | " cell boundary (optionally after "(<272K context length)") so -mini, -nano, -pro and dated snapshots cannot match.

Models on platform.openai.com/docs/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
GPT-4.1$2.00$0.50$8.00$2.00$0.50$8.00Agree
GPT-4.1 mini$0.40$0.10$1.60$0.40$0.10$1.60Agree
GPT-5.4$2.50$0.25$15.00$2.50$0.25$15.00Agree
GPT-5.4 mini$0.75$0.075$4.50$0.75$0.075$4.50Agree
GPT-5.5$5.00$0.50$30.00$5.00$0.50$30.00Agree
GPT-5.6 Sol$4.00$0.40$20.00$4.00$0.40$20.00Agree
GPT-5.6 Terra$2.00$0.20$12.00$2.00$0.20$12.00Agree
GPT-5.6 Luna$0.20$0.02$1.20$0.20$0.02$1.20Agree
GPT-Image-2$8.00$2.00$30.00$8.00$2.00$30.00Agree
text-embedding-3-large$0.13not listednot listed$0.13not listednot listedAgree
text-embedding-3-small$0.02not listednot listed$0.02not listednot listedAgree
GPT-5.5 Pro$30.00not listed$180.00$30.00not listed$180.00Agree
GPT-5.4 Pro$30.00not listed$180.00$30.00not listed$180.00Agree
GPT-5.4 nano$0.20$0.02$1.25$0.20$0.02$1.25Agree
GPT-5.2$1.75$0.175$14.00$1.75$0.175$14.00Agree
GPT-5.2 Pro$21.00not listed$168.00$21.00not listed$168.00Agree
Relace relace.ai/pricing OK

Format markdown

Surface: the official Markdown twin of this Framer-hosted page, served at the same URL (after the 308 to https://relace.ai/pricing) for 'Accept: text/markdown'. Framer documents it in 'Make your site readable by AI agents' (www.framer.com/help, updated 2026-09-15): every optimized page has a markdown version, also reachable with '?md'. The header is pinned because the server sends Vary: Accept and honors q-values, so the engine default Accept (text/html first, text/markdown;q=0.9) gets the 476KB HTML instead. On 2026-10-08 the twin (3.4KB) and the HTML had the same publish stamp (Sep 23, 2026, 1:23 AM UTC) and the same four cards. The twin's YAML front matter carries that 'published' stamp, so the page hash also changes on each Framer publish, even with no visible edit. The twin renders the 'Individual models (based on token usage)' grid once, one card per line: '<name>, $X, /million, (Input Tokens), $Y, /million, (Output Tokens)'; the first card line starts with the tab labels 'Models, Repos, Jacq, '. Recipes anchor the exact name between a line start or ', ' and the next separator, and check the '/million' unit and the Input/Output labels. The section ends at the 'For information about our policies' line. The page publishes no cached-input, batch or other tier rates. relace-compact ($0.20 in, $0.20 out) and relace-rank ($0.05 input, no output) are on the page but not in the MPW hint set. Fallback if the twin stops being served (Framer says markdown can be missing for pages that are not optimized or sites that are rate-limited; the fetch then returns HTML and the recipes fail instead of misreading): set format html with accept text/html, or fetch https://relace.ai/pricing?md. The patterns accept ', ' or a newline between the card parts, so they read the HTML text (one part per line) without edits. The HTML repeats the grid for the desktop, tablet and phone breakpoints (Framer ssr-variant blocks, desktop first); section_end also stops at the second copy's heading.

Models on relace.ai/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Relace Search$1.00not listed$3.00$1.00not listed$3.00Agree
Relace Apply 3$0.80not listed$1.20$0.80not listed$1.20Agree
Sakana AI sakana.ai/fugu OK

Format html

Server-rendered HTML; prices are in the bilingual (EN/JA) 'Token Plan' (pay-as-you-go) cards, one value per line. Each recipe is limited to its own card (Fugu Ultra card ends at 'Fugu Max'; Fugu Max card ends at 'Fugu Cyber'), so the $20/$100/$200 subscription plans and the FAQ prose cannot match. The page also has a plain 'Fugu' card (no fixed rate; bills at the routed model's standard rate) and a 'Fugu Cyber' card (contact sales); neither has a per-token price and MPW has no record for them. The same rates also appear on console.sakana.ai/pricing (source console-sakana-ai-pricing). No machine-readable twin exists (checked 2026-10-08): sakana.ai/llms.txt, /fugu.md and /fugu/index.md return 404, and 'accept: text/markdown' still returns HTML.

Models on sakana.ai/fugu
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Fugu Ultra$5.00$0.50$30.00$5.00$0.50$30.00Agree
Fugu Max$2.00$0.25$6.00$2.00$0.25$6.00Agree
Kwaipilot streamlake.com/product/wanqing OK

Format html, CNY converted at 0.148201 USD

StreamLake (Kuaishou Wanqing) product page, server-rendered HTML. Prices are read from the model price table (header 模型名称 | 分段 | 输入价格 | 输出价格 | 缓存命中 | 价格单位), quoted in 元/百万 tokens (CNY per 1M tokens). usd_rate 0.148201 = 1/6.7476, the CNY/USD rate Model Price Watch cites in price_note (2026-08-07); update it if MPW re-derives at a new FX rate. Each table cell is one line in the normalized text.

Models on streamlake.com/product/wanqing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
KAT-Coder-Pro V2.5$0.741$0.1482$2.96$0.741$0.148$2.96Agree
KAT-Coder-Air V2.5$0.1482$0.02964$0.5928$0.148$0.03$0.593Agree
Thinking Machines tinker-docs.thinkingmachines.ai/tinker/models OK

Format json

Fetches serverless.json, the official machine-readable twin of the page's Serverless Inference (Beta) table. The page links it as ../../serverless.json, its markdown twin (index.md) gives this absolute URL, and its 'Machine-Readable Pricing' section names models.json (Training table) and serverless.json (Serverless table) as the stable interface and says not to scrape the HTML tables. The page_url redirects to /tinker/models/models_and_pricing/. Root is an array of {name, tinker_id, context, url, input, cached_input, output}; prices are per 1M tokens as '$x.xx' strings. The HTML columns are 'Prefill (Input)', with the '(cached)' prompt-cache-hit price under it, and 'Sample (Output)', so input, cached_input and output map to input, cached and output. MPW tracks the serverless sampling-nvfp4 256K endpoint, not the Training table (prefill/sample/train legs, currently showing a limited-time 50% discount). Recipes select by tinker_id, not by name: models.json uses the bare name 'Inkling' for the 64K training tier and 'Inkling (256K)' for the 256K tier, so a name select could move silently to another tier if a second serverless endpoint is added. A tinker_id select reads the same endpoint or fails loudly. If Thinking Machines changes the endpoint suffix (for example a new quantization), re-check the serverless table and update the tinker_id.

Models on tinker-docs.thinkingmachines.ai/tinker/models
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Inkling$1.00$0.17$4.05$1.00$0.17$4.05Agree
Inkling Small$0.30$0.06$1.20$0.30$0.06$1.20Agree
Together together.ai/pricing OK

Format markdown

Surface: the official Markdown twin of www.together.ai/pricing, served for 'Accept: text/markdown'. www.together.ai/llms.txt ('For AI agents') says every www.together.ai page returns Markdown for that header and that .md paths return 404. Webflow builds the twin from the same CMS data and edge cache as the HTML (same surrogate keys and cache age on 2026-10-08), and all 28 Chat rows match the HTML rows. The header is pinned because the engine's default Accept also lists text/markdown;q=0.9 and the server varies on Accept. Each Chat row is one line: '| [![](logo) <name>](https://www.together.ai/models/<slug>) | $<input>$<cached> (cached) | $<output> |'. Recipes anchor the exact name between the logo (or '[') and ']'. The section starts at the Chat header ('Price per 1M tokens' / 'Batch API price' / 'Model | Input | output', which also checks the column order) and ends at the next 'Displayed prices refer' or 'Price per ' line, so the Vision, Image and Fine-Tuning tables, which repeat several names, are excluded. 'Batch API price' is a client-side toggle: the page script discounts only cells whose data-batch attribute holds a positive percentage, and on 2026-10-08 every tracked Chat cell had data-promo="false" and no data-batch value, so the static prices are serverless pay-as-you-go. A promo row (Kimi K3 on 2026-10-08) appends 'PROMO...' to the name and shows two prices per cell, so a tracked model on promo fails loudly instead of being misread. Fallback if the twin stops being served: the HTML of the same URL with accept text/html and format html has the same table, one cell per line. Official cross-checks: www.together.ai/models/<slug> (all nine cards agree with the table on 2026-10-08, one serverless tier each) and docs.together.ai/docs/serverless/models.md. The docs table does not always agree with the pricing page (Qwen3.8 Flash 0.09/0.282 vs 0.15/0.47; cached 0.50 vs 0.30 for Qwen3.7-Max and 0.50 vs 0.25 for Qwen3.8-2.4T-A95B), so the pricing page stays the source. The models API (api.together.ai/v1/models) needs an API key (401).

Models on together.ai/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
GLM-5.3$1.40$0.26$4.40$1.40$0.26$4.40Agree
DeepSeek V4 Pro$1.32$0.13$3.96$1.32$0.13$3.96Agree
GLM-5.2$1.40$0.26$4.40$1.40$0.26$4.40Agree
Muse Glimmer 30B$0.35$0.04$1.50$0.35$0.04$1.50Agree
Llama 3.3 70B$1.04not listed$1.04$1.04not listed$1.04Agree
MiniMax-M3$0.30$0.06$1.20$0.30$0.06$1.20Agree
Qwen3.7-Max$1.50$0.30$4.50$2.50$0.25$7.50Mismatch
Qwen3.7-Plus$0.32not listed$1.28$0.32not listed$1.28Agree
GPT-OSS 120B$0.15not listed$0.60$0.15not listed$0.60Agree

Differences from Model Price Watch

  • Qwen3.7-Max: Input page 1.5 vs MPW 2.5; Output page 4.5 vs MPW 7.5; Cached input page 0.3 vs MPW 0.25
    Together cut this SKU's price after MPW's 2026-09-21 read. On 2026-10-08 the Chat row 'Qwen3.7-Max' (link /models/qwen37-max) reads $1.50 / $0.30 (cached) / $4.50, with no PROMO tag (HTML data-promo="false"). The model card www.together.ai/models/qwen37-max agrees: 'Endpoint Qwen/Qwen3.7-Max ... Input price $1.50 / 1M tokens, $0.30 (cached)/1M, Output price $4.50 / 1M tokens', with one serverless tier and no long-context tier. docs.together.ai/docs/serverless/models.md shows '| Qwen | Qwen3.7 Max | Qwen/Qwen3.7-Max | - | $1.50 | $0.50 | $4.50 |'. Wayback snapshots of the same pricing row (history evidence only) match the history in MPW's note: $2.00/$0.25/$6.00 on 2026-09-11, $2.50/$0.25/$7.50 on 2026-09-18 (MPW's figures), and $1.50/$0.30/$4.50 on every snapshot from 2026-09-23 to 2026-10-07. It is the same SKU (the only Qwen3.7-Max row; Qwen3.8-2.4T-A95B is a separate endpoint) and the same tier (serverless pay-as-you-go). MPW's 2.50/0.25/7.50 equals Alibaba's first-party rate and is out of date for Together. Cached input is $0.30 on the pricing page and the model card but $0.50 in the docs table; the tracker follows the pricing page.
Unbiased unbiased.ai/pricing OK

Format html

Static HTML; the 'Pay as you go' rate block (Input / Cached input / Output, each '$x / Mtok') is in the visible text. Unbiased sells only Pareto, so the block is not labelled with the model name. The Personal and Personal Max subscriptions are weekly token allowances, not per-token prices, and are ignored. A machine-readable restatement exists at https://unbiased.ai/llms.txt ('Pareto rates: $0.80 in / $0.03 cached / $3.20 out per MTok.') if the HTML layout breaks.

Models on unbiased.ai/pricing
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Pareto$0.80$0.03$3.20$0.80$0.03$3.20Agree
Upstage upstage.ai/pricing/api OK

Format html

Surface: the official HTML page (Webflow, server-rendered). No official machine-readable surface carries the current price: the console.upstage.ai/docs/models/<model>.md twins say the content is 'rendered on the docs page', the console docs HTML shows only the standard list prices, and api.upstage.ai/v1/models needs an API key. Each .pricing-card-v2 card's markup holds the STANDARD price. The inline promo script (v3.3) repaints it in the browser with the schedule window open at the current time. The schedule is a hidden data-promo="schedule" element at the end of each card, one window per line: startISO|endISO|input=..|cached=..|output=..[|note=short]. A window is open when start <= now < end; an empty endISO is open-ended; overlapping windows: the later start wins; a malformed line is dropped; with no open window the standard card price shows. note=short only shortens the footnote. In the normalized text these lines follow the card's Output row. Schedule keys map to the data-rate rows: input = 'Input', cached = 'Input(Cached)' (cache read), output = 'Output'. The engine cannot pick a window by fetch time, so the Solar Pro 4 and Solar Mini 4 recipes are pinned to the windows open on 2026-10-08. Both windows end at 2026-10-11T00:00:00Z, and expired lines stay in the schedule, so after that instant the recipes keep reading the old window WITHOUT failing. MAINTAINER ACTION on 2026-10-11 UTC: (1) point Solar Pro 4 at '2026-10-11T00:00:00Z||input=..' (0.15/0.03/0.60 on 2026-10-08, '50% off until further notice'; MPW's note expects the standard 0.30/0.06/1.20 from Oct 10, so expect a mismatch then); (2) point Solar Mini 4 at '2026-10-11T00:00:00Z|2026-10-23T00:00:00Z|..' and remove its accept_mismatch (MPW's 0.05/0.005/0.20 should then match). On 2026-10-23 UTC: point Solar Mini 4 at the standard card rows (0.10/0.01/0.40 on 2026-10-08). Solar Pro 3, Solar Pro 2 and Solar Mini reach end of service on 2026-10-30 (KST); expect their cards to go. MPW cites only Solar Pro 3, Solar Pro 4 and Solar Mini 4 from this page; Solar Pro 2, Solar Mini, Embed, Embed 2, File Search and the Document AI products have no recipes. All prices exclude 10% VAT.

Models on upstage.ai/pricing/api
ModelTracked inputTracked cachedTracked outputMPW inputMPW cachedMPW outputAgreement
Solar Pro 3$0.15$0.015$0.60$0.15$0.015$0.60Agree
Solar Pro 4$0.09$0.018$0.36$0.09$0.018$0.36Agree
Solar Mini 4$0.03$0.003$0.12$0.05$0.005$0.20Mismatch

Differences from Model Price Watch

  • Solar Mini 4: Input page 0.03 vs MPW 0.05; Output page 0.12 vs MPW 0.2; Cached input page 0.003 vs MPW 0.005
    On 2026-10-08 the page shows 0.03 / 0.003 / 0.12 for Solar Mini 4, not MPW's 0.05 / 0.005 / 0.20. The card's schedule has two windows: '2026-09-22T05:00:00Z|2026-10-11T00:00:00Z|input=0.03|cached=0.003|output=0.12' and '2026-10-11T00:00:00Z|2026-10-23T00:00:00Z|input=0.05|cached=0.005|output=0.20'. The page's own promo script (parseSchedule and windowOf, run on the fetched HTML at 2026-10-08T15:43Z) selects the first window: 70% off the standard card price 0.10 / 0.01 / 0.40. The page banner says 'Solar Mini 4 is live on Console. 70% off through Oct 10 (UTC)', and the launch blog (upstage.ai/blog/en/solar-mini-4) says 'Upstage Console offer a 70% launch discount through October 10 UTC'. MPW's figures are the second (50% off) window, which opens on 2026-10-11. Same SKU and tier: Solar Mini 4 card, pay-as-you-go API, per 1M tokens, excluding VAT.