Add balance when needed. Funds are reserved before dispatch and settled from terminal usage.
Compare API costs before you ship.
A practical comparison of prepaid blended pricing, official split tariffs and subscription-style access. Start with your real token mix, not a headline number.
Model.sale public rates
Prices are per 1M charged tokens. The initial reservation is temporary and released after terminal usage; it is not a per-request fee. Prices remain listed when a model is temporarily unavailable.
| Model | Blended / 1M | Initial reservation | Status |
|---|---|---|---|
| gpt-5.5 | $0.45 | $0.01 (released) | LIVE |
| gpt-5.6-luna | $0.09 | $0.01 (released) | LIVE |
| gpt-5.6-sol | $0.45 | $0.01 (released) | LIVE |
| gpt-5.6-terra | $0.15 | $0.01 (released) | LIVE |
| gpt-6-astra | $1.125 | $0.01 (released) | LIVE |
| gpt-6-luna | $0.08 | $0.01 (released) | LIVE |
| gpt-6-sol | $0.38 | $0.01 (released) | LIVE |
| gpt-6.1-sol | $0.40 | $0.01 (released) | LIVE |
| claude-fable-5-1 | $1.00 | $0.01 (released) | LIVE |
| claude-haiku-4-5-20251001 | $0.20 | $0.01 (released) | LIVE |
| claude-opus-4-6 | $0.50 | $0.01 (released) | LIVE |
| claude-opus-4-7 | $0.50 | $0.01 (released) | LIVE |
| claude-opus-4-8 | $0.60 | $0.01 (released) | LIVE |
| claude-opus-5 | $0.60 | $0.01 (released) | LIVE |
| claude-opus-5-5 | $1.00 | $0.01 (released) | LIVE |
| claude-sonnet-5-5 | $0.55 | $0.01 (released) | LIVE |
| glm-5.2 | $0.15 | $0.01 (released) | LIVE |
| glm-5.3 | $0.15 | $0.01 (released) | LIVE |
| glm-5.3-flash | $0.10 | $0.01 (released) | LIVE |
| gpt-5.4 | $0.10 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| gpt-5.4-20 | $0.10 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| gpt-5.4-mini | $0.09 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| claude-haiku-4-5 | $0.075 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| claude-opus-4-5 | $0.30 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| claude-sonnet-4-5 | $0.15 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| claude-sonnet-4-6 | $0.25 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| deepseek-flash | $0.03 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| deepseek-v4-flash | $0.20 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| deepseek-v4-pro | $0.02 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| glm-5 | $0.12 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| glm-5.1 | $0.16 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| qwen-3.8 | $0.25 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| qwen3.5-plus | $0.05 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| qwen3.6-plus | $0.04 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| qwen3.7-max | $0.11 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| qwen3.7-plus | $0.03 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| kimi-k2.5 | $0.07 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| kimi-k2.6 | $0.11 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| kimi-k3 | $0.24 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| mimo-v2-omni | $0.05 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| mimo-v2-pro | $0.05 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| mimo-v2.5 | $0.01 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| mimo-v2.5-pro | $0.03 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| minimax-m2.5 | $0.02 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| minimax-m2.7 | $0.04 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| minimax-m3 | $0.04 | $0.01 (released) | TEMPORARILY UNAVAILABLE |
| gpt-4o-transcribe | $0.10 | $0.01 (released) | NOT IN THIS API RELEASE |
| gpt-image-1.5 | $20.00 | $0.01 (released) | NOT IN THIS API RELEASE |
| gpt-image-2 | $0.10 | $0.01 (released) | NOT IN THIS API RELEASE |
| gpt-image-2.5 | $20.00 | $0.01 (released) | NOT IN THIS API RELEASE |
| gpt-image-2.5-flare | $20.00 | $0.01 (released) | NOT IN THIS API RELEASE |
| gpt-image-2.5-sunburst | $20.00 | $0.01 (released) | NOT IN THIS API RELEASE |
Input, cached input and output can have different prices and context tiers.
A temporary reservation protects the wallet, then the final charge uses only terminal tokens returned by the API.
Popular API alternatives
Model.sale is a prepaid single-service catalog. The products below are useful alternatives, but they are not identical offers: some are multi-provider routers, some are direct open-model platforms, and some bill after usage. We compare the buying model and the documented trade-offs, not model quality or a universal “cheapest” claim.
| Service | Billing | Catalog / routing | Best fit | Checked |
|---|---|---|---|---|
| OpenRouter Multi-provider aggregator | Standard pay-as-you-go lists a 5.5% platform fee; Business lists 8%. The current pricing page also documents a BYOK allowance and a fee after the allowance. | The current Standard plan lists 500+ models and 80+ providers; its Free plan lists 25+ models and 4 providers. Standard includes auto-routing, preferred-provider selection, budgets and spend controls. | Teams that want breadth and provider choice in one account. Trade-off: A broad catalog and routing controls add choices; the platform fee and exact model billing still belong in the workload estimate. | 2026-10-01 · source |
| Together AI Open-model inference platform | Serverless rates are listed per 1M input and output tokens, with separate batch pricing on eligible models; dedicated and provisioned capacity have different billing units. | The live pricing page lists model-specific rates across open-model families, including DeepSeek, Qwen, Kimi, GLM and others. Choose a model from its catalog; provisioned throughput and dedicated inference are separate capacity options. | Developers who want a direct API and a broad open-model catalog. Trade-off: Compare input, cached input, output and batch rates separately; dedicated capacity is not comparable to serverless token billing. | 2026-10-01 · source |
| Fireworks AI Inference platform | Self-serve accounts moved to prepaid credits on July 1, 2026; contracted accounts are excluded from that migration. Serverless inference is per token, and on-demand deployments are priced per GPU-second. | Serverless open-model inference plus separate training and dedicated deployment products. Choose a serverless model and tier, or provision a deployment; these are distinct products rather than one cross-provider router. | Teams that want token-billed serverless inference and a path to dedicated GPU deployment. Trade-off: Self-serve prepaid credits are closer to a prepaid wallet, but credit rules and GPU-second deployment pricing still differ; compare the same workload and capacity assumptions. | 2026-10-01 · source |
| Groq Low-latency inference platform | Developer-tier usage is invoiced monthly or at progressive usage thresholds; a payment method is required to upgrade from Free. | A focused list of production and preview models, with model-specific prices, limits and lifecycle notes. Select a Groq-hosted model; the product is an inference platform, not a broad multi-provider router. | Latency-sensitive workloads that fit the currently hosted models and account limits. Trade-off: Postpaid billing, model-specific limits and preview lifecycle differ from a prepaid multi-family catalog. | 2026-10-01 · source |
Source pages and checked dates are part of the comparison. Terms, fees, model catalogs and limits can change; verify the linked documentation before making a production decision. Model.sale does not imply affiliation with any listed service.
Which API fits your workload?
Start with the operating model you need. Model catalogs, plan fees, payment rules and capabilities change; the dated source links above are the authority for each alternative.
| If you need… | Compare | Why it may fit | Check before choosing |
|---|---|---|---|
| One prepaid balance and a small tested catalog | Model.sale | One account, one API key and request-level usage records. | Confirm the model is live, its endpoint is supported and the current rate fits your token mix. |
| Many models, providers and routing controls | OpenRouter | A broad catalog with provider choice and documented routing controls. | Platform fees, route selection, provider-specific limits and fallback behavior. |
| Direct access to open-model inference | Together AI | Model-specific serverless pricing with separate capacity options. | Input/output/cache rates and whether serverless or reserved capacity is being compared. |
| Prepaid inference credits or dedicated GPU deployments | Fireworks AI | Serverless token usage and a separate path to dedicated deployments. | Credit and auto-reload rules; GPU-time costs are not per-token rates. |
| A focused catalog for latency-sensitive inference | Groq | Production and preview models with published limits. | Current model lifecycle, account limits and postpaid billing terms. |
Detailed comparisons
Model.sale vs OpenRouter
Prepaid simplicity versus catalog breadth, provider choice and routing.
Read the dated comparison →Together vs Fireworks vs Groq
Compare serving models, payment timing, capacity and operational trade-offs.
Read the decision guide →Compare the same token workload
Use a representative input/output mix instead of comparing unlike headline rates.
Read the cost guide →Illustrative workload examples
These examples compare the same 80K input + 20K output workload (100K total tokens), using standard short-context text rates. They do not claim that similarly named models have identical quality or behavior.
| Reference family | Model.sale blended | Model.sale example | Official split example |
|---|---|---|---|
| GPT-5.6 Terra | $0.15/M | $0.015 | $0.40 · $0.16 input + $0.24 output |
| GPT-5.6 Sol | $0.45/M | $0.045 | $0.72 · $0.32 input + $0.40 output |
| GPT-6 Astra | $1.125/M | $0.1125 | $1.80 · $0.80 input + $1.00 output |
The Model.sale example applies one tenth of the displayed blended rate to 100K charged tokens. The official example applies standard short-context input/output rates to the same token counts; cached input, reasoning surcharges and long-context tiers are excluded. Change the token mix and the result changes.
Prepaid vs subscription access
| Question | Prepaid API | Subscription product |
|---|---|---|
| How you pay | Deposit balance, then pay usage | Recurring seat or plan fee |
| Best for | Variable workloads, scripts and coding agents | Predictable seats and bundled product features |
| Cost visibility | Request-level token and ledger records | Usually plan usage limits or quotas |
| Risk to budget | Set key and daily spend limits | Recurring charge until cancelled |
How to calculate your effective cost
- Export a representative hour or day of input, cached and output tokens.
- Separate retries, failed calls and streaming disconnects.
- Apply each provider’s exact input/output/cache rules and context tier.
- Apply Model.sale blended price to the measured charged-token total; the initial reservation is released after settlement.
- Compare the resulting total, latency, availability and model capability together.
Official reference rows below are informational, dated and linked to their source. They are not a promise of parity or a benchmark.
Official published rates
USD per million tokens, standard API rates for the source's standard/short-context tier. Cached input means OpenAI cached input or Anthropic cache-read pricing as applicable; cache writes, long context, batch, priority modes and other surcharges are excluded. These reference rates do not mean a model is currently available through Model.sale. Availability is shown separately above.
| Model | Input | Cached input / cache read | Output | Source |
|---|---|---|---|---|
| OpenAI GPT-6 Astra | $10.00 | $1.00 | $50.00 | Official pricing |
| OpenAI GPT-6.1 Sol | $2.00 | $0.10 | $10.00 | Official pricing |
| OpenAI GPT-6 Luna | $0.10 | $0.01 | $0.50 | Official pricing |
| OpenAI GPT-5.6 Sol | $4.00 | $0.40 | $20.00 | Official pricing |
| OpenAI GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | Official pricing |
| Anthropic Claude Opus 5.5 | $4.00 | $0.20 | $20.00 | Official pricing |
| Anthropic Claude Sonnet 5.5 | $2.00 | $0.20 | $10.00 | Official pricing |
| Anthropic Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | Official pricing |
| Anthropic Claude Opus 4.8 | $5.00 | $0.50 | $25.00 | Official pricing |
| Anthropic Claude Sonnet 4.6 | $3.00 | $0.30 | $15.00 | Official pricing |
Not an equivalence claim
Reference prices are not a claim that the same model is enabled in the Model.sale catalog. Names, context limits and quality tiers can differ; treat the table as a billing-shape reference.
Rates can change
Check the source date and verify official pages before committing to a workload.
Need exact math?
Read the pricing methodology for measured usage, reservations and settlement.