ChatGPT, Codex and GLM through one API.
Model.sale exposes a familiar OpenAI-compatible base URL for ChatGPT-compatible GPT, Codex clients, GLM and DeepSeek models. Keep your key server-side, use a live model ID and inspect the request ID and usage returned with every call.
/v1/responses.Authorization: Bearer ms_live_…. Keys are shown once and can be revoked from the dashboard.GET /v1/models. A request is never silently routed to another model.Endpoints
| Method | Path | Purpose |
|---|---|---|
| GET | /v1/models | Published live catalog |
| GET | /v1/models/{id} | One callable model (OpenAI models.retrieve) |
| GET | /v1/catalog | Full priced registry with live/unavailable status |
| POST | /v1/responses | Responses JSON and SSE (GPT/Codex) |
| POST | /v1/chat/completions | Chat Completions JSON and SSE (GPT/GLM/DeepSeek) |
| POST | /v1/messages | Only when Anthropic Messages validation is published |
/v1/models intentionally contains only models admitted for customer traffic right now. /v1/catalog keeps every priced ID visible with its probe result and supported endpoints, including models that passed a live check but still await an audited price/margin publication. A model is admitted only after a fresh JSON/SSE check, terminal usage and publication approval; it is never silently substituted.Protocol capabilities
Use /v1/catalog to see each model's current supported_endpoints. Codex requires a live model with /v1/responses; GLM and DeepSeek currently use /v1/chat/completions. Claude IDs remain visible when priced, but Messages access is enabled only after its validation passes.
Minimal Responses request
Current published Responses example: gpt-5.5. Confirm it still appears in GET /v1/models.
curl https://api.model.sale/v1/responses \
-H "Authorization: Bearer $MODEL_SALE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-5.5","input":"Reply exactly OK"}'Streaming
Set stream: true. Responses streams finish with a terminal response event; Chat Completions streams finish with [DONE]. The gateway adds x-model-sale-request-id to every response. Usage is required for exact settlement; incomplete usage is marked for reconciliation.
Authentication
Send the key as Authorization: Bearer ms_live_… (OpenAI SDKs, Codex) or x-api-key: ms_live_… (Anthropic SDKs). Browser apps may call the API directly: CORS is enabled for /v1/* without cookies — only ship a key to a browser you control.
Errors and limits
Errors use the OpenAI shape {"error":{"message","type","param","code"}} (Anthropic shape on /v1/messages). type is the category and code the specific reason.
| Status | type | Typical code | What to do |
|---|---|---|---|
| 400 | invalid_request_error | invalid_json, model_required | Fix the request body; the message names the problem. |
| 401 | authentication_error | invalid_api_key | Check the key or create a new one in the dashboard. |
| 402 | insufficient_quota | insufficient_balance | Top up in Billing, then retry. |
| 403 | permission_error | model_not_allowed, ip_not_allowed | Adjust the key's model or IP restrictions. |
| 404 | not_found_error | model_not_found | Use an ID from GET /v1/models. |
| 413 | invalid_request_error | request_too_large | Request bodies are limited to 8 MB. |
| 429 | rate_limit_error | rate_limit_exceeded, spend_limit_exceeded | Wait for Retry-After seconds or raise the key's limits. |
| 502/503 | server_error | upstream_error, model_unavailable | Retry with backoff; failed requests are not charged. |