OpenAI-compatible API documentation

ChatGPT, Codex and GLM through one API.

Model.sale exposes a familiar OpenAI-compatible base URL for ChatGPT-compatible GPT, Codex clients, GLM and DeepSeek models. Keep your key server-side, use a live model ID and inspect the request ID and usage returned with every call.

ChatGPT-compatible GPTUse Responses or Chat Completions from an OpenAI SDK. The model ID stays unchanged end to end.
CodexConnect Codex CLI, App or VS Code with the Responses protocol and select a model whose catalog row includes /v1/responses.
GLM and DeepSeekCall the live open-model IDs with Chat Completions and optional SSE streaming.
1. AuthenticateSend Authorization: Bearer ms_live_…. Keys are shown once and can be revoked from the dashboard.
2. Add balanceDeposit at least $5. Funds are reserved before dispatch and settled from terminal usage.
3. Call a modelUse a model from GET /v1/models. A request is never silently routed to another model.

Endpoints

MethodPathPurpose
GET/v1/modelsPublished live catalog
GET/v1/models/{id}One callable model (OpenAI models.retrieve)
GET/v1/catalogFull priced registry with live/unavailable status
POST/v1/responsesResponses JSON and SSE (GPT/Codex)
POST/v1/chat/completionsChat Completions JSON and SSE (GPT/GLM/DeepSeek)
POST/v1/messagesOnly when Anthropic Messages validation is published
Why are some catalog models missing from /v1/models? /v1/models intentionally contains only models admitted for customer traffic right now. /v1/catalog keeps every priced ID visible with its probe result and supported endpoints, including models that passed a live check but still await an audited price/margin publication. A model is admitted only after a fresh JSON/SSE check, terminal usage and publication approval; it is never silently substituted.

Protocol capabilities

Use /v1/catalog to see each model's current supported_endpoints. Codex requires a live model with /v1/responses; GLM and DeepSeek currently use /v1/chat/completions. Claude IDs remain visible when priced, but Messages access is enabled only after its validation passes.

Minimal Responses request

Current published Responses example: gpt-5.5. Confirm it still appears in GET /v1/models.

curl https://api.model.sale/v1/responses \
  -H "Authorization: Bearer $MODEL_SALE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-5.5","input":"Reply exactly OK"}'

Streaming

Set stream: true. Responses streams finish with a terminal response event; Chat Completions streams finish with [DONE]. The gateway adds x-model-sale-request-id to every response. Usage is required for exact settlement; incomplete usage is marked for reconciliation.

Authentication

Send the key as Authorization: Bearer ms_live_… (OpenAI SDKs, Codex) or x-api-key: ms_live_… (Anthropic SDKs). Browser apps may call the API directly: CORS is enabled for /v1/* without cookies — only ship a key to a browser you control.

Errors and limits

Errors use the OpenAI shape {"error":{"message","type","param","code"}} (Anthropic shape on /v1/messages). type is the category and code the specific reason.

StatustypeTypical codeWhat to do
400invalid_request_errorinvalid_json, model_requiredFix the request body; the message names the problem.
401authentication_errorinvalid_api_keyCheck the key or create a new one in the dashboard.
402insufficient_quotainsufficient_balanceTop up in Billing, then retry.
403permission_errormodel_not_allowed, ip_not_allowedAdjust the key's model or IP restrictions.
404not_found_errormodel_not_foundUse an ID from GET /v1/models.
413invalid_request_errorrequest_too_largeRequest bodies are limited to 8 MB.
429rate_limit_errorrate_limit_exceeded, spend_limit_exceededWait for Retry-After seconds or raise the key's limits.
502/503server_errorupstream_error, model_unavailableRetry with backoff; failed requests are not charged.

Your first request in under a minute.

Create a key, top up from $5 and swap the base URL. No subscription, no hidden fees.

Get your API keyRead the docs