glm-5.3-flash API
A prepaid Model.sale endpoint for glm-5.3-flash: one API key, explicit limits and a live compatibility check. The requested model is never silently replaced.
Quick start
Create an API key, add balance and use the base URL below. Keep the key in an environment variable.
curl https://api.model.sale/v1/chat/completions \
-H "Authorization: Bearer $MODEL_SALE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"glm-5.3-flash","messages":[{"role":"user","content":"Reply exactly OK"}]}'import os
from openai import OpenAI
client = OpenAI(base_url="https://api.model.sale/v1", api_key=os.environ["MODEL_SALE_API_KEY"])
response = client.chat.completions.create(model="glm-5.3-flash", messages=[{"role": "user", "content": "Reply exactly OK"}])
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.model.sale/v1", apiKey: process.env.MODEL_SALE_API_KEY });
const response = await client.chat.completions.create({ model: "glm-5.3-flash", messages: [{ role: "user", content: "Reply exactly OK" }] });
console.log(response.choices[0]?.message.content);The wallet reserves before dispatch and settles from terminal usage. Cached and reasoning token treatment follows the current billing policy.
How to estimate a request
The public rate is a blended USD amount per one million charged tokens. It is not a separate input-token or output-token tariff. For a quick estimate, multiply the charged-token count by the rate and divide by 1,000,000. The request record is authoritative for actual usage and settlement.
| Charged tokens | Illustrative charge |
|---|---|
| 1,000 | $0.0001 |
| 10,000 | $0.001 |
| 100,000 | $0.01 |
These are arithmetic examples, not a promise of a fixed per-request price. Actual token counts depend on your input, output and the model's reported usage. Any minimum request amount shown above is an initial reservation; after terminal usage is received, the wallet settles the measured charge and releases unused reserved funds.
Frequently asked questions
How much does the glm-5.3-flash API cost?
$0.10 per 1M charged tokens on Model.sale, billed from a prepaid USD balance (minimum top-up $5). Cached input tokens are charged at 40% of that price, and every request shows its exact charge in the dashboard.
Which endpoints does glm-5.3-flash support?
Chat Completions. Support is verified per model, and streaming is tested separately for each route.
How do I call glm-5.3-flash?
Create an API key, then send requests to https://api.model.sale/v1 with the model ID "glm-5.3-flash" using the OpenAI SDK, curl or any HTTP client. The curl, Python and TypeScript examples are above.
Does Model.sale replace the model I request?
No. The requested model ID is preserved. If a model is unavailable, the request fails with an error instead of being switched to a different model.
Before using this model in production
- Confirm the exact model ID and one of the verified endpoints listed above.
- Start with a dedicated API key and a low single-request and daily spend cap.
- Send one short JSON request, then verify the completed request and charge in Usage.
- If your client depends on streaming, test an SSE request and confirm that it reaches a terminal event.
- Handle 401, 429, timeout and 5xx responses explicitly. A failed request is not automatically retried or switched to a different model by Model.sale.
Availability is a recent observation, not an uptime guarantee. Check the full model catalog, pricing table, client setup guides and billing methodology before choosing a production configuration.