GLM 5.3 Uncensored: Messages & Chat Completions
Configure both GLM 5.3 models on LLMsRelay: API keys, cash balance, cURL, Python, TypeScript, streaming, web search and tool calling.
LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility.
1. Account, wallet and key
- Sign in to LLMsRelay. If you do not have an account, register first, then return to Billing > Uncensored.
- Fund the separate GLM cash wallet: $45, $100, $500 or $1,000. Funding is 1:1 USD, with no signup, referral or promotional bonuses. Standard API-equivalent packs do not fund this wallet.
- Open API Keys, create a key and select GLM 5.3. Copy the secret once and keep it local. Use the full dashboard-issued sk- key, not a key from another service.
- Check both the cash balance and the key's remaining allowance. A positive standard balance is not sufficient. A key has one group; use separate keys for simultaneous Claude, GPT and GLM access.
2. Choose a public model ID
| Public model ID | Context (tokens) | Input |
|---|---|---|
| glm-5-3-uncensored | 262,144 | Text and images |
| glm-5-3-uncensored-1m | 1,000,000 | Text only |
curl https://api.llmsrelay.com/v1/models \
-H "Authorization: Bearer $LLMSRELAY_API_KEY"3. Store your credentials
Set LLMSRELAY_API_KEY in your local shell or secret manager. The examples use a placeholder; replace it locally, never in a shared chat, URL, browser bundle or repository. Do not add /v1 to the Anthropic SDK base URL; the OpenAI SDK needs /v1 exactly once.
export LLMSRELAY_API_KEY="YOUR_LLMSRELAY_KEY"4. Anthropic-compatible Messages
Messages uses x-api-key and anthropic-version. max_tokens is required. Read text blocks from response.content, not choices. The system instruction belongs in the top-level system field. To use the 1M model, replace only model with glm-5-3-uncensored-1m.
curl https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "Explain what an API gateway does."
}
]
}'5. OpenAI-compatible Chat Completions
Chat Completions uses Authorization: Bearer. Put system instructions in a system message and read choices[0].message.content. Use the exact public model ID. GET /v1/models lists the models accessible to a valid, funded key and may return 402 when no effective credit remains.
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"messages": [
{
"role": "system",
"content": "You are a concise assistant."
},
{
"role": "user",
"content": "Explain what an API gateway does."
}
]
}'6. Python and TypeScript SDKs
Install only the SDK you need. Keep the key in the server environment. The Anthropic SDK returns content blocks; the OpenAI SDK returns choices. These examples also work with the 1M public ID after changing model.
pip install anthropic openaiimport os
from anthropic import Anthropic
client = Anthropic(
api_key=os.environ["LLMSRELAY_API_KEY"],
base_url="https://api.llmsrelay.com",
)
response = client.messages.create(
model="glm-5-3-uncensored",
max_tokens=512,
messages=[{"role": "user", "content": "Hello!"}],
)
print("".join(block.text for block in response.content if block.type == "text"))import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LLMSRELAY_API_KEY"],
base_url="https://api.llmsrelay.com/v1",
)
response = client.chat.completions.create(
model="glm-5-3-uncensored",
max_tokens=512,
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)npm install @anthropic-ai/sdk openaiimport Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: process.env.LLMSRELAY_API_KEY,
baseURL: "https://api.llmsrelay.com",
});
const response = await client.messages.create({
model: "glm-5-3-uncensored",
max_tokens: 512,
messages: [{ role: "user", content: "Hello!" }],
});
for (const block of response.content) {
if (block.type === "text") console.log(block.text);
}import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.LLMSRELAY_API_KEY,
baseURL: "https://api.llmsrelay.com/v1",
});
const response = await client.chat.completions.create({
model: "glm-5-3-uncensored",
max_tokens: 512,
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);7. Claude Code and compatible IDEs
For Claude Code, use a dedicated GLM profile or shell with the variables below. Do not silently overwrite existing profiles. Remove conflicting ANTHROPIC_AUTH_TOKEN credentials in that profile. Map default and fast models to GLM so the client does not request Claude with a GLM key. For Cline, Roo Code or Continue, choose OpenAI Compatible, base URL https://api.llmsrelay.com/v1, your GLM key and one of the public IDs. Client tool compatibility varies; test a short request before a long session.
export ANTHROPIC_BASE_URL="https://api.llmsrelay.com"
export ANTHROPIC_API_KEY="$LLMSRELAY_API_KEY"
export ANTHROPIC_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_OPUS_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_SONNET_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="glm-5-3-uncensored"
export ANTHROPIC_SMALL_FAST_MODEL="glm-5-3-uncensored"
# Only in this dedicated GLM shell/profile:
unset ANTHROPIC_AUTH_TOKEN
claude8. Streaming (SSE)
Add stream:true to either request and use curl -N. Chat Completions emits data chunks: accumulate choices[].delta.content and tool_calls arguments until [DONE]. stream_options.include_usage:true requests a final usage chunk; its choices may be empty.
Messages emits named events including message_start, content_block_delta, message_delta and message_stop. Handle text, thinking and tool JSON deltas separately. Inspect error events even if HTTP status is 200; an opened stream is not proof of a completed response. Never concatenate raw SSE as plain JSON.
curl -N https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"messages": [
{
"role": "system",
"content": "You are a concise assistant."
},
{
"role": "user",
"content": "Explain what an API gateway does."
}
],
"stream": true,
"stream_options": {
"include_usage": true
}
}'curl -N https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "Explain what an API gateway does."
}
],
"stream": true
}'9. Built-in web search
Built-in search runs remotely; you do not implement a local web_search function. For Chat Completions, add web_search_options:{} to the request. For Messages, use the versioned web_search tool below. Both public models support these formats, including SSE.
Preserve returned citation metadata, URLs and server tool result blocks. A search request can produce several tool events before the answer. Verify that a search actually ran and inspect sources; a fluent answer alone is not evidence of a search.
Run custom function tools and built-in Chat web search as separate requests unless the combination has been explicitly validated for your client. Do not assume that accepting both parameters means both tools executed.
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 2048,
"messages": [
{
"role": "user",
"content": "Search the web for recent Python releases. Cite your sources."
}
],
"web_search_options": {}
}'curl https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 2048,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "Search the web for recent Python releases. Cite your sources."
}
],
"tools": [
{
"type": "web_search_2025_03_05",
"name": "web_search",
"max_uses": 3
}
]
}'10. Custom function tools
Define the function's JSON Schema and let the model request it. Your server validates arguments, performs the allowed operation, then sends its result back. Never execute arbitrary shell commands or untrusted arguments from a model.
Chat: append the assistant message containing tool_calls, then one role:tool message for each tool_call_id. Messages: append the complete assistant content, then a user message with tool_result blocks matching tool_use_id. Preserve reasoning/signature blocks when present. Continue until the model returns text; handle multiple calls and impose a loop limit.
The examples below show the follow-up shape, not executable weather services. Replace call IDs and full assistant payloads with the actual response and substitute a real, validated local result. In SSE, collect all argument fragments before parsing JSON.
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"messages": [
{
"role": "user",
"content": "What is the weather in Paris?"
}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string"
}
},
"required": [
"city"
]
}
}
}
]
}'# messages for the next Chat request; use the ACTUAL assistant message and call ID:
[
{"role":"user","content":"What is the weather in Paris?"},
{"role":"assistant","content":null,"tool_calls":[
{"id":"ACTUAL_CALL_ID","type":"function","function":{"name":"get_weather","arguments":"{\"city\":\"Paris\"}"}}
]},
{"role":"tool","tool_call_id":"ACTUAL_CALL_ID","content":"{\"temperature_c\":18}"}
]curl https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "What is the weather in Paris?"
}
],
"tools": [
{
"name": "get_weather",
"description": "Get current weather for a city",
"input_schema": {
"type": "object",
"properties": {
"city": {
"type": "string"
}
},
"required": [
"city"
]
}
}
]
}'# messages for the next Messages request; preserve the ACTUAL full assistant content:
[
{"role":"user","content":"What is the weather in Paris?"},
{"role":"assistant","content":[
{"type":"tool_use","id":"ACTUAL_TOOL_USE_ID","name":"get_weather","input":{"city":"Paris"}}
]},
{"role":"user","content":[
{"type":"tool_result","tool_use_id":"ACTUAL_TOOL_USE_ID","content":"{\"temperature_c\":18}"}
]}
]11. Reasoning and structured output
For Chat use reasoning_effort (for example high); for Messages use output_config.effort or thinking.budget_tokens. Hiding reasoning with include_reasoning:false does not eliminate reasoning cost. Budget enough output tokens for reasoning plus the answer. Chat JSON Schema is shown below; validate the returned JSON in your application. Do not assume every SDK accepts nonstandard fields: use extra_body in Python or raw HTTP.
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 2048,
"messages": [
{
"role": "user",
"content": "Return a short greeting as JSON."
}
],
"reasoning_effort": "high",
"include_reasoning": false,
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "greeting",
"strict": true,
"schema": {
"type": "object",
"properties": {
"greeting": {
"type": "string"
}
},
"required": [
"greeting"
],
"additionalProperties": false
}
}
}
}'12. Token counting, caching and images
Use POST /v1/messages/count_tokens with the same model and input before a long request. It estimates input tokens, not the eventual output or final charge. Keep the full input and output within the model's context window.
Keep stable instructions and tool schemas at the beginning of the prompt. Cache reuse is best-effort, not guaranteed on every repeat. Inspect usage for cache reads rather than assuming a hit. Cache creation is billed at the input rate; reasoning counts toward output.
Only glm-5-3-uncensored accepts images. For Chat use image_url content; for Messages use an image block with a supported URL or base64 source. The 1M model is text-only and rejects images with HTTP 400.
curl https://api.llmsrelay.com/v1/messages/count_tokens \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"messages": [
{
"role": "user",
"content": "Explain what an API gateway does."
}
]
}'13. Troubleshooting
| Status / symptom | Action |
|---|---|
| 401 | Check the full LLMsRelay secret, header and environment variable; rotate exposed keys. |
| 403 / model unavailable | Select GLM 5.3 and check the key's allowed models. A Basic or Codex key cannot access GLM. |
| 402 / insufficient_credits | Check the GLM cash wallet AND key allowance. Funds must cover the estimated reservation, including max_tokens; reduce the output budget or add funds. |
| 400 | Check model ID, JSON, max_tokens and tool schema. Remove images for the 1M model. |
| 404 / wrong route | Messages: /v1/messages. Chat: /v1/chat/completions. Remove duplicate /v1. |
| 429 | Respect Retry-After when present; use bounded exponential backoff and reduce concurrency. |
| 5xx / SSE error | Record the request ID, inspect the terminal stream event and contact support if repeated. A retry is a new request and may incur usage. |
| Wallet unavailable | Do not pay repeatedly or assume the balance is zero. Wait for balance verification or contact support. |
14. Responses compatibility notes
Prefer Messages or Chat Completions for this guide. Native Codex uses Responses; configuring an OpenAI-compatible IDE is not the same as configuring Codex CLI.
GLM Responses is stateless: previous_response_id is unsupported. Replay prior input/output and function_call_output instead.
In checks on October 9, 2026, base-model Responses web-search SSE failed, while non-streaming search completed. Prefer Chat or Messages for web search. The 1M Responses search returned source URLs but omitted final citation annotations.
These are compatibility observations, not a guarantee for every future request. Keep error handling and verify your exact client. Full-context load tests and all third-party clients are not covered.
15. Final USD token prices
USD per 1 million tokens, final customer rates. No additional multiplier is applied to this table. Actual input, output and cached tokens are charged to the separate GLM cash wallet; there are no bonuses in this flow.
| Model | Input | Output | Cache read | Cache write |
|---|---|---|---|---|
| glm-5-3-uncensored | $4.00 | $12.00 | $0.40 | $4.00 |
| glm-5-3-uncensored-1m | $12.00 | $20.00 | $1.20 | $12.00 |
Fund GLM wallet
Fund the separate GLM cash wallet: $45, $100, $500 or $1,000. Funding is 1:1 USD, with no signup, referral or promotional bonuses. Standard API-equivalent packs do not fund this wallet.
Fund GLM wallet