Back to dashboard

GLM 5.3 Uncensored: Messages & Chat Completions

Configure both GLM 5.3 models on LLMsRelay: API keys, cash balance, cURL, Python, TypeScript, streaming, web search and tool calling.

LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility.

1. Account, wallet and key

  1. Sign in to LLMsRelay. If you do not have an account, register first, then return to Billing > Uncensored.
  2. Fund the separate GLM cash wallet: $45, $100, $500 or $1,000. Funding is 1:1 USD, with no signup, referral or promotional bonuses. Standard API-equivalent packs do not fund this wallet.
  3. Open API Keys, create a key and select GLM 5.3. Copy the secret once and keep it local. Use the full dashboard-issued sk- key, not a key from another service.
  4. Check both the cash balance and the key's remaining allowance. A positive standard balance is not sufficient. A key has one group; use separate keys for simultaneous Claude, GPT and GLM access.

2. Choose a public model ID

Public model IDContext (tokens)Input
glm-5-3-uncensored262,144Text and images
glm-5-3-uncensored-1m1,000,000Text only
GET /v1/modelsbash
curl https://api.llmsrelay.com/v1/models \
  -H "Authorization: Bearer $LLMSRELAY_API_KEY"

3. Store your credentials

Set LLMSRELAY_API_KEY in your local shell or secret manager. The examples use a placeholder; replace it locally, never in a shared chat, URL, browser bundle or repository. Do not add /v1 to the Anthropic SDK base URL; the OpenAI SDK needs /v1 exactly once.

Environmentbash
export LLMSRELAY_API_KEY="YOUR_LLMSRELAY_KEY"

4. Anthropic-compatible Messages

Messages uses x-api-key and anthropic-version. max_tokens is required. Read text blocks from response.content, not choices. The system instruction belongs in the top-level system field. To use the 1M model, replace only model with glm-5-3-uncensored-1m.

POST /v1/messagesbash
curl https://api.llmsrelay.com/v1/messages \
  -H "x-api-key: $LLMSRELAY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "system": "You are a concise assistant.",
  "messages": [
    {
      "role": "user",
      "content": "Explain what an API gateway does."
    }
  ]
}'

5. OpenAI-compatible Chat Completions

Chat Completions uses Authorization: Bearer. Put system instructions in a system message and read choices[0].message.content. Use the exact public model ID. GET /v1/models lists the models accessible to a valid, funded key and may return 402 when no effective credit remains.

POST /v1/chat/completionsbash
curl https://api.llmsrelay.com/v1/chat/completions \
  -H "Authorization: Bearer $LLMSRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "messages": [
    {
      "role": "system",
      "content": "You are a concise assistant."
    },
    {
      "role": "user",
      "content": "Explain what an API gateway does."
    }
  ]
}'

6. Python and TypeScript SDKs

Install only the SDK you need. Keep the key in the server environment. The Anthropic SDK returns content blocks; the OpenAI SDK returns choices. These examples also work with the 1M public ID after changing model.

Pythonbash
pip install anthropic openai
Python: Messagespython
import os
from anthropic import Anthropic

client = Anthropic(
    api_key=os.environ["LLMSRELAY_API_KEY"],
    base_url="https://api.llmsrelay.com",
)
response = client.messages.create(
    model="glm-5-3-uncensored",
    max_tokens=512,
    messages=[{"role": "user", "content": "Hello!"}],
)
print("".join(block.text for block in response.content if block.type == "text"))
Python: Chat Completionspython
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["LLMSRELAY_API_KEY"],
    base_url="https://api.llmsrelay.com/v1",
)
response = client.chat.completions.create(
    model="glm-5-3-uncensored",
    max_tokens=512,
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
Node.jsbash
npm install @anthropic-ai/sdk openai
TypeScript: Messagestypescript
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  apiKey: process.env.LLMSRELAY_API_KEY,
  baseURL: "https://api.llmsrelay.com",
});
const response = await client.messages.create({
  model: "glm-5-3-uncensored",
  max_tokens: 512,
  messages: [{ role: "user", content: "Hello!" }],
});
for (const block of response.content) {
  if (block.type === "text") console.log(block.text);
}
TypeScript: Chat Completionstypescript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.LLMSRELAY_API_KEY,
  baseURL: "https://api.llmsrelay.com/v1",
});
const response = await client.chat.completions.create({
  model: "glm-5-3-uncensored",
  max_tokens: 512,
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);

7. Claude Code and compatible IDEs

For Claude Code, use a dedicated GLM profile or shell with the variables below. Do not silently overwrite existing profiles. Remove conflicting ANTHROPIC_AUTH_TOKEN credentials in that profile. Map default and fast models to GLM so the client does not request Claude with a GLM key. For Cline, Roo Code or Continue, choose OpenAI Compatible, base URL https://api.llmsrelay.com/v1, your GLM key and one of the public IDs. Client tool compatibility varies; test a short request before a long session.

Claude Code: GLMbash
export ANTHROPIC_BASE_URL="https://api.llmsrelay.com"
export ANTHROPIC_API_KEY="$LLMSRELAY_API_KEY"
export ANTHROPIC_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_OPUS_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_SONNET_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="glm-5-3-uncensored"
export ANTHROPIC_SMALL_FAST_MODEL="glm-5-3-uncensored"
# Only in this dedicated GLM shell/profile:
unset ANTHROPIC_AUTH_TOKEN
claude

8. Streaming (SSE)

Add stream:true to either request and use curl -N. Chat Completions emits data chunks: accumulate choices[].delta.content and tool_calls arguments until [DONE]. stream_options.include_usage:true requests a final usage chunk; its choices may be empty.

Messages emits named events including message_start, content_block_delta, message_delta and message_stop. Handle text, thinking and tool JSON deltas separately. Inspect error events even if HTTP status is 200; an opened stream is not proof of a completed response. Never concatenate raw SSE as plain JSON.

Chat Completions: SSEbash
curl -N https://api.llmsrelay.com/v1/chat/completions \
  -H "Authorization: Bearer $LLMSRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "messages": [
    {
      "role": "system",
      "content": "You are a concise assistant."
    },
    {
      "role": "user",
      "content": "Explain what an API gateway does."
    }
  ],
  "stream": true,
  "stream_options": {
    "include_usage": true
  }
}'
Messages: SSEbash
curl -N https://api.llmsrelay.com/v1/messages \
  -H "x-api-key: $LLMSRELAY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "system": "You are a concise assistant.",
  "messages": [
    {
      "role": "user",
      "content": "Explain what an API gateway does."
    }
  ],
  "stream": true
}'

10. Custom function tools

Define the function's JSON Schema and let the model request it. Your server validates arguments, performs the allowed operation, then sends its result back. Never execute arbitrary shell commands or untrusted arguments from a model.

Chat: append the assistant message containing tool_calls, then one role:tool message for each tool_call_id. Messages: append the complete assistant content, then a user message with tool_result blocks matching tool_use_id. Preserve reasoning/signature blocks when present. Continue until the model returns text; handle multiple calls and impose a loop limit.

The examples below show the follow-up shape, not executable weather services. Replace call IDs and full assistant payloads with the actual response and substitute a real, validated local result. In SSE, collect all argument fragments before parsing JSON.

Chat Completions: functionbash
curl https://api.llmsrelay.com/v1/chat/completions \
  -H "Authorization: Bearer $LLMSRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "messages": [
    {
      "role": "user",
      "content": "What is the weather in Paris?"
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string"
            }
          },
          "required": [
            "city"
          ]
        }
      }
    }
  ]
}'
Chat Completions: continuationtext
# messages for the next Chat request; use the ACTUAL assistant message and call ID:
[
  {"role":"user","content":"What is the weather in Paris?"},
  {"role":"assistant","content":null,"tool_calls":[
    {"id":"ACTUAL_CALL_ID","type":"function","function":{"name":"get_weather","arguments":"{\"city\":\"Paris\"}"}}
  ]},
  {"role":"tool","tool_call_id":"ACTUAL_CALL_ID","content":"{\"temperature_c\":18}"}
]
Messages: functionbash
curl https://api.llmsrelay.com/v1/messages \
  -H "x-api-key: $LLMSRELAY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "system": "You are a concise assistant.",
  "messages": [
    {
      "role": "user",
      "content": "What is the weather in Paris?"
    }
  ],
  "tools": [
    {
      "name": "get_weather",
      "description": "Get current weather for a city",
      "input_schema": {
        "type": "object",
        "properties": {
          "city": {
            "type": "string"
          }
        },
        "required": [
          "city"
        ]
      }
    }
  ]
}'
Messages: continuationtext
# messages for the next Messages request; preserve the ACTUAL full assistant content:
[
  {"role":"user","content":"What is the weather in Paris?"},
  {"role":"assistant","content":[
    {"type":"tool_use","id":"ACTUAL_TOOL_USE_ID","name":"get_weather","input":{"city":"Paris"}}
  ]},
  {"role":"user","content":[
    {"type":"tool_result","tool_use_id":"ACTUAL_TOOL_USE_ID","content":"{\"temperature_c\":18}"}
  ]}
]

11. Reasoning and structured output

For Chat use reasoning_effort (for example high); for Messages use output_config.effort or thinking.budget_tokens. Hiding reasoning with include_reasoning:false does not eliminate reasoning cost. Budget enough output tokens for reasoning plus the answer. Chat JSON Schema is shown below; validate the returned JSON in your application. Do not assume every SDK accepts nonstandard fields: use extra_body in Python or raw HTTP.

Chat Completions: JSON Schemabash
curl https://api.llmsrelay.com/v1/chat/completions \
  -H "Authorization: Bearer $LLMSRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 2048,
  "messages": [
    {
      "role": "user",
      "content": "Return a short greeting as JSON."
    }
  ],
  "reasoning_effort": "high",
  "include_reasoning": false,
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "greeting",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": {
          "greeting": {
            "type": "string"
          }
        },
        "required": [
          "greeting"
        ],
        "additionalProperties": false
      }
    }
  }
}'

12. Token counting, caching and images

Use POST /v1/messages/count_tokens with the same model and input before a long request. It estimates input tokens, not the eventual output or final charge. Keep the full input and output within the model's context window.

Keep stable instructions and tool schemas at the beginning of the prompt. Cache reuse is best-effort, not guaranteed on every repeat. Inspect usage for cache reads rather than assuming a hit. Cache creation is billed at the input rate; reasoning counts toward output.

Only glm-5-3-uncensored accepts images. For Chat use image_url content; for Messages use an image block with a supported URL or base64 source. The 1M model is text-only and rejects images with HTTP 400.

POST /v1/messages/count_tokensbash
curl https://api.llmsrelay.com/v1/messages/count_tokens \
  -H "x-api-key: $LLMSRELAY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "messages": [
    {
      "role": "user",
      "content": "Explain what an API gateway does."
    }
  ]
}'

13. Troubleshooting

Status / symptomAction
401Check the full LLMsRelay secret, header and environment variable; rotate exposed keys.
403 / model unavailableSelect GLM 5.3 and check the key's allowed models. A Basic or Codex key cannot access GLM.
402 / insufficient_creditsCheck the GLM cash wallet AND key allowance. Funds must cover the estimated reservation, including max_tokens; reduce the output budget or add funds.
400Check model ID, JSON, max_tokens and tool schema. Remove images for the 1M model.
404 / wrong routeMessages: /v1/messages. Chat: /v1/chat/completions. Remove duplicate /v1.
429Respect Retry-After when present; use bounded exponential backoff and reduce concurrency.
5xx / SSE errorRecord the request ID, inspect the terminal stream event and contact support if repeated. A retry is a new request and may incur usage.
Wallet unavailableDo not pay repeatedly or assume the balance is zero. Wait for balance verification or contact support.

14. Responses compatibility notes

Prefer Messages or Chat Completions for this guide. Native Codex uses Responses; configuring an OpenAI-compatible IDE is not the same as configuring Codex CLI.

GLM Responses is stateless: previous_response_id is unsupported. Replay prior input/output and function_call_output instead.

In checks on October 9, 2026, base-model Responses web-search SSE failed, while non-streaming search completed. Prefer Chat or Messages for web search. The 1M Responses search returned source URLs but omitted final citation annotations.

These are compatibility observations, not a guarantee for every future request. Keep error handling and verify your exact client. Full-context load tests and all third-party clients are not covered.

15. Final USD token prices

USD per 1 million tokens, final customer rates. No additional multiplier is applied to this table. Actual input, output and cached tokens are charged to the separate GLM cash wallet; there are no bonuses in this flow.

ModelInputOutputCache readCache write
glm-5-3-uncensored$4.00$12.00$0.40$4.00
glm-5-3-uncensored-1m$12.00$20.00$1.20$12.00

Fund GLM wallet

Fund the separate GLM cash wallet: $45, $100, $500 or $1,000. Funding is 1:1 USD, with no signup, referral or promotional bonuses. Standard API-equivalent packs do not fund this wallet.

Fund GLM wallet