返回控制台

GLM 5.3 Uncensored:Messages 与 Chat Completions

通过 LLMsRelay 配置两款 GLM 5.3:密钥、现金余额、cURL、Python、TypeScript、流式输出、网络搜索与工具调用。

LLMsRelay 是独立运营的 API 网关,与 Anthropic, PBC 无隶属或授权关系。产品和模型名称仅用于说明兼容性。

1. 账户、钱包和密钥

  1. 登录 LLMsRelay。如无账户,请先注册,再返回 Billing > Uncensored。
  2. 为独立的 GLM 现金钱包充值 $45、$100、$500 或 $1,000。按美元 1:1 入账,无注册、推荐或促销奖励。标准 API 等价用量包不会充值此钱包。
  3. 进入 API Keys,创建密钥并选择 GLM 5.3 分组。创建时复制完整密钥并保存在本地。使用本平台签发的完整 sk- 密钥,不要使用其他服务的密钥。
  4. 同时检查现金余额与密钥剩余额度。标准余额为正并不足够。每个密钥只有一个分组;同时使用 Claude、GPT 和 GLM 时请创建独立密钥。

2. 选择公开模型 ID

公开模型 ID上下文(Token)输入
glm-5-3-uncensored262,144文本与图片
glm-5-3-uncensored-1m1,000,000仅文本
GET /v1/modelsbash
curl https://api.llmsrelay.com/v1/models \
  -H "Authorization: Bearer $LLMSRELAY_API_KEY"

3. 安全保存密钥

在本地 shell 或密钥管理器中设置 LLMSRELAY_API_KEY。示例使用占位符,请仅在本地替换,不要放入共享聊天、URL、浏览器代码或仓库。Anthropic SDK 的 base URL 不加 /v1,OpenAI SDK 则必须且只能包含一次 /v1。

Environmentbash
export LLMSRELAY_API_KEY="YOUR_LLMSRELAY_KEY"

4. Anthropic 兼容 Messages

Messages 使用 x-api-key 与 anthropic-version。必须提供 max_tokens。文本在 response.content 的内容块中,不在 choices 中。系统提示使用顶层 system 字段。使用 1M 模型时只需将 model 改为 glm-5-3-uncensored-1m。

POST /v1/messagesbash
curl https://api.llmsrelay.com/v1/messages \
  -H "x-api-key: $LLMSRELAY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "system": "You are a concise assistant.",
  "messages": [
    {
      "role": "user",
      "content": "Explain what an API gateway does."
    }
  ]
}'

5. OpenAI 兼容 Chat Completions

Chat Completions 使用 Authorization: Bearer。系统提示放入 system 消息,答案读取 choices[0].message.content。请使用准确的公开模型 ID。有效且有余额的密钥可通过 GET /v1/models 查看模型;有效可用额度为零时可能返回 402。

POST /v1/chat/completionsbash
curl https://api.llmsrelay.com/v1/chat/completions \
  -H "Authorization: Bearer $LLMSRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "messages": [
    {
      "role": "system",
      "content": "You are a concise assistant."
    },
    {
      "role": "user",
      "content": "Explain what an API gateway does."
    }
  ]
}'

6. Python 与 TypeScript SDK

只安装所需 SDK,将密钥保存在服务器环境。Anthropic SDK 返回 content 块,OpenAI SDK 返回 choices。将示例中的 model 改为 1M 公开 ID 即可使用长上下文模型。

Pythonbash
pip install anthropic openai
Python: Messagespython
import os
from anthropic import Anthropic

client = Anthropic(
    api_key=os.environ["LLMSRELAY_API_KEY"],
    base_url="https://api.llmsrelay.com",
)
response = client.messages.create(
    model="glm-5-3-uncensored",
    max_tokens=512,
    messages=[{"role": "user", "content": "Hello!"}],
)
print("".join(block.text for block in response.content if block.type == "text"))
Python: Chat Completionspython
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["LLMSRELAY_API_KEY"],
    base_url="https://api.llmsrelay.com/v1",
)
response = client.chat.completions.create(
    model="glm-5-3-uncensored",
    max_tokens=512,
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
Node.jsbash
npm install @anthropic-ai/sdk openai
TypeScript: Messagestypescript
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  apiKey: process.env.LLMSRELAY_API_KEY,
  baseURL: "https://api.llmsrelay.com",
});
const response = await client.messages.create({
  model: "glm-5-3-uncensored",
  max_tokens: 512,
  messages: [{ role: "user", content: "Hello!" }],
});
for (const block of response.content) {
  if (block.type === "text") console.log(block.text);
}
TypeScript: Chat Completionstypescript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.LLMSRELAY_API_KEY,
  baseURL: "https://api.llmsrelay.com/v1",
});
const response = await client.chat.completions.create({
  model: "glm-5-3-uncensored",
  max_tokens: 512,
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);

7. Claude Code 与兼容 IDE

Claude Code 请使用独立 GLM 配置或 shell,设置下方变量,不要直接覆盖已有配置。在该配置中移除冲突的 ANTHROPIC_AUTH_TOKEN。将默认和快速模型都映射到 GLM,避免用 GLM 密钥请求 Claude。在 Cline、Roo Code 或 Continue 中选择 OpenAI Compatible,base URL 为 https://api.llmsrelay.com/v1,并填写 GLM 密钥和公开 ID。工具兼容性因客户端而异,长会话前先做短请求测试。

Claude Code: GLMbash
export ANTHROPIC_BASE_URL="https://api.llmsrelay.com"
export ANTHROPIC_API_KEY="$LLMSRELAY_API_KEY"
export ANTHROPIC_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_OPUS_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_SONNET_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="glm-5-3-uncensored"
export ANTHROPIC_SMALL_FAST_MODEL="glm-5-3-uncensored"
# Only in this dedicated GLM shell/profile:
unset ANTHROPIC_AUTH_TOKEN
claude

8. 流式输出(SSE)

在任一请求中添加 stream:true,并使用 curl -N。Chat Completions 需累积 choices[].delta.content 和 tool_calls 参数片段,直到 [DONE]。stream_options.include_usage:true 请求最终 usage 块,其中 choices 可能为空。

Messages 使用 message_start、content_block_delta、message_delta、message_stop 等命名事件。分别处理文本、thinking 和工具 JSON 增量。即使 HTTP 为 200,也要检查 error 事件;打开流不代表成功完成。不要把原始 SSE 当作普通 JSON 拼接。

Chat Completions: SSEbash
curl -N https://api.llmsrelay.com/v1/chat/completions \
  -H "Authorization: Bearer $LLMSRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "messages": [
    {
      "role": "system",
      "content": "You are a concise assistant."
    },
    {
      "role": "user",
      "content": "Explain what an API gateway does."
    }
  ],
  "stream": true,
  "stream_options": {
    "include_usage": true
  }
}'
Messages: SSEbash
curl -N https://api.llmsrelay.com/v1/messages \
  -H "x-api-key: $LLMSRELAY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "system": "You are a concise assistant.",
  "messages": [
    {
      "role": "user",
      "content": "Explain what an API gateway does."
    }
  ],
  "stream": true
}'

10. 自定义函数工具

用 JSON Schema 定义函数,由模型发起调用。服务器验证参数、执行允许的操作,再将结果传回。不要执行模型生成的任意 shell 命令或未经验证的参数。

Chat:追加包含 tool_calls 的完整 assistant 消息,再为每个 tool_call_id 追加 role:tool。Messages:追加完整 assistant content,再用 user 消息发送匹配 tool_use_id 的 tool_result。保留存在的推理与签名块。循环到模型返回文本,支持多工具并限制迭代次数。

下方示例仅展示续接结构,不是可运行的天气服务。请替换为真实调用 ID、完整 assistant 响应和经过验证的本地执行结果。SSE 参数必须收集完整后再解析 JSON。

Chat Completions: functionbash
curl https://api.llmsrelay.com/v1/chat/completions \
  -H "Authorization: Bearer $LLMSRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "messages": [
    {
      "role": "user",
      "content": "What is the weather in Paris?"
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string"
            }
          },
          "required": [
            "city"
          ]
        }
      }
    }
  ]
}'
Chat Completions: continuationtext
# messages for the next Chat request; use the ACTUAL assistant message and call ID:
[
  {"role":"user","content":"What is the weather in Paris?"},
  {"role":"assistant","content":null,"tool_calls":[
    {"id":"ACTUAL_CALL_ID","type":"function","function":{"name":"get_weather","arguments":"{\"city\":\"Paris\"}"}}
  ]},
  {"role":"tool","tool_call_id":"ACTUAL_CALL_ID","content":"{\"temperature_c\":18}"}
]
Messages: functionbash
curl https://api.llmsrelay.com/v1/messages \
  -H "x-api-key: $LLMSRELAY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 512,
  "system": "You are a concise assistant.",
  "messages": [
    {
      "role": "user",
      "content": "What is the weather in Paris?"
    }
  ],
  "tools": [
    {
      "name": "get_weather",
      "description": "Get current weather for a city",
      "input_schema": {
        "type": "object",
        "properties": {
          "city": {
            "type": "string"
          }
        },
        "required": [
          "city"
        ]
      }
    }
  ]
}'
Messages: continuationtext
# messages for the next Messages request; preserve the ACTUAL full assistant content:
[
  {"role":"user","content":"What is the weather in Paris?"},
  {"role":"assistant","content":[
    {"type":"tool_use","id":"ACTUAL_TOOL_USE_ID","name":"get_weather","input":{"city":"Paris"}}
  ]},
  {"role":"user","content":[
    {"type":"tool_result","tool_use_id":"ACTUAL_TOOL_USE_ID","content":"{\"temperature_c\":18}"}
  ]}
]

11. 推理与结构化输出

Chat 使用 reasoning_effort(如 high),Messages 使用 output_config.effort 或 thinking.budget_tokens。include_reasoning:false 只隐藏推理,不免除其费用。输出预算需覆盖推理与答案。下方是 Chat JSON Schema 示例,应用仍须验证返回的 JSON。非标准 SDK 参数请通过 Python extra_body 或直接 HTTP 传入。

Chat Completions: JSON Schemabash
curl https://api.llmsrelay.com/v1/chat/completions \
  -H "Authorization: Bearer $LLMSRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "max_tokens": 2048,
  "messages": [
    {
      "role": "user",
      "content": "Return a short greeting as JSON."
    }
  ],
  "reasoning_effort": "high",
  "include_reasoning": false,
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "greeting",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": {
          "greeting": {
            "type": "string"
          }
        },
        "required": [
          "greeting"
        ],
        "additionalProperties": false
      }
    }
  }
}'

12. Token 计数、缓存与图片

长请求前可用相同模型与输入调用 POST /v1/messages/count_tokens。它估算输入 Token,而非最终输出或总费用。输入与输出总量应在模型上下文范围内。

将稳定的指令和工具 Schema 放在前缀。缓存复用是尽力而为,并非每次重复都命中。检查 usage 中的缓存读取,不要自行假定。缓存写入按输入计费,推理计入输出。

仅 glm-5-3-uncensored 接受图片。Chat 使用 image_url,Messages 使用带支持 URL 或 base64 来源的 image 块。1M 为纯文本模型,图片输入返回 HTTP 400。

POST /v1/messages/count_tokensbash
curl https://api.llmsrelay.com/v1/messages/count_tokens \
  -H "x-api-key: $LLMSRELAY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "glm-5-3-uncensored",
  "messages": [
    {
      "role": "user",
      "content": "Explain what an API gateway does."
    }
  ]
}'

13. 故障排查

状态 / 现象处理方式
401检查完整 LLMsRelay 密钥、请求头和环境变量;泄露密钥应轮换。
403 / 模型不可用选择 GLM 5.3,检查密钥允许的模型。Basic 或 Codex 密钥不能访问 GLM。
402 / insufficient_credits同时检查 GLM 现金钱包与密钥额度。余额须覆盖包含 max_tokens 的预留;降低输出预算或充值。
400检查模型 ID、JSON、max_tokens、工具 Schema;1M 请求需移除图片。
404 / 路径错误Messages 使用 /v1/messages,Chat 使用 /v1/chat/completions。移除重复 /v1。
429遵守 Retry-After,使用有限指数退避并降低并发。
5xx / SSE 错误记录 request ID,检查终止事件;反复发生请联系支持。重试为新请求,可能单独计费。
钱包不可用不要重复付款,也不要推断余额为零。等待余额验证或联系支持。

14. Responses 兼容性说明

本指南优先使用 Messages 或 Chat Completions。原生 Codex 使用 Responses,配置 OpenAI 兼容 IDE 不等于配置 Codex CLI。

GLM Responses 为无状态接口,不支持 previous_response_id。请显式重传先前 input/output 和 function_call_output。

2026 年 10 月 9 日测试中,基础模型 Responses 搜索 SSE 失败,非流式搜索成功。搜索优先使用 Chat 或 Messages。1M Responses 搜索返回来源 URL,但缺少最终引用注释。

这些是兼容性观察,不是对所有未来请求的保证。仍需处理错误并验证实际客户端。未覆盖满上下文压力测试与所有第三方客户端。

15. 最终美元 Token 价格

每百万 Token 的美元最终客户价格,表格无需再乘任何倍率。实际输入、输出与缓存 Token 从独立 GLM 现金钱包扣除,此流程没有奖励。

模型输入输出缓存读取缓存写入
glm-5-3-uncensored$4.00$12.00$0.40$4.00
glm-5-3-uncensored-1m$12.00$20.00$1.20$12.00

充值 GLM 钱包

为独立的 GLM 现金钱包充值 $45、$100、$500 或 $1,000。按美元 1:1 入账,无注册、推荐或促销奖励。标准 API 等价用量包不会充值此钱包。

充值 GLM 钱包