GLM 5.3 Uncensored:Messages 与 Chat Completions
通过 LLMsRelay 配置两款 GLM 5.3:密钥、现金余额、cURL、Python、TypeScript、流式输出、网络搜索与工具调用。
LLMsRelay 是独立运营的 API 网关,与 Anthropic, PBC 无隶属或授权关系。产品和模型名称仅用于说明兼容性。
1. 账户、钱包和密钥
- 登录 LLMsRelay。如无账户,请先注册,再返回 Billing > Uncensored。
- 为独立的 GLM 现金钱包充值 $45、$100、$500 或 $1,000。按美元 1:1 入账,无注册、推荐或促销奖励。标准 API 等价用量包不会充值此钱包。
- 进入 API Keys,创建密钥并选择 GLM 5.3 分组。创建时复制完整密钥并保存在本地。使用本平台签发的完整 sk- 密钥,不要使用其他服务的密钥。
- 同时检查现金余额与密钥剩余额度。标准余额为正并不足够。每个密钥只有一个分组;同时使用 Claude、GPT 和 GLM 时请创建独立密钥。
2. 选择公开模型 ID
| 公开模型 ID | 上下文(Token) | 输入 |
|---|---|---|
| glm-5-3-uncensored | 262,144 | 文本与图片 |
| glm-5-3-uncensored-1m | 1,000,000 | 仅文本 |
curl https://api.llmsrelay.com/v1/models \
-H "Authorization: Bearer $LLMSRELAY_API_KEY"3. 安全保存密钥
在本地 shell 或密钥管理器中设置 LLMSRELAY_API_KEY。示例使用占位符,请仅在本地替换,不要放入共享聊天、URL、浏览器代码或仓库。Anthropic SDK 的 base URL 不加 /v1,OpenAI SDK 则必须且只能包含一次 /v1。
export LLMSRELAY_API_KEY="YOUR_LLMSRELAY_KEY"4. Anthropic 兼容 Messages
Messages 使用 x-api-key 与 anthropic-version。必须提供 max_tokens。文本在 response.content 的内容块中,不在 choices 中。系统提示使用顶层 system 字段。使用 1M 模型时只需将 model 改为 glm-5-3-uncensored-1m。
curl https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "Explain what an API gateway does."
}
]
}'5. OpenAI 兼容 Chat Completions
Chat Completions 使用 Authorization: Bearer。系统提示放入 system 消息,答案读取 choices[0].message.content。请使用准确的公开模型 ID。有效且有余额的密钥可通过 GET /v1/models 查看模型;有效可用额度为零时可能返回 402。
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"messages": [
{
"role": "system",
"content": "You are a concise assistant."
},
{
"role": "user",
"content": "Explain what an API gateway does."
}
]
}'6. Python 与 TypeScript SDK
只安装所需 SDK,将密钥保存在服务器环境。Anthropic SDK 返回 content 块,OpenAI SDK 返回 choices。将示例中的 model 改为 1M 公开 ID 即可使用长上下文模型。
pip install anthropic openaiimport os
from anthropic import Anthropic
client = Anthropic(
api_key=os.environ["LLMSRELAY_API_KEY"],
base_url="https://api.llmsrelay.com",
)
response = client.messages.create(
model="glm-5-3-uncensored",
max_tokens=512,
messages=[{"role": "user", "content": "Hello!"}],
)
print("".join(block.text for block in response.content if block.type == "text"))import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LLMSRELAY_API_KEY"],
base_url="https://api.llmsrelay.com/v1",
)
response = client.chat.completions.create(
model="glm-5-3-uncensored",
max_tokens=512,
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)npm install @anthropic-ai/sdk openaiimport Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: process.env.LLMSRELAY_API_KEY,
baseURL: "https://api.llmsrelay.com",
});
const response = await client.messages.create({
model: "glm-5-3-uncensored",
max_tokens: 512,
messages: [{ role: "user", content: "Hello!" }],
});
for (const block of response.content) {
if (block.type === "text") console.log(block.text);
}import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.LLMSRELAY_API_KEY,
baseURL: "https://api.llmsrelay.com/v1",
});
const response = await client.chat.completions.create({
model: "glm-5-3-uncensored",
max_tokens: 512,
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);7. Claude Code 与兼容 IDE
Claude Code 请使用独立 GLM 配置或 shell,设置下方变量,不要直接覆盖已有配置。在该配置中移除冲突的 ANTHROPIC_AUTH_TOKEN。将默认和快速模型都映射到 GLM,避免用 GLM 密钥请求 Claude。在 Cline、Roo Code 或 Continue 中选择 OpenAI Compatible,base URL 为 https://api.llmsrelay.com/v1,并填写 GLM 密钥和公开 ID。工具兼容性因客户端而异,长会话前先做短请求测试。
export ANTHROPIC_BASE_URL="https://api.llmsrelay.com"
export ANTHROPIC_API_KEY="$LLMSRELAY_API_KEY"
export ANTHROPIC_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_OPUS_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_SONNET_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="glm-5-3-uncensored"
export ANTHROPIC_SMALL_FAST_MODEL="glm-5-3-uncensored"
# Only in this dedicated GLM shell/profile:
unset ANTHROPIC_AUTH_TOKEN
claude8. 流式输出(SSE)
在任一请求中添加 stream:true,并使用 curl -N。Chat Completions 需累积 choices[].delta.content 和 tool_calls 参数片段,直到 [DONE]。stream_options.include_usage:true 请求最终 usage 块,其中 choices 可能为空。
Messages 使用 message_start、content_block_delta、message_delta、message_stop 等命名事件。分别处理文本、thinking 和工具 JSON 增量。即使 HTTP 为 200,也要检查 error 事件;打开流不代表成功完成。不要把原始 SSE 当作普通 JSON 拼接。
curl -N https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"messages": [
{
"role": "system",
"content": "You are a concise assistant."
},
{
"role": "user",
"content": "Explain what an API gateway does."
}
],
"stream": true,
"stream_options": {
"include_usage": true
}
}'curl -N https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "Explain what an API gateway does."
}
],
"stream": true
}'9. 内置网络搜索
内置搜索在远端执行,无需实现本地 web_search 函数。Chat Completions 添加 web_search_options:{},Messages 使用下方带版本的 web_search 工具。两款模型均支持这两种格式与 SSE。
保留引用元数据、URL 和服务器工具结果块。答案前可能出现多次搜索事件。确认工具实际执行并检查来源,流畅的回答本身不能证明搜索发生。
除非已针对客户端验证组合行为,否则将自定义函数与 Chat 内置搜索分开请求。接受两种参数不代表两种工具都执行。
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 2048,
"messages": [
{
"role": "user",
"content": "Search the web for recent Python releases. Cite your sources."
}
],
"web_search_options": {}
}'curl https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 2048,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "Search the web for recent Python releases. Cite your sources."
}
],
"tools": [
{
"type": "web_search_2025_03_05",
"name": "web_search",
"max_uses": 3
}
]
}'10. 自定义函数工具
用 JSON Schema 定义函数,由模型发起调用。服务器验证参数、执行允许的操作,再将结果传回。不要执行模型生成的任意 shell 命令或未经验证的参数。
Chat:追加包含 tool_calls 的完整 assistant 消息,再为每个 tool_call_id 追加 role:tool。Messages:追加完整 assistant content,再用 user 消息发送匹配 tool_use_id 的 tool_result。保留存在的推理与签名块。循环到模型返回文本,支持多工具并限制迭代次数。
下方示例仅展示续接结构,不是可运行的天气服务。请替换为真实调用 ID、完整 assistant 响应和经过验证的本地执行结果。SSE 参数必须收集完整后再解析 JSON。
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"messages": [
{
"role": "user",
"content": "What is the weather in Paris?"
}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string"
}
},
"required": [
"city"
]
}
}
}
]
}'# messages for the next Chat request; use the ACTUAL assistant message and call ID:
[
{"role":"user","content":"What is the weather in Paris?"},
{"role":"assistant","content":null,"tool_calls":[
{"id":"ACTUAL_CALL_ID","type":"function","function":{"name":"get_weather","arguments":"{\"city\":\"Paris\"}"}}
]},
{"role":"tool","tool_call_id":"ACTUAL_CALL_ID","content":"{\"temperature_c\":18}"}
]curl https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "What is the weather in Paris?"
}
],
"tools": [
{
"name": "get_weather",
"description": "Get current weather for a city",
"input_schema": {
"type": "object",
"properties": {
"city": {
"type": "string"
}
},
"required": [
"city"
]
}
}
]
}'# messages for the next Messages request; preserve the ACTUAL full assistant content:
[
{"role":"user","content":"What is the weather in Paris?"},
{"role":"assistant","content":[
{"type":"tool_use","id":"ACTUAL_TOOL_USE_ID","name":"get_weather","input":{"city":"Paris"}}
]},
{"role":"user","content":[
{"type":"tool_result","tool_use_id":"ACTUAL_TOOL_USE_ID","content":"{\"temperature_c\":18}"}
]}
]11. 推理与结构化输出
Chat 使用 reasoning_effort(如 high),Messages 使用 output_config.effort 或 thinking.budget_tokens。include_reasoning:false 只隐藏推理,不免除其费用。输出预算需覆盖推理与答案。下方是 Chat JSON Schema 示例,应用仍须验证返回的 JSON。非标准 SDK 参数请通过 Python extra_body 或直接 HTTP 传入。
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 2048,
"messages": [
{
"role": "user",
"content": "Return a short greeting as JSON."
}
],
"reasoning_effort": "high",
"include_reasoning": false,
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "greeting",
"strict": true,
"schema": {
"type": "object",
"properties": {
"greeting": {
"type": "string"
}
},
"required": [
"greeting"
],
"additionalProperties": false
}
}
}
}'12. Token 计数、缓存与图片
长请求前可用相同模型与输入调用 POST /v1/messages/count_tokens。它估算输入 Token,而非最终输出或总费用。输入与输出总量应在模型上下文范围内。
将稳定的指令和工具 Schema 放在前缀。缓存复用是尽力而为,并非每次重复都命中。检查 usage 中的缓存读取,不要自行假定。缓存写入按输入计费,推理计入输出。
仅 glm-5-3-uncensored 接受图片。Chat 使用 image_url,Messages 使用带支持 URL 或 base64 来源的 image 块。1M 为纯文本模型,图片输入返回 HTTP 400。
curl https://api.llmsrelay.com/v1/messages/count_tokens \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"messages": [
{
"role": "user",
"content": "Explain what an API gateway does."
}
]
}'13. 故障排查
| 状态 / 现象 | 处理方式 |
|---|---|
| 401 | 检查完整 LLMsRelay 密钥、请求头和环境变量;泄露密钥应轮换。 |
| 403 / 模型不可用 | 选择 GLM 5.3,检查密钥允许的模型。Basic 或 Codex 密钥不能访问 GLM。 |
| 402 / insufficient_credits | 同时检查 GLM 现金钱包与密钥额度。余额须覆盖包含 max_tokens 的预留;降低输出预算或充值。 |
| 400 | 检查模型 ID、JSON、max_tokens、工具 Schema;1M 请求需移除图片。 |
| 404 / 路径错误 | Messages 使用 /v1/messages,Chat 使用 /v1/chat/completions。移除重复 /v1。 |
| 429 | 遵守 Retry-After,使用有限指数退避并降低并发。 |
| 5xx / SSE 错误 | 记录 request ID,检查终止事件;反复发生请联系支持。重试为新请求,可能单独计费。 |
| 钱包不可用 | 不要重复付款,也不要推断余额为零。等待余额验证或联系支持。 |
14. Responses 兼容性说明
本指南优先使用 Messages 或 Chat Completions。原生 Codex 使用 Responses,配置 OpenAI 兼容 IDE 不等于配置 Codex CLI。
GLM Responses 为无状态接口,不支持 previous_response_id。请显式重传先前 input/output 和 function_call_output。
2026 年 10 月 9 日测试中,基础模型 Responses 搜索 SSE 失败,非流式搜索成功。搜索优先使用 Chat 或 Messages。1M Responses 搜索返回来源 URL,但缺少最终引用注释。
这些是兼容性观察,不是对所有未来请求的保证。仍需处理错误并验证实际客户端。未覆盖满上下文压力测试与所有第三方客户端。
15. 最终美元 Token 价格
每百万 Token 的美元最终客户价格,表格无需再乘任何倍率。实际输入、输出与缓存 Token 从独立 GLM 现金钱包扣除,此流程没有奖励。
| 模型 | 输入 | 输出 | 缓存读取 | 缓存写入 |
|---|---|---|---|---|
| glm-5-3-uncensored | $4.00 | $12.00 | $0.40 | $4.00 |
| glm-5-3-uncensored-1m | $12.00 | $20.00 | $1.20 | $12.00 |
充值 GLM 钱包
为独立的 GLM 现金钱包充值 $45、$100、$500 或 $1,000。按美元 1:1 入账,无注册、推荐或促销奖励。标准 API 等价用量包不会充值此钱包。
充值 GLM 钱包