GLM 5.3 Uncensored: Messages 및 Chat Completions
LLMsRelay의 두 GLM 5.3 모델 설정: API 키, 현금 잔액, cURL, Python, TypeScript, 스트리밍, 웹 검색 및 도구 호출.
LLMsRelay는 독립적으로 운영되는 API 게이트웨이이며 Anthropic, PBC와 제휴하거나 보증받지 않습니다. 제품 및 모델 이름은 호환성을 설명하기 위해서만 사용됩니다.
1. 계정, 지갑 및 키
- LLMsRelay에 로그인하세요. 계정이 없으면 가입한 후 Billing > Uncensored로 돌아오세요.
- 별도 GLM 현금 지갑을 $45, $100, $500 또는 $1,000로 충전하세요. USD 1:1로 적립되며 가입, 추천, 프로모션 보너스는 없습니다. 표준 API 등가 사용량 패키지는 이 지갑을 충전하지 않습니다.
- API Keys에서 키를 만들고 GLM 5.3 그룹을 선택하세요. 생성 시 전체 비밀 키를 복사해 로컬에 보관하세요. 다른 서비스 키가 아닌 대시보드에서 발급한 전체 sk- 키를 사용하세요.
- 현금 잔액과 키의 남은 한도를 모두 확인하세요. 표준 잔액이 있어도 충분하지 않습니다. 키마다 그룹은 하나이며 Claude, GPT, GLM 동시 사용에는 별도 키가 필요합니다.
2. 공개 모델 ID 선택
| 공개 모델 ID | 컨텍스트 (토큰) | 입력 |
|---|---|---|
| glm-5-3-uncensored | 262,144 | 텍스트 및 이미지 |
| glm-5-3-uncensored-1m | 1,000,000 | 텍스트 전용 |
curl https://api.llmsrelay.com/v1/models \
-H "Authorization: Bearer $LLMSRELAY_API_KEY"3. 자격 증명 보관
로컬 셸 또는 비밀 관리 도구에 LLMSRELAY_API_KEY를 설정하세요. 예제의 자리표시자는 로컬에서만 바꾸고 공유 채팅, URL, 브라우저 코드, 저장소에 넣지 마세요. Anthropic SDK base URL에는 /v1을 붙이지 않으며 OpenAI SDK에는 /v1이 정확히 한 번 필요합니다.
export LLMSRELAY_API_KEY="YOUR_LLMSRELAY_KEY"4. Anthropic 호환 Messages
Messages는 x-api-key와 anthropic-version을 사용하며 max_tokens가 필수입니다. choices가 아닌 response.content의 텍스트 블록을 읽으세요. 시스템 지시는 최상위 system 필드에 넣습니다. 1M 모델은 model만 glm-5-3-uncensored-1m으로 바꾸면 됩니다.
curl https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "Explain what an API gateway does."
}
]
}'5. OpenAI 호환 Chat Completions
Chat Completions는 Authorization: Bearer를 사용합니다. 시스템 지시는 system 메시지에 넣고 choices[0].message.content를 읽으세요. 정확한 공개 ID를 사용하세요. 유효하고 충전된 키의 GET /v1/models가 접근 가능 모델을 반환하며 유효 잔액이 없으면 402가 반환될 수 있습니다.
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"messages": [
{
"role": "system",
"content": "You are a concise assistant."
},
{
"role": "user",
"content": "Explain what an API gateway does."
}
]
}'6. Python 및 TypeScript SDK
필요한 SDK만 설치하고 키는 서버 환경에 보관하세요. Anthropic SDK는 content 블록, OpenAI SDK는 choices를 반환합니다. model을 1M 공개 ID로 바꾸면 같은 예제를 사용할 수 있습니다.
pip install anthropic openaiimport os
from anthropic import Anthropic
client = Anthropic(
api_key=os.environ["LLMSRELAY_API_KEY"],
base_url="https://api.llmsrelay.com",
)
response = client.messages.create(
model="glm-5-3-uncensored",
max_tokens=512,
messages=[{"role": "user", "content": "Hello!"}],
)
print("".join(block.text for block in response.content if block.type == "text"))import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LLMSRELAY_API_KEY"],
base_url="https://api.llmsrelay.com/v1",
)
response = client.chat.completions.create(
model="glm-5-3-uncensored",
max_tokens=512,
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)npm install @anthropic-ai/sdk openaiimport Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: process.env.LLMSRELAY_API_KEY,
baseURL: "https://api.llmsrelay.com",
});
const response = await client.messages.create({
model: "glm-5-3-uncensored",
max_tokens: 512,
messages: [{ role: "user", content: "Hello!" }],
});
for (const block of response.content) {
if (block.type === "text") console.log(block.text);
}import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.LLMSRELAY_API_KEY,
baseURL: "https://api.llmsrelay.com/v1",
});
const response = await client.chat.completions.create({
model: "glm-5-3-uncensored",
max_tokens: 512,
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);7. Claude Code 및 호환 IDE
Claude Code는 아래 변수를 설정한 별도 GLM 프로필 또는 셸을 사용하세요. 기존 프로필을 확인 없이 덮어쓰지 마세요. 해당 프로필의 충돌하는 ANTHROPIC_AUTH_TOKEN을 제거하세요. 기본 및 빠른 모델을 GLM으로 지정해 GLM 키로 Claude를 요청하지 않도록 하세요. Cline, Roo Code, Continue에서는 OpenAI Compatible, base URL https://api.llmsrelay.com/v1, GLM 키와 공개 ID를 선택하세요. 도구 호환성은 클라이언트마다 다르므로 긴 세션 전에 짧은 요청을 확인하세요.
export ANTHROPIC_BASE_URL="https://api.llmsrelay.com"
export ANTHROPIC_API_KEY="$LLMSRELAY_API_KEY"
export ANTHROPIC_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_OPUS_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_SONNET_MODEL="glm-5-3-uncensored"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="glm-5-3-uncensored"
export ANTHROPIC_SMALL_FAST_MODEL="glm-5-3-uncensored"
# Only in this dedicated GLM shell/profile:
unset ANTHROPIC_AUTH_TOKEN
claude8. 스트리밍 (SSE)
요청에 stream:true를 추가하고 curl -N을 사용하세요. Chat Completions는 [DONE]까지 choices[].delta.content와 tool_calls 인자 조각을 모읍니다. stream_options.include_usage:true는 최종 usage 청크를 요청하며 choices가 비어 있을 수 있습니다.
Messages는 message_start, content_block_delta, message_delta, message_stop 등의 이름 있는 이벤트를 보냅니다. 텍스트, thinking, 도구 JSON 델타를 분리하세요. HTTP 200이어도 error 이벤트를 확인하세요. 스트림 시작은 완료를 뜻하지 않습니다. 원시 SSE를 일반 JSON처럼 합치지 마세요.
curl -N https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"messages": [
{
"role": "system",
"content": "You are a concise assistant."
},
{
"role": "user",
"content": "Explain what an API gateway does."
}
],
"stream": true,
"stream_options": {
"include_usage": true
}
}'curl -N https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "Explain what an API gateway does."
}
],
"stream": true
}'9. 내장 웹 검색
내장 검색은 원격 실행되므로 로컬 web_search 함수를 구현할 필요가 없습니다. Chat Completions에는 web_search_options:{}, Messages에는 아래 버전 지정 도구를 사용하세요. 두 모델은 SSE를 포함해 두 형식을 지원합니다.
인용 메타데이터, URL, 서버 도구 결과 블록을 보존하세요. 최종 답변 전에 여러 검색 이벤트가 올 수 있습니다. 실제 검색 실행과 출처를 확인하세요. 자연스러운 답변만으로 검색 실행을 입증할 수 없습니다.
클라이언트에서 조합을 검증하기 전에는 사용자 함수와 Chat 내장 검색을 별도 요청으로 실행하세요. 두 매개변수 수락이 두 도구 실행을 의미하지는 않습니다.
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 2048,
"messages": [
{
"role": "user",
"content": "Search the web for recent Python releases. Cite your sources."
}
],
"web_search_options": {}
}'curl https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 2048,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "Search the web for recent Python releases. Cite your sources."
}
],
"tools": [
{
"type": "web_search_2025_03_05",
"name": "web_search",
"max_uses": 3
}
]
}'10. 사용자 정의 함수 도구
JSON Schema로 함수를 정의하고 모델의 호출을 기다리세요. 서버는 인자를 검증하고 허용된 작업을 실행한 뒤 결과를 반환합니다. 모델이 만든 임의 셸 명령이나 검증되지 않은 인자를 실행하지 마세요.
Chat은 tool_calls가 포함된 assistant 메시지 전체와 각 tool_call_id에 대응하는 role:tool 메시지를 추가합니다. Messages는 전체 assistant content 뒤에 tool_use_id가 일치하는 tool_result를 담은 user 메시지를 추가합니다. 추론과 서명 블록이 있으면 보존하세요. 텍스트 응답까지 계속하되 다중 호출을 지원하고 반복 횟수를 제한하세요.
아래 예제는 이어 보내는 구조이며 실제 날씨 서비스가 아닙니다. 실제 호출 ID, 전체 assistant 응답, 검증된 로컬 실행 결과로 바꾸세요. SSE 인자 조각을 모두 모은 뒤 JSON을 파싱하세요.
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"messages": [
{
"role": "user",
"content": "What is the weather in Paris?"
}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string"
}
},
"required": [
"city"
]
}
}
}
]
}'# messages for the next Chat request; use the ACTUAL assistant message and call ID:
[
{"role":"user","content":"What is the weather in Paris?"},
{"role":"assistant","content":null,"tool_calls":[
{"id":"ACTUAL_CALL_ID","type":"function","function":{"name":"get_weather","arguments":"{\"city\":\"Paris\"}"}}
]},
{"role":"tool","tool_call_id":"ACTUAL_CALL_ID","content":"{\"temperature_c\":18}"}
]curl https://api.llmsrelay.com/v1/messages \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 512,
"system": "You are a concise assistant.",
"messages": [
{
"role": "user",
"content": "What is the weather in Paris?"
}
],
"tools": [
{
"name": "get_weather",
"description": "Get current weather for a city",
"input_schema": {
"type": "object",
"properties": {
"city": {
"type": "string"
}
},
"required": [
"city"
]
}
}
]
}'# messages for the next Messages request; preserve the ACTUAL full assistant content:
[
{"role":"user","content":"What is the weather in Paris?"},
{"role":"assistant","content":[
{"type":"tool_use","id":"ACTUAL_TOOL_USE_ID","name":"get_weather","input":{"city":"Paris"}}
]},
{"role":"user","content":[
{"type":"tool_result","tool_use_id":"ACTUAL_TOOL_USE_ID","content":"{\"temperature_c\":18}"}
]}
]11. 추론 및 구조화 출력
Chat은 reasoning_effort(예: high), Messages는 output_config.effort 또는 thinking.budget_tokens를 사용합니다. include_reasoning:false는 추론 표시만 숨기며 비용은 남습니다. 추론과 답변 모두를 위한 출력 예산을 확보하세요. 아래 Chat JSON Schema 예제의 반환 JSON도 앱에서 검증하세요. SDK 비표준 필드는 Python extra_body 또는 직접 HTTP로 전달하세요.
curl https://api.llmsrelay.com/v1/chat/completions \
-H "Authorization: Bearer $LLMSRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"max_tokens": 2048,
"messages": [
{
"role": "user",
"content": "Return a short greeting as JSON."
}
],
"reasoning_effort": "high",
"include_reasoning": false,
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "greeting",
"strict": true,
"schema": {
"type": "object",
"properties": {
"greeting": {
"type": "string"
}
},
"required": [
"greeting"
],
"additionalProperties": false
}
}
}
}'12. 토큰 계산, 캐시 및 이미지
긴 요청 전에 같은 모델과 입력으로 POST /v1/messages/count_tokens를 호출할 수 있습니다. 입력 토큰 추정치이며 최종 출력이나 요금은 아닙니다. 전체 입력과 출력은 컨텍스트 한도 안에 있어야 합니다.
안정적인 지시와 도구 스키마는 프롬프트 앞부분에 두세요. 반복해도 캐시 적중은 보장되지 않습니다. usage의 캐시 읽기를 확인하세요. 캐시 쓰기는 입력 요금이며 추론은 출력에 포함됩니다.
이미지는 glm-5-3-uncensored만 지원합니다. Chat은 image_url, Messages는 지원 URL 또는 base64 소스의 image 블록을 사용합니다. 1M 모델은 텍스트 전용으로 이미지에 HTTP 400을 반환합니다.
curl https://api.llmsrelay.com/v1/messages/count_tokens \
-H "x-api-key: $LLMSRELAY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5-3-uncensored",
"messages": [
{
"role": "user",
"content": "Explain what an API gateway does."
}
]
}'13. 문제 해결
| 상태 / 증상 | 조치 |
|---|---|
| 401 | 전체 LLMsRelay 키, 헤더, 환경 변수를 확인하고 노출된 키는 교체하세요. |
| 403 / 모델 접근 불가 | GLM 5.3 그룹과 허용 모델을 확인하세요. Basic 또는 Codex 키로 GLM에 접근할 수 없습니다. |
| 402 / insufficient_credits | GLM 현금 지갑과 키 한도를 모두 확인하세요. max_tokens를 포함한 예상 예약액이 필요합니다. 출력 예산을 줄이거나 충전하세요. |
| 400 | 모델 ID, JSON, max_tokens, 도구 스키마를 확인하세요. 1M 입력에서 이미지를 제거하세요. |
| 404 / 잘못된 경로 | Messages는 /v1/messages, Chat은 /v1/chat/completions입니다. 중복 /v1을 제거하세요. |
| 429 | Retry-After를 따르고 제한된 지수 백오프를 적용하며 동시 요청을 줄이세요. |
| 5xx / SSE 오류 | request ID와 종료 이벤트를 기록하세요. 반복되면 지원팀에 문의하세요. 재시도는 별도 과금될 수 있습니다. |
| 지갑 확인 불가 | 결제를 반복하거나 잔액이 0이라고 판단하지 마세요. 잔액 확인을 기다리거나 지원팀에 문의하세요. |
14. Responses 호환성 안내
이 가이드는 Messages와 Chat Completions를 권장합니다. 네이티브 Codex는 Responses를 사용하며 OpenAI 호환 IDE와 Codex CLI 설정은 다릅니다.
GLM Responses는 상태를 저장하지 않으며 previous_response_id를 지원하지 않습니다. 이전 input/output 및 function_call_output을 명시적으로 다시 보내세요.
2026년 10월 9일 검사에서 기본 모델 Responses 검색 SSE는 실패했고 비스트리밍 검색은 완료되었습니다. 검색에는 Chat 또는 Messages를 권장합니다. 1M Responses 검색은 출처 URL을 반환했지만 최종 인용 주석은 없었습니다.
이는 호환성 관찰이며 모든 향후 요청에 대한 보장은 아닙니다. 오류를 처리하고 실제 클라이언트를 검증하세요. 최대 컨텍스트 부하 시험 및 모든 타사 클라이언트는 검사 범위가 아닙니다.
15. 최종 USD 토큰 요금
100만 토큰당 최종 고객 USD 요금이며 표에 추가 배수를 적용하지 않습니다. 실제 입력, 출력, 캐시 토큰은 별도 GLM 현금 지갑에서 차감되며 이 흐름에는 보너스가 없습니다.
| 모델 | 입력 | 출력 | 캐시 읽기 | 캐시 쓰기 |
|---|---|---|---|---|
| glm-5-3-uncensored | $4.00 | $12.00 | $0.40 | $4.00 |
| glm-5-3-uncensored-1m | $12.00 | $20.00 | $1.20 | $12.00 |
GLM 지갑 충전
별도 GLM 현금 지갑을 $45, $100, $500 또는 $1,000로 충전하세요. USD 1:1로 적립되며 가입, 추천, 프로모션 보너스는 없습니다. 표준 API 등가 사용량 패키지는 이 지갑을 충전하지 않습니다.
GLM 지갑 충전