LLMsRelay Platform Pricing Explained
A clear breakdown of LLMsRelay token pricing, platform usage, cache rates, and cost controls.
LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility.
| you pay | Anthropic-equivalent balance | discount |
|---|---|---|
| $45 | $500balance | −91% |
| $90popular | $1,000balance | −91% |
How Platform Pricing Works
LLMsRelay pricing is based on tokens, the units of text that a model processes. A request can include input, output, cache, and thinking token categories depending on the model and request.
Platform usage is deducted from your balance based on actual token consumption and the LLMsRelay rate card.
What Are Tokens?
A token is roughly 4 characters of English text, or about ¾ of a word. For example:
- "Hello, world!" ≈ 4 tokens
- A 1,000-word article ≈ 1,300 tokens
- A typical code file ≈ 500–2,000 tokens
Token Categories
| Category | What It Is | Relative Cost |
|---|---|---|
| Input (uncached) | Standard input tokens sent to the model | Base rate |
| Output | Tokens generated by Claude in the response | 3–5× input rate |
| Cache Write | Input tokens written to prompt cache (first request) | 1.25× input rate |
| Cache Read | Input cached from a previous request | 0.1× input rate |
Per-Model Pricing
Different model IDs have different price points. The LLMsRelay rate card below is shown per million tokens (MTok):
| Model | Input / 1M | Output / 1M | Best For |
|---|---|---|---|
| claude-opus-4.7 | $3.50 | $17.50 | Reasoning and complex refactors |
| claude-opus-4.6 | $3.50 | $17.50 | Long-running coding tasks |
| claude-sonnet-4.6 | $2.10 | $10.50 | Coding and general tasks |
| claude-haiku-4.5 | $0.70 | $3.50 | Fast responses and classification |
Cache pricing: Cache Write = 1.25× input, Cache Read = 0.10× input.
Cost Calculation Example
Here's how to estimate the cost of a typical request using Claude Sonnet 4.6 at LLMsRelay's public rate ($2.10/MTok input, $10.50/MTok output):
Input: 2,000 tokens × $2.10 / 1,000,000 = $0.0042
Output: 500 tokens × $10.50 / 1,000,000 = $0.00525
─────────────────────────────────────────────
Total: $0.00945 per request
At this rate, a $45 platform usage pack can cover about 529 similar requests.
That same request would deduct about 1.89 credits.Tips to Reduce Costs
1. Choose the Right Model
Don't use Opus for tasks that Sonnet or Haiku can handle. Sonnet is 5× cheaper than Opus and handles most coding and general tasks excellently.
2. Use Prompt Caching
If you send the same system prompt repeatedly, enable prompt caching. After the first request, cached tokens cost only 10% of the standard input rate.
3. Optimize Prompt Length
Remove unnecessary context from your prompts. Shorter prompts = fewer input tokens = lower cost.
4. Set Per-Key Limits
Use LLMsRelay's per-key credit limits to prevent unexpected spending, especially on development and testing keys.
5. Monitor Usage
Check the LLMsRelay dashboard regularly to understand your consumption patterns and optimize accordingly.
Cost Controls
Use prompt caching for repeated context, choose a model ID that fits the task, set per-key limits, and monitor usage from the dashboard. Check GET /v1/models and the current rate card before production rollout.