/Docs
Back to dashboard

LLMsRelay Platform Pricing Explained

A clear breakdown of LLMsRelay token pricing, platform usage, cache rates, and cost controls.

LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility.

llmsrelay --pricingup to −91%
$ Pay once. Get ~11× the Anthropic balance.
Same Anthropic per-token rates. Massive discount on the top-up.
you payAnthropic-equivalent balancediscount
$45$500balance−91%
$90popular$1,000balance−91%
Tokens are billed at Anthropic's exact per-token rates · Balance never expires · No subscription

How Platform Pricing Works

LLMsRelay pricing is based on tokens, the units of text that a model processes. A request can include input, output, cache, and thinking token categories depending on the model and request.

Platform usage is deducted from your balance based on actual token consumption and the LLMsRelay rate card.

What Are Tokens?

A token is roughly 4 characters of English text, or about ¾ of a word. For example:

  • "Hello, world!" ≈ 4 tokens
  • A 1,000-word article ≈ 1,300 tokens
  • A typical code file ≈ 500–2,000 tokens

Token Categories

CategoryWhat It IsRelative Cost
Input (uncached)Standard input tokens sent to the modelBase rate
OutputTokens generated by Claude in the response3–5× input rate
Cache WriteInput tokens written to prompt cache (first request)1.25× input rate
Cache ReadInput cached from a previous request0.1× input rate
Cache read tokens are up to 90% cheaper than uncached input. Use prompt caching for repeated system prompts to dramatically reduce costs.

Per-Model Pricing

Different model IDs have different price points. The LLMsRelay rate card below is shown per million tokens (MTok):

ModelInput / 1MOutput / 1MBest For
claude-opus-4.7$3.50$17.50Reasoning and complex refactors
claude-opus-4.6$3.50$17.50Long-running coding tasks
claude-sonnet-4.6$2.10$10.50Coding and general tasks
claude-haiku-4.5$0.70$3.50Fast responses and classification

Cache pricing: Cache Write = 1.25× input, Cache Read = 0.10× input.

Cost Calculation Example

Here's how to estimate the cost of a typical request using Claude Sonnet 4.6 at LLMsRelay's public rate ($2.10/MTok input, $10.50/MTok output):

Example: 2,000 input + 500 output tokenstext
Input: 2,000 tokens × $2.10 / 1,000,000 = $0.0042
Output: 500 tokens × $10.50 / 1,000,000 = $0.00525
─────────────────────────────────────────────
Total: $0.00945 per request

At this rate, a $45 platform usage pack can cover about 529 similar requests.
That same request would deduct about 1.89 credits.

Tips to Reduce Costs

1. Choose the Right Model

Don't use Opus for tasks that Sonnet or Haiku can handle. Sonnet is 5× cheaper than Opus and handles most coding and general tasks excellently.

2. Use Prompt Caching

If you send the same system prompt repeatedly, enable prompt caching. After the first request, cached tokens cost only 10% of the standard input rate.

3. Optimize Prompt Length

Remove unnecessary context from your prompts. Shorter prompts = fewer input tokens = lower cost.

4. Set Per-Key Limits

Use LLMsRelay's per-key credit limits to prevent unexpected spending, especially on development and testing keys.

5. Monitor Usage

Check the LLMsRelay dashboard regularly to understand your consumption patterns and optimize accordingly.

Cost Controls

Use prompt caching for repeated context, choose a model ID that fits the task, set per-key limits, and monitor usage from the dashboard. Check GET /v1/models and the current rate card before production rollout.

Ready to start?

Create a key and configure a compatible API route in under 2 minutes.

View Plans