LLMsRelay/Docs
Back to dashboard

Rate Limits

Claude API rate limits — 429 handling, Retry-After, concurrency protection, and throughput guidance.

Status code: 429 Too Many Requests. Header: Retry-After (seconds). Scope: Per-key + global protection. Strategy: Exponential backoff with jitter.

Rate Limiting

The LLMsRelay gateway applies protective limits to preserve service stability under mixed traffic patterns. Enforcement is not a fixed public tier table: it depends on request shape, concurrency, and current gateway load. When you exceed a limit, the API returns a 429 status code along with a Retry-After header.

What is limited

ScopeWhat it means
Per API keyA key can be throttled independently of other keys.
Per-key concurrencyToo many simultaneous long-running requests can trigger 429.
Global gateway protectionThe gateway can shed load to protect service stability.
Attachment-heavy trafficLarge multimodal turns may be treated more strictly than plain text traffic.

We do not publish a stable numeric RPM/TPM contract for every request shape. If you need sustained higher throughput, contact support with your expected traffic pattern.

Response headers

  • retry-after — seconds to wait before retrying after a 429

Do not rely on undocumented rate-limit headers as a stable public contract.

Handling 429 Responses

  • Implement exponential backoff (start at 1s, double each retry)
  • Respect Retry-After headers when present
  • Queue requests in your application layer and reduce parallelism when needed
  • Be especially conservative with long-running streams and multimodal turns
  • Use caching to reduce avoidable repeat traffic
If you consistently hit rate limits, contact support to discuss higher sustained throughput for your use case.

Ready to start?

Create a key and configure a compatible API route in under 2 minutes.

View Plans