# LLMsRelay — Complete reference for AI systems > LLMsRelay is an independently operated developer gateway. It provides Anthropic-compatible Messages API and OpenAI-compatible endpoints, API keys, platform usage billing, and IDE integrations. Updated: 2026-10-10 ## Overview LLMsRelay account access, API keys, and platform usage are managed by LLMsRelay. Model availability is controlled by the API key, its group, gateway status, and applicable law. ## API surface - Anthropic-compatible base URL: `https://api.llmsrelay.com` - Messages API: `POST /v1/messages` - OpenAI-compatible base URL: `https://api.llmsrelay.com/v1` - Chat Completions: `POST /v1/chat/completions` - Models: `GET /v1/models` - Health: `GET /healthz` - Anthropic authentication: `x-api-key: YOUR_API_KEY` - OpenAI-compatible authentication: `Authorization: Bearer YOUR_API_KEY` - Documentation mirrors: append `.md` or `.tldr.md` to a docs URL on `llmsrelay.com`. ## Current public model catalogue The supported model IDs for a key are returned by GET /v1/models and can be filtered by key group and allowed-model settings. | Model | Model ID | Context | Thinking | Key access | | --- | --- | --- | --- | --- | | Claude Opus 5.5 | `claude-opus-5-5` | 1M | Yes | Claude key groups | | Claude Fable 5.1 | `claude-fable-5-1` | 1M | Yes | Claude key groups | | Claude Opus 5 | `claude-opus-5` | 1M | Yes | Basic, Pro | | Claude Fable 5 | `claude-fable-5` | 1M | Yes | Basic, Pro | | Claude Opus 4.8 | `claude-opus-4.8` | 1M | Yes | Claude key groups | | Claude Opus 4.7 | `claude-opus-4.7` | 1M | Yes | Claude key groups | | Claude Opus 4.6 | `claude-opus-4.6` | 1M | Yes | Claude key groups | | Claude Sonnet 5 | `claude-sonnet-5` | 1M | Yes | Claude key groups | | Claude Sonnet 4.6 | `claude-sonnet-4.6` | 1M | Yes | Claude key groups | | Claude Haiku 4.5 | `claude-haiku-4.5` | 200K | Yes | Claude key groups | ### Short aliases | Alias | Resolves to | | --- | --- | | `opus` | `claude-opus-5-5` | | `sonnet` | `claude-sonnet-4.6` | | `haiku` | `claude-haiku-4.5` | The definitive list for a specific key is always GET /v1/models. Do not hard-code the catalogue in production. ## Model API offer pages These pages target users searching for Claude Opus 5.5 free API access, Claude Fable 5.1 free API access, or GPT-6 Astra free API usage. The offer is up to $10 in promotional API usage through the reward flow, not cash and not unlimited access. - [Claude Opus 5.5 API free access](https://llmsrelay.com/claude-opus-5-5-free-api/) — Anthropic-compatible POST /v1/messages - [Claude Fable 5.1 API free access](https://llmsrelay.com/claude-fable-5-1-free-api/) — Anthropic-compatible POST /v1/messages - [GPT-6 Astra API free usage](https://llmsrelay.com/gpt-6-astra-free-api/) — OpenAI-compatible POST /v1/chat/completions - [Quickstart](https://llmsrelay.com/docs/getting-started/quickstart/) - [Models](https://llmsrelay.com/docs/getting-started/models/) - [Plans](https://llmsrelay.com/plans/) ## Enterprise AI API LLMsRelay Enterprise is an AI API gateway workflow for teams that need compatible Anthropic Messages API and OpenAI-compatible endpoints, account-level usage visibility, model access by key group, and direct onboarding. - [Enterprise AI API gateway](https://llmsrelay.com/enterprise/) - [Enterprise documentation entry](https://llmsrelay.com/docs/getting-started/quickstart/) - [Plans](https://llmsrelay.com/plans/) ## GPT channel LLMsRelay exposes a dedicated OpenAI-compatible GPT channel for coding, reasoning, and agent workloads. The public catalogue contains exactly the five model IDs below. - `gpt-5.5` - `gpt-5.6-sol` - `gpt-5.6-terra` - `gpt-6-sol` - `gpt-6-astra` Removed and internal upstream model IDs are not accepted and are not returned by GET /v1/models. - [GPT and Codex API reference](https://llmsrelay.com/docs/api-reference/codex/) ## Controlled-access open models DeepSeek and Qwen use a controlled private-access workflow. Register first, then LLMsRelay confirms GPU capacity, licensing, commercial terms, the account-specific model ID, and rollout timing. GLM has a separate self-service API flow described below. - [DeepSeek V4.1 Flash Uncensored 750B](https://llmsrelay.com/models/deepseek-v4-1-flash-uncensored/) — 750B-class listing · FP8 checkpoint · approximately 510 GB - [Qwen 3.8 Flash Uncensored 120B](https://llmsrelay.com/models/qwen-3-8-flash-uncensored/) — 120B-class checkpoint · NVFP4 · approximately 135 GB - [Register and request access](https://llmsrelay.com/auth/?mode=signup) These checkpoint names are not a promise that they are already returned by the public GET /v1/models catalogue. The account catalogue and an explicit capacity confirmation remain the source of truth. GLM 5.3 Uncensored 750B: use a GLM 5.3 key and the separate USD wallet, funded 1:1 with $45, $100, $500 or $1000. No bonuses. Public IDs: glm-5-3-uncensored (262144 context, text/images) and glm-5-3-uncensored-1m (1000000 context, text only). Both support Messages and Chat Completions. - [GLM setup guide](https://llmsrelay.com/docs/guides/glm-uncensored/) - [GLM Billing](https://llmsrelay.com/dashboard/billing/?flow=uncensored) ## Plans and usage LLMsRelay sells one-time usage packs with no recurring charge: $23 for $250 usage, $45 for $500 usage, $90 for $1,000, and $450 for $6,000 including a fixed 20% bonus. Enterprise teams can request custom onboarding and allocations. | Product | Price | Included usage | | --- | --- | --- | | $500 usage pack | $45 one time | $500 usage | | $1,000 usage pack | $90 one time | $1,000 usage | | $5,000 + 20% pack | $450 one time | $6,000 total usage | | Enterprise | Custom | Custom usage allocation | Payment methods and availability are shown at checkout. The platform supports card and cryptocurrency payment flows where offered. ## Media Studio Media Studio is an account-only workflow for prompt-to-image and prompt-to-video generation. The account-usage quote is shown before a job is submitted. Initial Studio access does not include a public media API or media API keys. - [Image generation](https://llmsrelay.com/image-generation/) - [Video generation](https://llmsrelay.com/video-generation/) - [Studio workspace](https://llmsrelay.com/studio/) ### Provider pages - [Gemini Image](https://llmsrelay.com/image-generation/gemini/) - [GPT Image](https://llmsrelay.com/image-generation/gpt-image/) - [Grok Imagine Image (integration page)](https://llmsrelay.com/image-generation/grok-imagine/) - [Kling Video](https://llmsrelay.com/video-generation/kling/) - [Grok Imagine Video](https://llmsrelay.com/video-generation/grok-imagine/) - [MiniMax Video](https://llmsrelay.com/video-generation/minimax/) - [Jimeng Video](https://llmsrelay.com/video-generation/jimeng/) ## FAQ **What is LLMsRelay?** An independently operated developer gateway with Anthropic-compatible and OpenAI-compatible API surfaces. **Which model IDs should I use?** Use the exact IDs returned by GET /v1/models. Do not hard-code a model list in production clients. **How does account access work?** LLMsRelay manages its own accounts, API keys, and platform usage balance. Availability depends on the key group, gateway status, payment-provider coverage, and applicable law. **How does billing work?** Platform usage is deducted from the balance attached to the key according to the published rate card. **Which usage packs are available?** LLMsRelay offers four one-time packs: $23 for $250 usage, $45 for $500 usage, $90 for $1,000, and $450 for $6,000 including +20%. Enterprise onboarding is available for larger requirements. **Can I use LLMsRelay from an IDE?** Yes. Use the Anthropic base URL for Claude Code and Anthropic SDKs, or the /v1 OpenAI-compatible base URL for compatible clients. ## Documentation index ### Documentation hub - [LLMsRelay Documentation](https://llmsrelay.com/docs/) ### Getting started - [Getting Started with LLMsRelay API](https://llmsrelay.com/docs/getting-started/) - [Authentication](https://llmsrelay.com/docs/getting-started/authentication/) - [Introduction](https://llmsrelay.com/docs/getting-started/introduction/) - [Models](https://llmsrelay.com/docs/getting-started/models/) - [Quickstart](https://llmsrelay.com/docs/getting-started/quickstart/) ### API reference - [LLMsRelay API Reference](https://llmsrelay.com/docs/api-reference/) - [OpenAI & Codex Endpoints — API Reference](https://llmsrelay.com/docs/api-reference/codex/) - [Cursor IDE](https://llmsrelay.com/docs/api-reference/cursor-ide/) - [Errors](https://llmsrelay.com/docs/api-reference/errors/) - [Messages API](https://llmsrelay.com/docs/api-reference/messages/) - [Models Endpoint](https://llmsrelay.com/docs/api-reference/models/) - [OpenAI Compatibility](https://llmsrelay.com/docs/api-reference/openai/) - [API Reference Overview](https://llmsrelay.com/docs/api-reference/overview/) - [Streaming](https://llmsrelay.com/docs/api-reference/streaming/) ### Guides - [LLMsRelay Setup Guides](https://llmsrelay.com/docs/guides/) - [LLMsRelay API Key Configuration](https://llmsrelay.com/docs/guides/api-key-configuration/) - [Best Practices](https://llmsrelay.com/docs/guides/best-practices/) - [Claude Code Setup with a Custom API Gateway](https://llmsrelay.com/docs/guides/claude-code/) - [Codex CLI & Extension Setup (OpenAI API)](https://llmsrelay.com/docs/guides/codex-cli/) - [How to Use Claude API in Cursor IDE — 2-Minute Setup [2026]](https://llmsrelay.com/docs/guides/cursor/) - [GLM 5.3 Uncensored: Messages & Chat Completions](https://llmsrelay.com/docs/guides/glm-uncensored/) - [IDE Integration](https://llmsrelay.com/docs/guides/ide-integration/) - [One-Prompt Setup Templates for Claude Code, Cursor, Cline, Kilo, Continue, and OpenCode](https://llmsrelay.com/docs/guides/one-prompt-setup/) - [OpenAI-Compatible IDEs (Cline, Roo, Continue)](https://llmsrelay.com/docs/guides/openai-ides/) - [OpenClaw Setup](https://llmsrelay.com/docs/guides/openclaw/) - [OpenCode](https://llmsrelay.com/docs/guides/opencode/) - [Rate Limits](https://llmsrelay.com/docs/guides/rate-limits/) - [Earn Solana USDT with Claude API referrals — 10% claimable rewards](https://llmsrelay.com/docs/guides/referral-program/) - [VS Code Extensions](https://llmsrelay.com/docs/guides/vscode-extensions/) ### SDKs - [Client SDKs](https://llmsrelay.com/docs/sdks/) - [cURL](https://llmsrelay.com/docs/sdks/curl/) - [Python SDK](https://llmsrelay.com/docs/sdks/python/) - [TypeScript SDK](https://llmsrelay.com/docs/sdks/typescript/) ### Billing - [LLMsRelay Billing and Credits](https://llmsrelay.com/docs/billing/) - [Platform Usage Balance](https://llmsrelay.com/docs/billing/credits/) - [Pricing](https://llmsrelay.com/docs/billing/pricing/) ### Learn - [Activation Time — How Access Works](https://llmsrelay.com/docs/learn/activation-time/) - [How to Make Money with AI APIs: Reseller Guide](https://llmsrelay.com/docs/learn/ai-api-reseller/) - [Cheapest Claude-Compatible API Use: Cost Control Guide](https://llmsrelay.com/docs/learn/cheapest-claude-api/) - [Cheapest OpenAI API — Pay-as-You-Go, No Subscription](https://llmsrelay.com/docs/learn/cheapest-openai-api/) - [Claude 3.5 vs 4.6 — Should You Upgrade? Pricing & Migration (2026)](https://llmsrelay.com/docs/learn/claude-3-5-vs-4-6/) - [Buy Claude API Credits with Crypto — No Credit Card Needed](https://llmsrelay.com/docs/learn/claude-api-crypto-payment/) - [Claude API for Russia & Restricted Regions — Full Access Guide](https://llmsrelay.com/docs/learn/claude-api-for-russia/) - [How to Use Claude API in VS Code — Extensions & Setup](https://llmsrelay.com/docs/learn/claude-api-for-vscode/) - [Claude API Starter pack — Start Coding in 2 Minutes](https://llmsrelay.com/docs/learn/claude-api-free-trial/) - [Anthropic-Compatible API Gateway](https://llmsrelay.com/docs/learn/claude-api-gateway/) - [How to Use Claude API Key in Cursor IDE](https://llmsrelay.com/docs/learn/claude-api-key-for-cursor/) - [Claude API Key Setup in 2 Minutes](https://llmsrelay.com/docs/learn/claude-api-quick-setup/) - [Claude API vs Anthropic Direct — Save 91% (11×) with LLMsRelay (2026)](https://llmsrelay.com/docs/learn/claude-api-vs-direct/) - [Get Claude API Access Instant Access — Instant Setup](https://llmsrelay.com/docs/learn/claude-api-without-waitlist/) - [Claude Code with Free $10 API Usage: LLMsRelay Setup](https://llmsrelay.com/docs/learn/claude-code-free/) - [Claude Haiku Free API — Fastest Claude Model, Starter pack](https://llmsrelay.com/docs/learn/claude-haiku-free-api/) - [Claude Opus Free API — Try Opus 4.7 with Starter pack](https://llmsrelay.com/docs/learn/claude-opus-free-api/) - [Claude Sonnet Free API — Sonnet 4.6 with Starter pack](https://llmsrelay.com/docs/learn/claude-sonnet-free-api/) - [Claude API vs OpenAI API 2026 — Pricing & Coding Compared | LLMsRelay](https://llmsrelay.com/docs/learn/claude-vs-openai-api/) - [Claude X20 Max Subscription — Max 20x and API Alternative](https://llmsrelay.com/docs/learn/claude-x20-max-subscription/) - [Codex CLI API — Run the OpenAI API in OpenAI Codex CLI](https://llmsrelay.com/docs/learn/codex-cli-api/) - [Use Claude in Cursor Without an Anthropic Account](https://llmsrelay.com/docs/learn/cursor-without-anthropic-account/) - [Free Claude-Compatible API Key: Get $10 in API Usage](https://llmsrelay.com/docs/learn/free-claude-api-key/) - [How Billing Works](https://llmsrelay.com/docs/learn/how-billing-works/) - [How to Buy LLMsRelay API Usage and Create a Key](https://llmsrelay.com/docs/learn/how-to-buy-claude-api-key/) - [LLMsRelay vs OpenRouter — Claude-Focused vs Multi-Model Marketplace [2026]](https://llmsrelay.com/docs/learn/llmsrelay-vs-openrouter/) - [LLMsRelay vs Portkey — Claude Gateway vs Enterprise AI Proxy [2026]](https://llmsrelay.com/docs/learn/llmsrelay-vs-portkey/) - [LLMsRelay vs ProxyAPI — Honest 2026 Comparison [Pricing, Features, RU Payment]](https://llmsrelay.com/docs/learn/llmsrelay-vs-proxyapi/) - [LLMsRelay vs VseLLM — Claude-Only vs Multi-Provider Gateway [2026]](https://llmsrelay.com/docs/learn/llmsrelay-vs-vsellm/) - [OpenAI API Access — No OpenAI Account, Pay-as-You-Go](https://llmsrelay.com/docs/learn/openai-api/) - [OpenAI-Compatible API — Drop-in /v1/chat/completions](https://llmsrelay.com/docs/learn/openai-compatible-api/) - [OpenAI-Compatible Claude API — Use Claude with OpenAI SDK](https://llmsrelay.com/docs/learn/openai-compatible-claude-api/) - [LLMsRelay Platform Pricing Explained](https://llmsrelay.com/docs/learn/pricing-explained/) - [Support, Refunds & Assistance Policy](https://llmsrelay.com/docs/learn/refund-policy/) - [How to Save Tokens on the Claude API — Practical Tips](https://llmsrelay.com/docs/learn/save-tokens-claude-api/) - [LLMsRelay Regional Availability and API Access](https://llmsrelay.com/docs/learn/supported-countries/) - [Why Choose LLMsRelay](https://llmsrelay.com/docs/learn/why-llmsrelay/) ## Full documentation content ### GLM 5.3 Uncensored: Messages & Chat Completions Canonical URL: https://llmsrelay.com/docs/guides/glm-uncensored/ # GLM 5.3 Uncensored: Messages & Chat Completions Configure both GLM 5.3 models on LLMsRelay: API keys, cash balance, cURL, Python, TypeScript, streaming, web search and tool calling. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** Create a GLM 5.3 key and fund the separate GLM USD wallet. Use https://api.llmsrelay.com for the Anthropic SDK or https://api.llmsrelay.com/v1 for the OpenAI SDK. Both public GLM models support Messages and Chat Completions; the 1M model accepts text only. ## 1\. Account, wallet and key 1. Sign in to LLMsRelay. If you do not have an account, register first, then return to Billing > Uncensored. 2. Fund the separate GLM cash wallet: $45, $100, $500 or $1,000. Funding is 1:1 USD, with no signup, referral or promotional bonuses. Standard API-equivalent packs do not fund this wallet. 3. Open API Keys, create a key and select GLM 5.3. Copy the secret once and keep it local. Use the full dashboard-issued sk- key, not a key from another service. 4. Check both the cash balance and the key's remaining allowance. A positive standard balance is not sufficient. A key has one group; use separate keys for simultaneous Claude, GPT and GLM access. [Fund GLM wallet](/dashboard/billing/?flow=uncensored&lang=en)[Open API Keys](/dashboard/keys/?lang=en) ## 2\. Choose a public model ID | Public model ID | Context (tokens) | Input | | --- | --- | --- | | glm-5-3-uncensored | 262,144 | Text and images | | glm-5-3-uncensored-1m | 1,000,000 | Text only | GET /v1/modelsbash ``` curl https://api.llmsrelay.com/v1/models \ -H "Authorization: Bearer $LLMSRELAY_API_KEY" ``` ## 3\. Store your credentials Set LLMSRELAY\_API\_KEY in your local shell or secret manager. The examples use a placeholder; replace it locally, never in a shared chat, URL, browser bundle or repository. Do not add /v1 to the Anthropic SDK base URL; the OpenAI SDK needs /v1 exactly once. Environmentbash ``` export LLMSRELAY_API_KEY="YOUR_LLMSRELAY_KEY" ``` ## 4\. Anthropic-compatible Messages Messages uses x-api-key and anthropic-version. max\_tokens is required. Read text blocks from response.content, not choices. The system instruction belongs in the top-level system field. To use the 1M model, replace only model with glm-5-3-uncensored-1m. POST /v1/messagesbash ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: $LLMSRELAY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_tokens": 512, "system": "You are a concise assistant.", "messages": [ { "role": "user", "content": "Explain what an API gateway does." } ] }' ``` ## 5\. OpenAI-compatible Chat Completions Chat Completions uses Authorization: Bearer. Put system instructions in a system message and read choices\[0\].message.content. Use the exact public model ID. GET /v1/models lists the models accessible to a valid, funded key and may return 402 when no effective credit remains. POST /v1/chat/completionsbash ``` curl https://api.llmsrelay.com/v1/chat/completions \ -H "Authorization: Bearer $LLMSRELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_tokens": 512, "messages": [ { "role": "system", "content": "You are a concise assistant." }, { "role": "user", "content": "Explain what an API gateway does." } ] }' ``` ## 6\. Python and TypeScript SDKs Install only the SDK you need. Keep the key in the server environment. The Anthropic SDK returns content blocks; the OpenAI SDK returns choices. These examples also work with the 1M public ID after changing model. Pythonbash ``` pip install anthropic openai ``` Python: Messagespython ``` import os from anthropic import Anthropic client = Anthropic( api_key=os.environ["LLMSRELAY_API_KEY"], base_url="https://api.llmsrelay.com", ) response = client.messages.create( model="glm-5-3-uncensored", max_tokens=512, messages=[{"role": "user", "content": "Hello!"}], ) print("".join(block.text for block in response.content if block.type == "text")) ``` Python: Chat Completionspython ``` import os from openai import OpenAI client = OpenAI( api_key=os.environ["LLMSRELAY_API_KEY"], base_url="https://api.llmsrelay.com/v1", ) response = client.chat.completions.create( model="glm-5-3-uncensored", max_tokens=512, messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` Node.jsbash ``` npm install @anthropic-ai/sdk openai ``` TypeScript: Messagestypescript ``` import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ apiKey: process.env.LLMSRELAY_API_KEY, baseURL: "https://api.llmsrelay.com", }); const response = await client.messages.create({ model: "glm-5-3-uncensored", max_tokens: 512, messages: [{ role: "user", content: "Hello!" }], }); for (const block of response.content) { if (block.type === "text") console.log(block.text); } ``` TypeScript: Chat Completionstypescript ``` import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.LLMSRELAY_API_KEY, baseURL: "https://api.llmsrelay.com/v1", }); const response = await client.chat.completions.create({ model: "glm-5-3-uncensored", max_tokens: 512, messages: [{ role: "user", content: "Hello!" }], }); console.log(response.choices[0].message.content); ``` ## 7\. Claude Code and compatible IDEs For Claude Code, use a dedicated GLM profile or shell with the variables below. Do not silently overwrite existing profiles. Remove conflicting ANTHROPIC\_AUTH\_TOKEN credentials in that profile. Map default and fast models to GLM so the client does not request Claude with a GLM key. For Cline, Roo Code or Continue, choose OpenAI Compatible, base URL https://api.llmsrelay.com/v1, your GLM key and one of the public IDs. Client tool compatibility varies; test a short request before a long session. Claude Code: GLMbash ``` export ANTHROPIC_BASE_URL="https://api.llmsrelay.com" export ANTHROPIC_API_KEY="$LLMSRELAY_API_KEY" export ANTHROPIC_MODEL="glm-5-3-uncensored" export ANTHROPIC_DEFAULT_OPUS_MODEL="glm-5-3-uncensored" export ANTHROPIC_DEFAULT_SONNET_MODEL="glm-5-3-uncensored" export ANTHROPIC_DEFAULT_HAIKU_MODEL="glm-5-3-uncensored" export ANTHROPIC_SMALL_FAST_MODEL="glm-5-3-uncensored" # Only in this dedicated GLM shell/profile: unset ANTHROPIC_AUTH_TOKEN claude ``` ## 8\. Streaming (SSE) Add stream:true to either request and use curl -N. Chat Completions emits data chunks: accumulate choices\[\].delta.content and tool\_calls arguments until \[DONE\]. stream\_options.include\_usage:true requests a final usage chunk; its choices may be empty. Messages emits named events including message\_start, content\_block\_delta, message\_delta and message\_stop. Handle text, thinking and tool JSON deltas separately. Inspect error events even if HTTP status is 200; an opened stream is not proof of a completed response. Never concatenate raw SSE as plain JSON. Chat Completions: SSEbash ``` curl -N https://api.llmsrelay.com/v1/chat/completions \ -H "Authorization: Bearer $LLMSRELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_tokens": 512, "messages": [ { "role": "system", "content": "You are a concise assistant." }, { "role": "user", "content": "Explain what an API gateway does." } ], "stream": true, "stream_options": { "include_usage": true } }' ``` Messages: SSEbash ``` curl -N https://api.llmsrelay.com/v1/messages \ -H "x-api-key: $LLMSRELAY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_tokens": 512, "system": "You are a concise assistant.", "messages": [ { "role": "user", "content": "Explain what an API gateway does." } ], "stream": true }' ``` ## 9\. Built-in web search Use Messages web\_search\_2025\_03\_05 or Responses web\_search. Chat web\_search\_options is temporarily unavailable because exact search usage is not reported. No local search function is needed. Preserve returned citation metadata, URLs and server tool result blocks. A search request can produce several tool events before the answer. Verify that a search actually ran and inspect sources; a fluent answer alone is not evidence of a search. Each actual search costs the upstream price of $0.008, plus input/output/cache tokens. Messages max\_uses and Responses max\_tool\_calls default to 3; use 1-10. Calls can perform several searches. Domain restrictions: Messages allowed\_domains/blocked\_domains; Responses filters.allowed\_domains. Web fetch is Messages-only: web\_fetch\_2025\_03\_05. Retrieved content consumes tokens; no separate fetch fee currently. Tool/image workloads reserve up to the model context ceiling plus output/search fees; unused reservation is released on settlement. Preserve sources for display, but remove server\_tool\_use, web\_search\_tool\_result and web\_fetch\_tool\_result blocks when replaying Messages history. Responses: web searchbash ``` curl https://api.llmsrelay.com/v1/responses \ -H "Authorization: Bearer $LLMSRELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_output_tokens": 2048, "max_tool_calls": 3, "tools": [ { "type": "web_search" } ], "input": "Search the web for recent Python releases. Cite your sources." }' ``` Messages: web searchbash ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: $LLMSRELAY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_tokens": 2048, "system": "You are a concise assistant.", "messages": [ { "role": "user", "content": "Search the web for recent Python releases. Cite your sources." } ], "tools": [ { "type": "web_search_2025_03_05", "name": "web_search", "max_uses": 3 } ] }' ``` Messages: web fetchbash ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: $LLMSRELAY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_tokens": 2048, "system": "You are a concise assistant.", "messages": [ { "role": "user", "content": "Fetch https://www.iana.org/help/example-domains and summarize it." } ], "tools": [ { "type": "web_fetch_2025_03_05", "name": "web_fetch", "max_uses": 1 } ] }' ``` ## 10\. Custom function tools Define the function's JSON Schema and let the model request it. Your server validates arguments, performs the allowed operation, then sends its result back. Never execute arbitrary shell commands or untrusted arguments from a model. Chat: append the assistant message containing tool\_calls, then one role:tool message for each tool\_call\_id. Messages: append the complete assistant content, then a user message with tool\_result blocks matching tool\_use\_id. Preserve reasoning/signature blocks when present. Continue until the model returns text; handle multiple calls and impose a loop limit. The examples below show the follow-up shape, not executable weather services. Replace call IDs and full assistant payloads with the actual response and substitute a real, validated local result. In SSE, collect all argument fragments before parsing JSON. Chat Completions: functionbash ``` curl https://api.llmsrelay.com/v1/chat/completions \ -H "Authorization: Bearer $LLMSRELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_tokens": 512, "messages": [ { "role": "user", "content": "What is the weather in Paris?" } ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Get current weather for a city", "parameters": { "type": "object", "properties": { "city": { "type": "string" } }, "required": [ "city" ] } } } ] }' ``` Chat Completions: continuationtext ``` # messages for the next Chat request; use the ACTUAL assistant message and call ID: [ {"role":"user","content":"What is the weather in Paris?"}, {"role":"assistant","content":null,"tool_calls":[ {"id":"ACTUAL_CALL_ID","type":"function","function":{"name":"get_weather","arguments":"{\"city\":\"Paris\"}"}} ]}, {"role":"tool","tool_call_id":"ACTUAL_CALL_ID","content":"{\"temperature_c\":18}"} ] ``` Messages: functionbash ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: $LLMSRELAY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_tokens": 512, "system": "You are a concise assistant.", "messages": [ { "role": "user", "content": "What is the weather in Paris?" } ], "tools": [ { "name": "get_weather", "description": "Get current weather for a city", "input_schema": { "type": "object", "properties": { "city": { "type": "string" } }, "required": [ "city" ] } } ] }' ``` Messages: continuationtext ``` # messages for the next Messages request; preserve the ACTUAL full assistant content: [ {"role":"user","content":"What is the weather in Paris?"}, {"role":"assistant","content":[ {"type":"tool_use","id":"ACTUAL_TOOL_USE_ID","name":"get_weather","input":{"city":"Paris"}} ]}, {"role":"user","content":[ {"type":"tool_result","tool_use_id":"ACTUAL_TOOL_USE_ID","content":"{\"temperature_c\":18}"} ]} ] ``` ## 11\. Reasoning and structured output For Chat use reasoning\_effort (for example high); for Messages use output\_config.effort or thinking.budget\_tokens. Hiding reasoning with include\_reasoning:false does not eliminate reasoning cost. Budget enough output tokens for reasoning plus the answer. Chat JSON Schema is shown below; validate the returned JSON in your application. Do not assume every SDK accepts nonstandard fields: use extra\_body in Python or raw HTTP. Chat Completions: JSON Schemabash ``` curl https://api.llmsrelay.com/v1/chat/completions \ -H "Authorization: Bearer $LLMSRELAY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_tokens": 2048, "messages": [ { "role": "user", "content": "Return a short greeting as JSON." } ], "reasoning_effort": "high", "response_format": { "type": "json_schema", "json_schema": { "name": "greeting", "strict": true, "schema": { "type": "object", "properties": { "greeting": { "type": "string" } }, "required": [ "greeting" ], "additionalProperties": false } } } }' ``` ## 12\. Token counting, caching and images Use POST /v1/messages/count\_tokens with the same model and input before a long request. It estimates input tokens, not the eventual output or final charge. Keep the full input and output within the model's context window. Keep stable instructions and tool schemas at the beginning of the prompt. Cache reuse is best-effort, not guaranteed on every repeat. Inspect usage for cache reads rather than assuming a hit. Cache creation is billed at the input rate; reasoning counts toward output. Only glm-5-3-uncensored accepts images: PNG, JPEG, WebP or GIF, up to 12 MB each. Chat uses image\_url, Messages uses image source URL/base64, Responses uses input\_image. Images are input-token billed, not billed as base64 text. The 1M model rejects images; video is unsupported. Caching is automatic and best-effort on an unchanged prefix in the same model/project. prompt\_cache\_key is supported; prompt\_cache\_retention is not. cache\_control/ttl hints do not guarantee retention. OpenAI input totals include cache reads/writes; Messages input\_tokens excludes them. created\_cache\_tokens counts as cache write, not a second input charge. POST /v1/messages/count\_tokensbash ``` curl https://api.llmsrelay.com/v1/messages/count_tokens \ -H "x-api-key: $LLMSRELAY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "messages": [ { "role": "user", "content": "Explain what an API gateway does." } ] }' ``` Messages: imagebash ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: $LLMSRELAY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "glm-5-3-uncensored", "max_tokens": 512, "system": "You are a concise assistant.", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image." }, { "type": "image", "source": { "type": "url", "url": "https://YOUR_HOST/image.png" } } ] } ] }' ``` ## 13\. Troubleshooting | Status / symptom | Action | | --- | --- | | 401 | Check the full LLMsRelay secret, header and environment variable; rotate exposed keys. | | 403 / model unavailable | Select GLM 5.3 and check the key's allowed models. A Basic or Codex key cannot access GLM. | | 402 / insufficient\_credits | Check the GLM cash wallet AND key allowance. Funds must cover the estimated reservation, including max\_tokens; reduce the output budget or add funds. | | 400 | Check model ID, JSON, max\_tokens and tool schema. Remove images for the 1M model. | | 404 / wrong route | Messages: /v1/messages. Chat: /v1/chat/completions. Remove duplicate /v1. | | 429 | Respect Retry-After when present; use bounded exponential backoff and reduce concurrency. | | 5xx / SSE error | Record the request ID, inspect the terminal stream event and contact support if repeated. A retry is a new request and may incur usage. | | Wallet unavailable | Do not pay repeatedly or assume the balance is zero. Wait for balance verification or contact support. | ## 14\. Responses compatibility notes Prefer Messages or Chat Completions for this guide. Native Codex uses Responses; configuring an OpenAI-compatible IDE is not the same as configuring Codex CLI. GLM Responses is stateless: previous\_response\_id is unsupported. Replay prior input/output and function\_call\_output instead. In checks on October 9, 2026, base-model Responses web-search SSE failed, while non-streaming search completed. Prefer Messages or non-streaming Responses for search. The 1M Responses search returned source URLs but omitted final citation annotations. Avoid include\_reasoning:false with JSON Schema: checks returned empty content. These are compatibility observations, not a guarantee for every future request. Keep error handling and verify your exact client. Full-context load tests and all third-party clients are not covered. ## 15\. Final USD token prices USD per 1 million tokens, final customer rates. No additional multiplier is applied to this table. Actual input, output and cached tokens are charged to the separate GLM cash wallet; there are no bonuses in this flow. | Model | Input | Output | Cache read | Cache write | | --- | --- | --- | --- | --- | | glm-5-3-uncensored | $4.00 | $12.00 | $0.40 | $4.00 | | glm-5-3-uncensored-1m | $12.00 | $20.00 | $1.20 | $12.00 | Fund GLM wallet Fund the separate GLM cash wallet: $45, $100, $500 or $1,000. Funding is 1:1 USD, with no signup, referral or promotional bonuses. Standard API-equivalent packs do not fund this wallet. [Fund GLM wallet](/dashboard/billing/?flow=uncensored&lang=en) [PreviousAPI Key Configuration](/docs/guides/api-key-configuration/)[Next Best Practices](/docs/guides/best-practices/) --- ### Introduction Canonical URL: https://llmsrelay.com/docs/getting-started/introduction/ # Introduction Set up the LLMsRelay developer gateway with Anthropic-compatible and OpenAI-compatible API formats. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Base URL:** https://api.llmsrelay.com - **Models:** Returned by GET /v1/models for your key - **API formats:** Anthropic Messages + OpenAI Chat Completions - **Auth:** API3 key - **Setup time:** ~2 minutes ## What the LLMsRelay Gateway Provides The LLMsRelay gateway gives you compatible API access through one API key from the dashboard. It works with both the **Anthropic Messages API** and the **OpenAI Chat Completions API**, so you can plug it into SDKs, editors, agents, and automation tools that support a custom base URL. | Area | What you get | | --- | --- | | Anthropic-compatible | POST /v1/messages with Anthropic-style request and SSE streaming | | OpenAI-compatible | POST /v1/chat/completions and GET /v1/models | | Model IDs | Returned by GET /v1/models for the authenticated key | | Tooling | Works with Cursor, Claude Code, VS Code extensions, SDKs, and cURL | | Billing | Platform usage is deducted according to the published rate card | | Key controls | Allowed models, per-key rate limit (requests/min), and monthly credit cap | ## Recommended Model IDs These are the stable IDs we recommend in docs and clients: - `claude-opus-5-5` — default starting point for complex work - `claude-fable-5-1` — high-capability coding and agent workflows - `claude-opus-4.8` — newest Opus model for long, complex coding and agent sessions - `claude-opus-4.7` - `claude-sonnet-4.6` — lower-cost daily workloads - `claude-haiku-4.5` Use the IDs returned by GET /v1/models for new setups. Do not depend on an assumed catalogue in production. ## Short Aliases For quick CLI experiments you can use short aliases that always resolve to the latest version of each tier: - `opus` → `claude-opus-5-5` - `sonnet` → `claude-sonnet-4.6` - `haiku` → `claude-haiku-4.5` For production use pinned IDs — aliases will silently upgrade across major versions. See [Models](/docs/getting-started/models/) for full table. ## Base URL Rules - **Anthropic SDK / raw Anthropic API**: use `https://api.llmsrelay.com/v1` - **OpenAI SDK / Cursor / VS Code OpenAI-compatible tools**: use `https://api.llmsrelay.com/v1` - **Claude Code**: use `https://api.llmsrelay.com` for `ANTHROPIC_BASE_URL` (without `/v1`) ## Quick Links [Quickstart — First request with Anthropic and OpenAI formats](/docs/getting-started/quickstart)[Models — Opus 4.8, Opus 4.7, aliases, context sizes, rates](/docs/getting-started/models)[Authentication — API key usage and per-key controls](/docs/getting-started/authentication)[IDE Integration — Cursor, Claude Code, VS Code extensions, OpenCode](/docs/guides/ide-integration)[Pricing — Direct per-model input / cache / output rates](/docs/billing/pricing) Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousEnterprise AI API](/enterprise/)[Next Models](/docs/getting-started/models/) --- ### Quickstart Canonical URL: https://llmsrelay.com/docs/getting-started/quickstart/ # Quickstart Send your first LLMsRelay API request with Anthropic-compatible or OpenAI-compatible formats. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Base URL:** https://api.llmsrelay.com - **Endpoint:** POST /v1/messages - **Auth header:** x-api-key: YOUR_KEY - **Default model:** claude-opus-5-5 - **Time to first call:** ~2 minutes 01 Create a key From your LLMsRelay dashboard 02 Send a request Use Messages or Chat Completions 03 Connect your tools Cursor, Claude Code, or your SDK ## 1\. Get Your API Key Sign up at [llmsrelay.com](https://llmsrelay.com) and create an API key from your dashboard. Your key will look like `sk-cs4-...`. ## Validate gateway access Before sending real requests, run two cheap checks: a health probe and a models list. They confirm the base URL, your network, and (for /v1/models) your API key — without spending any credits. ### 1\. Liveness probe — /healthz Public, unauthenticated. Returns 200 and a JSON status payload when the gateway is up. Use it from CI, monitoring, or any shell. curl (status)curl (body)PythonNode.js ``` curl -sS -o /dev/null -w "HTTP %{http_code}\n" \ https://api.llmsrelay.com/healthz ``` ### 2\. Auth + catalogue — /v1/models Returns the list of model IDs you can use with the messages API. Requires a valid x-api-key. If this works, your key, base URL, and headers are all correct. curlPythonNode.js ``` curl -sS https://api.llmsrelay.com/v1/models \ -H "x-api-key: $ANTHROPIC_API_KEY" | jq '.data[].id' ``` What a healthy response looks like - /healthz → HTTP 200 with {"status":"ok", ...}. Anything else (timeout, 5xx, HTML page) means you hit the wrong host or there is a network/proxy issue. - /v1/models → HTTP 200 with {"object":"list","data":\[...\]}. The data array contains the only model IDs accepted by /v1/messages. Common mistakes - Wrong base URL. The only correct host is https://api.llmsrelay.com. Do not use api.anthropic.com, api.llmsrelay.com (old) or any other variant. - Missing /v1 prefix. Endpoints live under /v1/\* (e.g. /v1/messages, /v1/models). /healthz is the one exception — it is at the root. - Wrong auth header. Use x-api-key: sk-cs4-... (Anthropic-style). Do not send Authorization: Bearer ... unless the SDK builds it for you. - Invalid model ID. Only IDs returned by /v1/models work. The public catalogue includes claude-opus-5-5, claude-fable-5-1, claude-opus-5, claude-fable-5, claude-opus-4.8, claude-opus-4.7, claude-opus-4.6, claude-sonnet-5, claude-sonnet-4.6, and claude-haiku-4.5, subject to key filtering. Tip: use /v1/models as the source of truth for model IDs in your scripts and CI — never hard-code a list. ## 2\. Send Your First Request cURL · MessagescURL · OpenAI ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: YOUR_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-opus-5-5", "max_tokens": 1024, "messages": [ {"role": "user", "content": "What is the capital of France?"} ] }' ``` ## 3\. Set Up IDE Integration ### Cursor Cursor Pro Plan Required Custom API providers only work on Cursor's **paid Pro plan** or higher. The pay-as-you-go starter ($5) does not support third-party API keys. In Cursor Settings → Models → Add Provider: - Base URL: `https://api.llmsrelay.com/v1` - API Key: your LLMsRelay API key - Add a model ID returned by `GET /v1/models`, for example `claude-opus-5-5` ### Claude Code Claude Code CLI uses `api.llmsrelay.com` with `sk-cs4-*` keys. Use `ANTHROPIC_API_KEY` only — never `ANTHROPIC_AUTH_TOKEN`. Linux / macOS: Set environment variablesbash ``` export ANTHROPIC_BASE_URL="https://api.llmsrelay.com" export ANTHROPIC_API_KEY="YOUR_SK_CS4_KEY" export ANTHROPIC_MODEL="claude-opus-5-5" export ANTHROPIC_SMALL_FAST_MODEL="claude-haiku-4.5" ``` Windows PowerShell: Set environment variables (PowerShell)powershell ``` $env:ANTHROPIC_BASE_URL = "https://api.llmsrelay.com" $env:ANTHROPIC_API_KEY = "YOUR_SK_CS4_KEY" $env:ANTHROPIC_MODEL = "claude-opus-5-5" $env:ANTHROPIC_SMALL_FAST_MODEL = "claude-haiku-4.5" ``` Important For Claude Code, the base URL must be set **without** `/v1`. The CLI appends `/v1/messages` automatically. See the dedicated [Claude Code guide](/docs/guides/claude-code/) for the full `~/.claude/settings.json` template and verification steps. ## 4\. List Available Models List the model IDs currently available to your key: ``` curl https://api.llmsrelay.com/v1/models \ -H "Authorization: Bearer YOUR_API_KEY" ``` ## Free model API offers See the current free API offer pages for supported model IDs and setup examples: - [Claude Opus 5.5 API free access](/claude-opus-5-5-free-api/) - [Claude Fable 5.1 API free access](/claude-fable-5-1-free-api/) - [GPT-6 Astra API free usage](/gpt-6-astra-free-api/) ## 5\. Billing Basics - Choose a one-time platform usage pack. Current prices and credited usage are shown on the Plans page. - You are billed per token (input + output) based on the model used - Check your balance anytime in the [dashboard](https://llmsrelay.com) - Purchased usage remains available until it is consumed See [Pricing](/docs/billing/pricing/) for full per-model token rates. ## Next Steps - Explore the [API Reference](/docs/api-reference/overview/) - Set up [Streaming](/docs/api-reference/streaming/) - Learn about [Pricing](/docs/billing/pricing/) - Configure [IDE Integrations](/docs/guides/ide-integration/) Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousModels](/docs/getting-started/models/)[Next Authentication](/docs/getting-started/authentication/) --- ### Models Canonical URL: https://llmsrelay.com/docs/getting-started/models/ # Models Current Claude models on LLMsRelay — Opus 5.5, Fable 5.1, Opus 5, Fable 5, Opus 4.8, Opus 4.7, Sonnet 5, Sonnet 4.6 and Haiku 4.5. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. llmsrelay --pricingup to −91% $ Pay once. Get ~11× the Anthropic balance. Same Anthropic per-token rates. Massive discount on the top-up. | you pay | Anthropic-equivalent balance | discount | | --- | --- | --- | | $45 | $500balance | −91% | | $90popular | $1,000balance | −91% | Tokens are billed at Anthropic's exact per-token rates · Balance never expires · No subscription ## Recommended Model IDs Use the stable model IDs listed below in your API requests. These IDs are consistent across both Anthropic and OpenAI-compatible endpoints. | Model | Model ID | Context | Thinking | Key access | | --- | --- | --- | --- | --- | | Claude Opus 5.5 | claude-opus-5-5 | 1M | Yes | Claude key groups | | Claude Fable 5.1 | claude-fable-5-1 | 1M | Yes | Claude key groups | | Claude Opus 5 | claude-opus-5 | 1M | Yes | Basic, Pro | | Claude Fable 5 | claude-fable-5 | 1M | Yes | Basic, Pro | | Claude Opus 4.8 | claude-opus-4.8 | 1M | Yes | Claude key groups | | Claude Opus 4.7 | claude-opus-4.7 | 1M | Yes | Claude key groups | | Claude Opus 4.6 | claude-opus-4.6 | 1M | Yes | Claude key groups | | Claude Sonnet 5 | claude-sonnet-5 | 1M | Yes | Claude key groups | | Claude Sonnet 4.6 | claude-sonnet-4.6 | 1M | Yes | Claude key groups | | Claude Haiku 4.5 | claude-haiku-4.5 | 200K | Yes | Claude key groups | The catalog can be filtered by API-key group and per-key allowed-model settings. Use the response from `GET /v1/models` before configuring a client. ## Short Aliases For quick experimentation you can use short aliases instead of pinned IDs. Aliases always resolve to the latest version of each tier — convenient for chat clients, risky for production. | Alias | Resolves to | Recommended for | | --- | --- | --- | | opus | claude-opus-5-5 | Quick CLI calls, sandboxes | | sonnet | claude-sonnet-4.6 | Quick CLI calls, sandboxes | | haiku | claude-haiku-4.5 | Quick CLI calls, sandboxes | For production deployments use pinned IDs. Aliases may resolve to a different model after a catalogue update, which can change behaviour, output length, and pricing without warning. ## Long Context (1M tokens) LLMsRelay exposes a **1 million token context window** for Opus 5.5, Fable 5.1, Opus 5, Fable 5, the Opus 4.x models, and Sonnet 5/4.6. Haiku 4.5 has a 200K context window. Enable the 1M context mode by sending the `anthropic-beta` header on each request: Enabling 1M contextbash ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: YOUR_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "anthropic-beta: context-1m-2025" \ -H "content-type: application/json" \ -d '{ "model": "claude-opus-5", "max_tokens": 4096, "messages": [{"role": "user", "content": "."}] }' ``` 1M context is billed at the same per-token rate as standard context — the cap is just larger. Useful for whole-codebase analysis, long-form research, and very large RAG payloads. ## LLMsRelay Platform Rates These are the platform per-token rates used for usage deductions. Cache write is billed at 1.25× input and cache read at 0.10× input. | Model | Input / 1M | Output / 1M | Cache write / 1M | Cache read / 1M | | --- | --- | --- | --- | --- | | Claude Opus 5.5 | $5.00 | $25.00 | $6.25 | $0.50 | | Claude Fable 5.1 | $10.00 | $50.00 | $12.50 | $1.00 | | Claude Opus 5 | $5.00 | $25.00 | $6.25 | $0.50 | | Claude Fable 5 | $10.00 | $50.00 | $12.50 | $1.00 | | Claude Opus 4.8 | $5.00 | $25.00 | $6.25 | $0.50 | | Claude Opus 4.7 | $5.00 | $25.00 | $6.25 | $0.50 | | Claude Opus 4.6 | $5.00 | $25.00 | $6.25 | $0.50 | | Claude Sonnet 5 | $2.00 | $10.00 | $2.50 | $0.20 | | Claude Sonnet 4.6 | $3.00 | $15.00 | $3.75 | $0.30 | | Claude Haiku 4.5 | $1.00 | $5.00 | $1.25 | $0.10 | Rates are shown per million tokens. The exact balance deduction also depends on the measured request usage and cache token categories. ## Choosing a Model - **Opus 5.5** — default starting point for highest-capability reasoning and agent workloads - **Fable 5.1** — high-capability coding and agent workloads - **Opus 4.8 / 4.7 / 4.6** — complex reasoning, research, and long coding sessions - **Sonnet 5 / 4.6** — general coding and daily development - **Haiku 4.5** — fast, lower-cost classification and automation See [Key Controls](/docs/guides/key-controls/) to restrict which models a specific API key can use. ## Listing Models via API You can retrieve the list of available models programmatically: GET https://api.llmsrelay.com/v1/models The response is filtered for the authenticated key and includes the model capabilities exposed by the gateway. Pricing rates are documented on the [Pricing](/docs/billing/pricing/) page. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousIntroduction](/docs/getting-started/introduction/)[Next Quickstart](/docs/getting-started/quickstart/) --- ### Authentication Canonical URL: https://llmsrelay.com/docs/getting-started/authentication/ # Authentication How to authenticate Claude API requests. Set up your API key, configure headers, and secure your integration. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Key prefix:** sk-cs4-* - **Anthropic header:** x-api-key: sk-cs4-... - **OpenAI header:** Authorization: Bearer sk-cs4-... - **Base URL:** https://api.llmsrelay.com - **Rotation:** Anytime from dashboard, no downtime ## API Key Format For the new gateway, LLMsRelay issues one API key family: - `sk-cs4-…` — keys for `api.llmsrelay.com`. Use them with Anthropic-compatible clients (Claude Code, Anthropic SDK, Continue, OpenCode) and OpenAI-compatible clients (Cursor, Cline, Roo Code, OpenAI SDK). Keep the key secret — never expose it in client-side code or public repositories. ## Authentication Methods The gateway supports two authentication shapes, depending on which API format you call: ### Anthropic Format Pass your API key in the `x-api-key` header along with the required `anthropic-version` header: ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: YOUR_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{"model": "claude-opus-5-5", "max_tokens": 1024, "messages": [{"role": "user", "content": "Hello"}]}' ``` ### OpenAI-Compatible Format Pass your API key as a Bearer token in the `Authorization` header: ``` curl https://api.llmsrelay.com/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "content-type: application/json" \ -d '{"model": "claude-opus-5-5", "messages": [{"role": "user", "content": "Hello"}]}' ``` ### Claude Code CLI / Extension The Claude Code CLI and VS Code extension authenticate via the `ANTHROPIC_API_KEY` environment variable. **Use `ANTHROPIC_API_KEY` only** — `ANTHROPIC_AUTH_TOKEN` is reserved for Anthropic's OAuth login (Claude Pro/Max) and will not authenticate against this gateway. Claude Code env varsbash ``` export ANTHROPIC_BASE_URL="https://api.llmsrelay.com" export ANTHROPIC_API_KEY="YOUR_SK_CS4_KEY" ``` ## Per-Key Controls Each API key can be configured with the following enforced controls in the dashboard: | Setting | Description | | --- | --- | | Rate limit | Max requests per minute for this key | | Allowed models | Restrict key to specific models (e.g. only claude-sonnet-4) | | Monthly credit cap | Hard cap on credits this key can spend within a UTC calendar month | | Revocation | Instantly disable a compromised key without affecting others | Per-key controls are configured in the LLMsRelay dashboard under **API Keys → Edit Key**. All settings are optional — by default keys have full access (no model restriction, no monthly cap), bounded only by your remaining credit balance. ## SDK Configuration Both the Anthropic and OpenAI SDKs accept the API key directly. See the [SDK guides](/docs/sdks/python/) for language-specific examples. ## Verify your setup ## Validate gateway access Before sending real requests, run two cheap checks: a health probe and a models list. They confirm the base URL, your network, and (for /v1/models) your API key — without spending any credits. ### 1\. Liveness probe — /healthz Public, unauthenticated. Returns 200 and a JSON status payload when the gateway is up. Use it from CI, monitoring, or any shell. curl (status)curl (body)PythonNode.js ``` curl -sS -o /dev/null -w "HTTP %{http_code}\n" \ https://api.llmsrelay.com/healthz ``` ### 2\. Auth + catalogue — /v1/models Returns the list of model IDs you can use with the messages API. Requires a valid x-api-key. If this works, your key, base URL, and headers are all correct. curlPythonNode.js ``` curl -sS https://api.llmsrelay.com/v1/models \ -H "x-api-key: $ANTHROPIC_API_KEY" | jq '.data[].id' ``` What a healthy response looks like - /healthz → HTTP 200 with {"status":"ok", ...}. Anything else (timeout, 5xx, HTML page) means you hit the wrong host or there is a network/proxy issue. - /v1/models → HTTP 200 with {"object":"list","data":\[...\]}. The data array contains the only model IDs accepted by /v1/messages. Common mistakes - Wrong base URL. The only correct host is https://api.llmsrelay.com. Do not use api.anthropic.com, api.llmsrelay.com (old) or any other variant. - Missing /v1 prefix. Endpoints live under /v1/\* (e.g. /v1/messages, /v1/models). /healthz is the one exception — it is at the root. - Wrong auth header. Use x-api-key: sk-cs4-... (Anthropic-style). Do not send Authorization: Bearer ... unless the SDK builds it for you. - Invalid model ID. Only IDs returned by /v1/models work. The public catalogue includes claude-opus-5-5, claude-fable-5-1, claude-opus-5, claude-fable-5, claude-opus-4.8, claude-opus-4.7, claude-opus-4.6, claude-sonnet-5, claude-sonnet-4.6, and claude-haiku-4.5, subject to key filtering. Tip: use /v1/models as the source of truth for model IDs in your scripts and CI — never hard-code a list. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousQuickstart](/docs/getting-started/quickstart/)[Next Overview](/docs/api-reference/overview/) --- ### API Reference Overview Canonical URL: https://llmsrelay.com/docs/api-reference/overview/ # API Reference Overview Claude API reference — endpoints, headers, request format. Anthropic-compatible and OpenAI-compatible modes. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Base URL:** https://api.llmsrelay.com - **Anthropic endpoint:** POST /v1/messages - **OpenAI endpoint:** POST /v1/chat/completions - **Streaming:** SSE on both endpoints - **Auth:** sk-cs4-* keys ## Base URL https://api.llmsrelay.com ## Endpoints | Method | Path | Description | | --- | --- | --- | | POST | /v1/messages | Send a message (Anthropic format) | | POST | /v1/chat/completions | Send a message (OpenAI format) | | GET | /v1/models | List available models | | GET | /health | Health check endpoint | ## Required Headers | Header | Required For | Value | | --- | --- | --- | | x-api-key | Anthropic format | Your API key | | anthropic-version | Anthropic format | 2023-06-01 | | Authorization | OpenAI format | Bearer YOUR\_API\_KEY | | content-type | All POST requests | application/json | ## Recommended Model IDs Use these stable model IDs in your requests: - `claude-opus-5-5`, `claude-fable-5-1` — newest high-capability models - `claude-opus-4.8`, `claude-opus-4.7`, `claude-opus-4.6` - `claude-sonnet-5`, `claude-sonnet-4.6` — lower-cost daily workloads - `claude-haiku-4.5` — fast, lower-cost automation Or use short aliases that always point to the latest version of each tier: - `opus` → `claude-opus-5-5` - `sonnet` → `claude-sonnet-4.6` - `haiku` → `claude-haiku-4.5` Use `GET /v1/models` to retrieve the current list of available models programmatically. Aliases resolve server-side before model whitelisting and billing. ## Smoke-test the gateway ## Validate gateway access Before sending real requests, run two cheap checks: a health probe and a models list. They confirm the base URL, your network, and (for /v1/models) your API key — without spending any credits. ### 1\. Liveness probe — /healthz Public, unauthenticated. Returns 200 and a JSON status payload when the gateway is up. Use it from CI, monitoring, or any shell. curl (status)curl (body)PythonNode.js ``` curl -sS -o /dev/null -w "HTTP %{http_code}\n" \ https://api.llmsrelay.com/healthz ``` ### 2\. Auth + catalogue — /v1/models Returns the list of model IDs you can use with the messages API. Requires a valid x-api-key. If this works, your key, base URL, and headers are all correct. curlPythonNode.js ``` curl -sS https://api.llmsrelay.com/v1/models \ -H "x-api-key: $ANTHROPIC_API_KEY" | jq '.data[].id' ``` What a healthy response looks like - /healthz → HTTP 200 with {"status":"ok", ...}. Anything else (timeout, 5xx, HTML page) means you hit the wrong host or there is a network/proxy issue. - /v1/models → HTTP 200 with {"object":"list","data":\[...\]}. The data array contains the only model IDs accepted by /v1/messages. Common mistakes - Wrong base URL. The only correct host is https://api.llmsrelay.com. Do not use api.anthropic.com, api.llmsrelay.com (old) or any other variant. - Missing /v1 prefix. Endpoints live under /v1/\* (e.g. /v1/messages, /v1/models). /healthz is the one exception — it is at the root. - Wrong auth header. Use x-api-key: sk-cs4-... (Anthropic-style). Do not send Authorization: Bearer ... unless the SDK builds it for you. - Invalid model ID. Only IDs returned by /v1/models work. The public catalogue includes claude-opus-5-5, claude-fable-5-1, claude-opus-5, claude-fable-5, claude-opus-4.8, claude-opus-4.7, claude-opus-4.6, claude-sonnet-5, claude-sonnet-4.6, and claude-haiku-4.5, subject to key filtering. Tip: use /v1/models as the source of truth for model IDs in your scripts and CI — never hard-code a list. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousAuthentication](/docs/getting-started/authentication/)[Next Messages](/docs/api-reference/messages/) --- ### Messages API Canonical URL: https://llmsrelay.com/docs/api-reference/messages/ # Messages API Send prompts to Claude Opus 5.5, Fable 5.1, Sonnet 4.6 and Haiku 4.5 via the Messages API. Code examples, request/response format, streaming support. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Method:** POST - **URL:** https://api.llmsrelay.com/v1/messages - **Required fields:** model, max_tokens, messages - **Streaming:** Set stream: true (SSE) - **Anthropic SDK:** 100% compatible ## Endpoint POST https://api.llmsrelay.com/v1/messages ## Request Body | Field | Type | Required | Description | | --- | --- | --- | --- | | model | string | Yes | Model ID (e.g., claude-opus-5-5) | | messages | array | Yes | Array of message objects with role and content | | max\_tokens | integer | Yes | Maximum tokens to generate | | system | string | No | System prompt | | stream | boolean | No | Enable streaming (default: false) | | top\_p | number | No | Nucleus sampling parameter | | top\_k | integer | No | Top-k sampling parameter | | stop\_sequences | array | No | Custom stop sequences | ## Example Request Messages API Requestbash ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: YOUR_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-opus-5-5", "max_tokens": 1024, "system": "You are a helpful assistant.", "messages": [ {"role": "user", "content": "Explain quantum computing in one paragraph."} ] }' ``` ## Response Example Responsejson ``` { "id": "msg_01XFDUDYJgAACzvnptvVoYEL", "type": "message", "role": "assistant", "content": [ { "type": "text", "text": "Quantum computing leverages quantum mechanical phenomena..." } ], "model": "claude-opus-5-5", "stop_reason": "end_turn", "usage": { "input_tokens": 25, "output_tokens": 150 } } ``` ## Streaming Set `"stream": true` in the request body to receive Server-Sent Events. See the [Streaming](/docs/api-reference/streaming/) page for details. ## Tool Use Tool use (function calling) is supported through the standard Anthropic `tools` parameter. Pass tool definitions in your request and Claude will generate structured tool calls when appropriate. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousOverview](/docs/api-reference/overview/)[Next Models](/docs/api-reference/models/) --- ### Models Endpoint Canonical URL: https://llmsrelay.com/docs/api-reference/models/ # Models Endpoint List the Claude API models available to an API key through the /v1/models endpoint. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Endpoint:** GET /v1/models - **Auth:** x-api-key: sk-cs4-* - **Public catalogue:** Opus 5.5, Fable 5.1, Opus 5, Fable 5, Sonnet 5, Sonnet 4.6, Haiku 4.5, GPT-6 Astra - **Aliases:** opus, sonnet, haiku ## List Models GET https://api.llmsrelay.com/v1/models ``` curl https://api.llmsrelay.com/v1/models \ -H "x-api-key: YOUR_API_KEY" ``` Responsejson ``` { "data": [ {"id": "claude-opus-5-5", "object": "model"}, {"id": "claude-fable-5-1", "object": "model"}, {"id": "claude-opus-5", "object": "model"}, {"id": "claude-fable-5", "object": "model"}, {"id": "claude-opus-4.8", "object": "model"}, {"id": "claude-opus-4.7", "object": "model"}, {"id": "claude-opus-4.6", "object": "model"}, {"id": "claude-sonnet-5", "object": "model"}, {"id": "claude-sonnet-4.6", "object": "model"}, {"id": "claude-haiku-4.5", "object": "model"}, {"id": "gpt-6-astra", "object": "model"}, ] } ``` The response contains canonical IDs only. Request aliases such as `opus`, `sonnet`, and `haiku` are resolved server-side but are not separate catalogue entries. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousMessages](/docs/api-reference/messages/)[Next OpenAI Compatibility](/docs/api-reference/openai/) --- ### OpenAI Compatibility Canonical URL: https://llmsrelay.com/docs/api-reference/openai/ # OpenAI Compatibility Use Claude models with the OpenAI SDK — drop-in compatible API. Works with any tool that supports OpenAI format. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Base URL:** https://api.llmsrelay.com/v1 - **Endpoint:** POST /chat/completions - **API key:** sk-cs4-* - **Models:** Use IDs returned by GET /v1/models - **Best for:** Existing OpenAI SDKs and OpenAI-compatible tools ## When to Use This Route - Use the OpenAI-compatible route with a **Basic** key for Claude models, or with a **Codex** key for OpenAI/GPT models. - Use the native `/v1/messages` route if you want the most Claude-native behavior, cleaner troubleshooting, or direct Anthropic-style SDK integration. - For Cursor, Cline, Roo Code, and similar OpenAI-style tools, choose the key group that matches the model family. Kilo Code has a separate OpenAI Responses provider for GPT/Codex models; see the [Kilo Code guide](/docs/guides/vscode-extensions/#kilo-code). ## Endpoint POST https://api.llmsrelay.com/v1/chat/completions ## Authentication Use the standard `Authorization: Bearer` header with your LLMsRelay API key: ``` curl https://api.llmsrelay.com/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "content-type: application/json" \ -d '{ "model": "claude-opus-5-5", "messages": [{"role": "user", "content": "Hello!"}] }' ``` ## Python (OpenAI SDK) OpenAI SDK for Pythonpython ``` from openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.llmsrelay.com/v1" ) response = client.chat.completions.create( model="claude-opus-5-5", messages=[ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is machine learning?"} ] ) print(response.choices[0].message.content) ``` When using the OpenAI SDK, set `base_url` to `https://api.llmsrelay.com/v1` (with `/v1`). ## TypeScript (OpenAI SDK) OpenAI SDK for TypeScripttypescript ``` import OpenAI from "openai"; const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.llmsrelay.com/v1", }); const response = await client.chat.completions.create({ model: "claude-opus-5-5", messages: [{ role: "user", content: "Hello!" }], }); console.log(response.choices[0].message.content); ``` ## Important Compatibility Notes - Use `Authorization: Bearer YOUR_API_KEY` and the `/v1` suffix. This route is not the same as the native Anthropic-format endpoint. - System and developer instructions may be treated as one combined top-level instruction rather than as fully independent turns. - OpenAI-specific fields that are not part of Claude's native behavior may be ignored. Do not assume every OpenAI-only knob changes model behavior. - For tool calling, you should not rely on strict schema guarantees just because your OpenAI client exposes a strict mode toggle. - If a feature feels awkward to express through Chat Completions, switch to the native `/v1/messages` API instead of fighting the compatibility layer. A practical rule: use OpenAI compatibility for migration and tooling convenience, and use the native Messages API when you want the cleanest Claude-specific integration path. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousModels](/docs/api-reference/models/)[Next OpenAI & Codex Endpoints](/docs/api-reference/codex/) --- ### Streaming Canonical URL: https://llmsrelay.com/docs/api-reference/streaming/ # Streaming Claude API streaming — real-time token-by-token responses. SSE format for both Anthropic and OpenAI-compatible endpoints. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Enable:** stream: true - **Format:** Server-Sent Events (text/event-stream) - **Anthropic events:** message_start, content_block_delta, message_stop - **OpenAI events:** choices[].delta.content chunks ending with [DONE] ## Anthropic Streaming Set `"stream": true` in your request body. The response uses Server-Sent Events (SSE): Streaming Requestbash ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: YOUR_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-opus-5-5", "max_tokens": 1024, "stream": true, "messages": [{"role": "user", "content": "Hello!"}] }' ``` ### Event Types SSE Eventstext ``` event: message_start data: {"type":"message_start","message":{"id":"msg_...","type":"message","role":"assistant","content":[],"model":"claude-opus-5-5"}} event: content_block_start data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}} event: content_block_delta data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}} event: content_block_delta data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"! How can I help?"}} event: content_block_stop data: {"type":"content_block_stop","index":0} event: message_stop data: {"type":"message_stop"} ``` ## OpenAI-Compatible Streaming Set `"stream": true` in the Chat Completions request: OpenAI Streaming Requestbash ``` curl https://api.llmsrelay.com/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "content-type: application/json" \ -d '{ "model": "claude-opus-5-5", "stream": true, "messages": [{"role": "user", "content": "Hello!"}] }' ``` ### Event Format SSE Eventstext ``` data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]} data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]} data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":null}]} data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]} data: [DONE] ``` The stream terminates with `data: [DONE]`. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousCursor IDE](/docs/api-reference/cursor-ide/)[Next Errors](/docs/api-reference/errors/) --- ### Errors Canonical URL: https://llmsrelay.com/docs/api-reference/errors/ # Errors Claude API error codes and troubleshooting — 401, 429, 500 status codes. How to handle rate limits and authentication errors. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **401:** Invalid or missing sk-cs4-* key - **402:** Out of credits — top up to continue - **429:** Rate limited — back off using Retry-After - **500/502/504:** Transient — retry with exponential backoff ## Anthropic Error Format Anthropic-style errorjson ``` { "type": "error", "error": { "type": "invalid_request_error", "message": "model: Invalid model ID" } } ``` ## OpenAI Error Format OpenAI-style errorjson ``` { "error": { "message": "Invalid model ID", "type": "invalid_request_error", "code": "invalid_model" } } ``` ## HTTP Status Codes | Status | Meaning | Action | | --- | --- | --- | | 400 | Bad Request — invalid parameters | Check request body and parameters | | 401 | Unauthorized — invalid or missing API key | Verify your API key | | 403 | Forbidden — insufficient permissions | Check key permissions and credit balance | | 429 | Rate Limited — too many requests | Reduce request rate, implement backoff | | 500 | Internal Server Error | Retry with exponential backoff | | 502 | Bad Gateway — upstream error | Retry after a brief delay | | 529 | Overloaded — service at capacity | Retry with exponential backoff | ## Debugging Tips - Always check the `error.type` and `error.message` fields for specific information - For 429 errors, implement exponential backoff starting at 1 second - For 5xx errors, retry up to 3 times with increasing delays - Use `GET /health` to check service status - Verify model IDs match exactly (e.g., `claude-sonnet-4.6`) Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousStreaming](/docs/api-reference/streaming/)[Next Rate Limits](/docs/guides/rate-limits/) --- ### Claude Code Setup with a Custom API Gateway Canonical URL: https://llmsrelay.com/docs/guides/claude-code/ # Claude Code Setup with a Custom API Gateway Install Claude Code, connect it to the LLMsRelay Anthropic-compatible endpoint, select the correct key group, and verify the configuration. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** Basic or Pro - **Base URL:** https://api.llmsrelay.com - **API key:** sk-cs4-* - **Key group:** Basic or Pro - **Billing:** One-time usage packs Use an Anthropic key group Claude Code uses the Anthropic Messages format. Create the key in the **Basic** or **Pro** group. The **Codex** group is reserved for OpenAI and Codex clients. ## Install Claude Code Anthropic recommends the native installer. It installs the CLI without requiring Node.js and keeps Claude Code updated automatically. ### macOS, Linux, or WSL Terminalbash ``` curl -fsSL https://claude.ai/install.sh | bash claude --version ``` ### Windows PowerShell PowerShellpowershell ``` irm https://claude.ai/install.ps1 | iex claude --version ``` Homebrew is also supported: `brew install --cask claude-code`. The old global npm installation is no longer the recommended setup. ## Create the correct API key 1. Open [API Keys](/dashboard/v2/keys/) in the dashboard. 2. Create an `sk-cs4-...` key and select **Basic** or **Pro**. 3. Copy the key when it is shown. It is not displayed again. 4. Add a one-time usage pack from [Plans](/plans/) if the account has no available balance. Do not use an old `api2` endpoint or a Codex-group key. Claude Code must connect to `api.llmsrelay.com` with an Anthropic-compatible key. ## Configure the gateway The shortest setup is to export two variables in the same terminal where you start Claude Code. macOS, Linux, or WSLbash ``` export ANTHROPIC_BASE_URL="https://api.llmsrelay.com" export ANTHROPIC_API_KEY="YOUR_SK_CS4_KEY" export CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1 unset ANTHROPIC_AUTH_TOKEN claude ``` Windows PowerShellpowershell ``` $env:ANTHROPIC_BASE_URL = "https://api.llmsrelay.com" $env:ANTHROPIC_API_KEY = "YOUR_SK_CS4_KEY" $env:CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY = "1" Remove-Item Env:ANTHROPIC_AUTH_TOKEN -ErrorAction SilentlyContinue claude ``` Base URL format Use `https://api.llmsrelay.com` exactly. Do not append `/v1`; Claude Code adds the API path itself. ## Make the configuration persistent Put the gateway variables in your user settings file so the terminal client and the official editor extension use the same configuration: ~/.claude/settings.jsonjson ``` { "env": { "ANTHROPIC_BASE_URL": "https://api.llmsrelay.com", "ANTHROPIC_API_KEY": "YOUR_SK_CS4_KEY", "CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY": "1" }, "model": "claude-sonnet-4.6" } ``` **macOS / Linux / WSL:** `~/.claude/settings.json` **Windows:** `%USERPROFILE%\.claude\settings.json` Keep the key out of Git Store the real key in your user-level file. If a project needs its own settings, use `.claude/settings.local.json` and make sure it is ignored by Git. ## Select a model without conflicts Check the model IDs available to your key, then use one source of truth for model selection. Model discovery adds the gateway catalog to Claude Code's model picker. List modelsbash ``` curl https://api.llmsrelay.com/v1/models \ -H "x-api-key: $ANTHROPIC_API_KEY" ``` - Use the top-level `"model"` value in `~/.claude/settings.json` for a persistent default. - Use `claude --model MODEL_ID` for a one-off override. - Use the Claude Code model picker when you want to switch interactively. Do not configure two different defaults Avoid setting both the top-level `"model"` field and a conflicting `ANTHROPIC_MODEL` environment variable. The environment override can make the UI show one model while the gateway receives another. ## Verify the complete setup Gateway healthbash ``` curl https://api.llmsrelay.com/healthz ``` Claude Code smoke testbash ``` claude -p "Reply with exactly: pong" ``` A `pong` response confirms that Claude Code loaded the gateway URL, authenticated the key, resolved the model, and completed a request. In an interactive session, run `/status` and confirm that the displayed Anthropic base URL is `https://api.llmsrelay.com` and the active credential is `ANTHROPIC_API_KEY`. ## Troubleshooting ### Claude Code opens the Anthropic login flow The variables were not loaded. Start Claude Code from the terminal where you exported them, or add them to `~/.claude/settings.json` and restart the CLI or editor. ### 401 or 403 Confirm the key starts with `sk-cs4-`, belongs to Basic or Pro, and has not been revoked. Remove stale `ANTHROPIC_AUTH_TOKEN` and old `api2` values from the shell and settings files. ### The wrong model is charged Check `~/.claude/settings.json`, project `.claude/settings.json`, local `.claude/settings.local.json`, and `env | grep ANTHROPIC`. Remove conflicting model overrides, then restart Claude Code. ### The editor extension uses different settings Fully restart the editor after changing environment variables. Prefer the user-level settings file when the CLI and extension must share one gateway configuration. ## Related guides - [API key configuration](/docs/guides/api-key-configuration/) - [Cursor setup](/docs/guides/cursor/) - [VS Code extensions](/docs/guides/vscode-extensions/) - [Available models](/docs/getting-started/models/) Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousCursor](/docs/guides/cursor/)[Next Codex CLI & Extension](/docs/guides/codex-cli/) --- ### How to Use Claude API in Cursor IDE — 2-Minute Setup [2026] Canonical URL: https://llmsrelay.com/docs/guides/cursor/ # How to Use Claude API in Cursor IDE — 2-Minute Setup \[2026\] Connect Claude Sonnet 4.6, Opus 4.7, and Haiku 4.5 to Cursor IDE via LLMsRelay. Step-by-step guide, no Anthropic account. Works with Chat, Composer, and Agent. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Provider type:** OpenAI-compatible (custom) - **Base URL:** https://api.llmsrelay.com/v1 - **API key:** sk-cs4-* - **Model id:** claude-sonnet-4.6 (or opus / haiku) - **Setup time:** ~30 seconds ## Quick Answer **To use Claude API in Cursor IDE:** open _Cursor Settings → Models_, enable **OpenAI API Key** and **OpenAI Base URL**, set the base URL to `https://api.llmsrelay.com/v1`, paste your LLMsRelay API key, and add `claude-sonnet-4.6` to the model list. Setup takes under 2 minutes. No Anthropic account, no phone verification — just a Cursor Pro subscription and a LLMsRelay key. ## Why LLMsRelay for Cursor? Cursor's built-in Claude access requires Anthropic's API key, which has account requirements, no Russian/CIS payment options, and limited per-key controls. LLMsRelay solves all three: instant signup, payment in RUB/crypto/cards, and per-key credit limits. - Same Claude models (Sonnet 4.6, Opus 4.7, Haiku 4.5) — same per-token rates as Anthropic direct - OpenAI-compatible endpoint — Cursor sees it as a generic OpenAI proxy - Pay with Mir card, SBP, USDT, BTC, or international Visa/Mastercard - Per-key credit limits — set $20 cap on a key for a junior dev, no surprise bills ## Required Plan Custom API access in Cursor requires a **paid Pro plan or higher**. **Pro trial (7 days)** does **not** unlock custom API endpoints. You must be on a paid Cursor subscription before connecting LLMsRelay. ## Exact Fields in Cursor - **OpenAI API Key**: your API key from the LLMsRelay dashboard (starts with `sk-`) - **OpenAI Base URL**: `https://api.llmsrelay.com/v1` (must include `/v1`) - **Verify**: click "Verify" — Cursor will list all available models ## Recommended Setup 1. Open Cursor settings (`Cmd/Ctrl + ,`) 2. Go to **Models** 3. Turn on **OpenAI API Key** 4. Paste your API key from [llmsrelay.com dashboard](/plans/) 5. Turn on **OpenAI Base URL** 6. Set it to `https://api.llmsrelay.com/v1` 7. Click **Verify** — should show "OpenAI key works" 8. Open Composer (`Cmd/Ctrl + I`), pick a Claude model from the dropdown, send a test prompt Use the OpenAI-compatible path, not Anthropic mode. The base URL for Cursor includes `/v1` at the end. ## Recommended Model IDs - `claude-opus-4.7` — most powerful, best for refactors and architecture - `claude-sonnet-4.6` — best value for daily coding (default) - `claude-haiku-4.5` — fastest, cheapest, ideal for batch edits ## Troubleshooting **"Invalid API Key" error:** Verify the key starts with `sk-` and the base URL ends with `/v1`. Cursor caches keys aggressively — restart Cursor after pasting. **"Model not found":** Make sure the model ID exactly matches (e.g. `claude-sonnet-4.6`, not `claude-sonnet-4.5`). Add the model in Cursor's "Add Model" field. **Slow responses:** Switch to Haiku 4.5 for fast inline edits, or check your network. LLMsRelay adds <100ms overhead vs Anthropic direct. ## Windows-specific notes The 8-step setup above is identical on Windows, macOS, and Linux. A few Windows extras: - **Installer**: download Cursor from [cursor.sh](https://cursor.sh). The MSI puts `cursor` on PATH and registers `%APPDATA%\Cursor` for settings. - **Hotkey**: Ctrl+, for Settings, Ctrl+I for Composer, Ctrl+L for Chat. - **Restart Cursor fully** after pasting the key — close all windows, then end any `Cursor.exe` in Task Manager. Cursor caches OpenAI credentials in memory. - **Corporate proxy**: `setx HTTPS_PROXY "http://proxy.company:8080"` in PowerShell, then reopen Cursor. - **SmartScreen / antivirus** blocking outbound calls — whitelist `api.llmsrelay.com`. - **WSL projects**: Cursor on Windows talks to LLMsRelay directly — env vars inside WSL do not affect Cursor's HTTP calls. Configure the key in Cursor Settings (the GUI), not in `~/.bashrc`. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousIDE Integration](/docs/guides/ide-integration/)[Next Claude Code](/docs/guides/claude-code/) --- ### VS Code Extensions Canonical URL: https://llmsrelay.com/docs/guides/vscode-extensions/ # VS Code Extensions Claude API for VS Code — configure Continue, Cline, and Roo Code extensions with your Claude API key. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. ## Continue Example `config.yaml` model entry (uses the latest flagship Claude Opus 4.7): config.yamlyaml ``` name: Claude API version: 1.0.0 schema: v1 models: - name: Claude Opus 4.7 provider: anthropic model: claude-opus-4.7 apiBase: https://api.llmsrelay.com apiKey: sk-cs4-YOUR_KEY roles: - chat - edit ``` ## Cline Use **OpenAI Compatible** provider: - **Custom BaseURL**: `https://api.llmsrelay.com/v1` - **API Key**: from the dashboard ## Roo Code Use **OpenAI Compatible** provider: - **Custom BaseURL**: `https://api.llmsrelay.com/v1` - **API Key**: from the dashboard ## Kilo Code Kilo Code can use LLMsRelay through two custom providers: Anthropic Messages for Claude models and OpenAI Responses for GPT/Codex models. Keep them as separate providers because the wire protocols are different. Choose the matching key group Create or select a LLMsRelay key in the **Basic** group for Claude/Anthropic models and a key in the **Codex** group for OpenAI/GPT/Codex models. One key has one active group at a time, so use two keys if you need both providers available simultaneously. ### 1\. Add the provider Open Kilo's provider settings or edit the user-level `kilo.jsonc` file. The example below is portable and contains no real credentials: kilo.jsoncjsonc ``` { "$schema": "https://app.kilo.ai/config.json", "provider": { "llmsrelay-anthropic": { "name": "LLMsRelay Anthropic Messages", "npm": "@ai-sdk/anthropic", "options": { "baseURL": "https://api.llmsrelay.com/v1" }, "models": { "claude-opus-5": { "name": "claude-opus-5", "reasoning": true }, "claude-fable-5": { "name": "claude-fable-5", "reasoning": true }, "claude-opus-4.8": { "name": "claude-opus-4.8", "reasoning": true }, "claude-opus-4.7": { "name": "claude-opus-4.7", "reasoning": true }, "claude-opus-4.6": { "name": "claude-opus-4.6", "reasoning": true }, "claude-sonnet-5": { "name": "claude-sonnet-5", "reasoning": true }, "claude-sonnet-4.6": { "name": "claude-sonnet-4.6", "reasoning": true }, "claude-haiku-4.5": { "name": "claude-haiku-4.5", "reasoning": true } } }, "llmsrelay-openai": { "name": "LLMsRelay OpenAI Responses", "npm": "@ai-sdk/openai", "env": ["OPENAI_API_KEY"], "options": { "baseURL": "https://api.llmsrelay.com/v1" }, "models": { "gpt-5.5": { "name": "gpt-5.5", "reasoning": true }, "gpt-5.6-sol": { "name": "gpt-5.6-sol", "reasoning": true }, "gpt-5.6-terra": { "name": "gpt-5.6-terra", "reasoning": true }, "gpt-6-sol": { "name": "gpt-6-sol", "reasoning": true }, "gpt-6-astra": { "name": "gpt-6-astra", "reasoning": true } } } } } ``` ### 2\. Add the API key securely Add a **Basic** key to the Anthropic provider and a separate**Codex** key to the OpenAI provider through Kilo's credential dialog or the `/connect` flow. Use separate placeholders such as `sk-cs4-YOUR_BASIC_KEY` and `sk-cs4-YOUR_CODEX_KEY`. Do not put a real key in `kilo.jsonc`, workspace files, screenshots, or a public repository. ### 3\. Select a model Select the exact ID returned by `GET /v1/models`. Claude IDs belong to the Anthropic provider; GPT/Codex IDs belong to the OpenAI Responses provider. For example, use `gpt-5.5`, not a shortened alias. OpenAI models are sent to `/v1/responses`; Claude models are sent to `/v1/messages`. After changing the provider or model list, reload the VS Code window and start a new Kilo session so the extension drops any cached model selection. ## Windows setup (all extensions) On Windows, Continue, Cline, Roo Code, and Kilo Code use the same gateway, but their provider settings differ. Follow the shared environment-variable steps below, then use the extension-specific notes. ### 1\. Prerequisites - VS Code 1.80+ (latest recommended) - Node.js LTS — required only if you use Continue's CLI/MCP features. Verify with `node -v` in PowerShell. - Your LLMsRelay API keys: `sk-cs4-YOUR_BASIC_KEY` for Claude and `sk-cs4-YOUR_CODEX_KEY` for OpenAI. ### 2\. Set environment variables (PowerShell) Open **PowerShell** (not cmd.exe) and run. `setx` writes to the user environment permanently — values become available in every _new_ terminal and process. PowerShellpowershell ``` setx ANTHROPIC_BASE_URL "https://api.llmsrelay.com" setx ANTHROPIC_API_KEY "sk-cs4-YOUR_BASIC_KEY" setx OPENAI_BASE_URL "https://api.llmsrelay.com/v1" setx OPENAI_API_KEY "sk-cs4-YOUR_CODEX_KEY" ``` GUI alternative: press Win+R → type `sysdm.cpl` → **Advanced** → **Environment Variables** → add the same five variables under "User variables". ### 3\. Fully restart VS Code VS Code caches environment variables at launch. After `setx`, close every VS Code window, then open Task Manager (Ctrl+Shift+Esc) and end any remaining `Code.exe` processes before reopening. Otherwise the extensions will still see old values. ### 4\. Per-extension Windows notes - **Continue** — config lives at `%USERPROFILE%\.continue\config.yaml`. Open it with `notepad %USERPROFILE%\.continue\config.yaml` in PowerShell. - **Cline / Roo Code** — open the extension's settings panel, choose **OpenAI Compatible**, paste Base URL `https://api.llmsrelay.com/v1` and your API key. The env vars from step 2 are used as defaults for new tasks but the UI values always win. - **Kilo Code** — add the two custom providers from the [Kilo Code section above](#kilo-code). Store the key for each provider through Kilo credentials or `/connect`, and select an exact model ID from `GET /v1/models`. ### 5\. Using WSL / Remote-WSL? Windows env vars do **not** propagate into the WSL distro. Add the same variables to `~/.bashrc` (or `~/.zshrc`) inside WSL: ~/.bashrc (inside WSL)bash ``` export ANTHROPIC_BASE_URL="https://api.llmsrelay.com" export ANTHROPIC_API_KEY="sk-cs4-YOUR_BASIC_KEY" export OPENAI_BASE_URL="https://api.llmsrelay.com/v1" export OPENAI_API_KEY="sk-cs4-YOUR_CODEX_KEY" ``` ### 6\. Troubleshooting - `setx` never affects the current terminal — open a new one. - In PowerShell always wrap values in double quotes, especially URLs with `/`. - Behind a corporate proxy: also set `setx HTTPS_PROXY "http://proxy.company:8080"`. - Antivirus / firewall blocking requests — whitelist the host `api.llmsrelay.com`. - Wrong model is being called — see [Common Claude Code configuration mistakes](/docs/guides/claude-code/#common-mistakes). ## Wrong model is being used? Check ANTHROPIC\_MODEL conflicts If the extension keeps calling a different model than the one you selected in the UI, see [Common Claude Code configuration mistakes](/docs/guides/claude-code/#common-mistakes) — the most frequent cause is a conflict between the UI-selected `model` and an `ANTHROPIC_MODEL` environment override. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousOpenAI-Compatible IDEs](/docs/guides/openai-ides/)[Next OpenCode](/docs/guides/opencode/) --- ### LLMsRelay API Key Configuration Canonical URL: https://llmsrelay.com/docs/guides/api-key-configuration/ # LLMsRelay API Key Configuration Create a LLMsRelay API key, set usage limits, rotate it, and store it securely. Works with tools that support Anthropic-compatible or OpenAI-compatible endpoints. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. ## What Is a LLMsRelay API Key? A LLMsRelay API key is a secret token that authenticates requests to the LLMsRelay developer gateway. It works with supported Anthropic-compatible and OpenAI-compatible request formats. Use `GET /v1/models` to check which model IDs the key can use. Use the key with tools that support a custom base URL, including [Cursor IDE](/docs/guides/cursor/), [Claude Code](/docs/guides/claude-code/), [VS Code extensions](/docs/guides/vscode-extensions/), and custom applications. ## Creating Your First API Key Creating a LLMsRelay API key takes a few steps: 1. Log in to your [LLMsRelay dashboard](https://llmsrelay.com) 2. Navigate to **API Keys** 3. Click **Create Key** 4. Set an optional name and per-key credit limit 5. Copy and securely store your key — it starts with `sk-cs4-` API keys are only shown once at creation. Store them securely — they cannot be retrieved later. ## Per-Key Controls & Usage Tracking LLMsRelay gives you fine-grained control over every API key: - **Credit limit** — Cap the maximum spend per key to prevent runaway costs - **Usage tracking** — Monitor token consumption, request counts, and cost per model in real time - **Instant revocation** — Disable a compromised or unused key without affecting others - **Multiple keys** — Create separate keys for development, staging, and production environments ## Storing API Keys Securely Follow these best practices to keep your API credentials safe: - Store keys in **environment variables** (`.env` files), never in source code - Add `.env` to your `.gitignore` to prevent accidental commits - Use a secrets manager (Vault, AWS Secrets Manager, Doppler) in production - Set per-key credit limits to cap exposure if a key leaks - Rotate keys periodically — create a new key, update configs, then revoke the old one - Never expose API keys in client-side JavaScript or public repositories ## Using Your Key with IDEs and Tools Configure your client to use the compatible base URL and your `sk-cs4-` key. The exact option names depend on the tool: - [Cursor](/docs/guides/cursor/) — Set the API key in Cursor Settings → Models → Anthropic - [Claude Code](/docs/guides/claude-code/) — Use a `sk-cs4-*` key, set `ANTHROPIC_BASE_URL=https://api.llmsrelay.com` and `ANTHROPIC_API_KEY` (not `ANTHROPIC_AUTH_TOKEN`) - [VS Code](/docs/guides/vscode-extensions/) — Configure in extension settings with custom endpoint - [OpenAI SDK](/docs/api-reference/openai-compatible/) — Use the `/v1/chat/completions` endpoint for tools expecting OpenAI format Ready to get started? [Buy platform usage](/plans/) and create your first key. See the [Quickstart guide](/docs/getting-started/quickstart/) for a complete walkthrough. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousOpenClaw](/docs/guides/openclaw/)[Next GLM 5.3 Uncensored](/docs/guides/glm-uncensored/) --- ### Rate Limits Canonical URL: https://llmsrelay.com/docs/guides/rate-limits/ # Rate Limits Claude API rate limits — 429 handling, Retry-After, concurrency protection, and throughput guidance. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Status code:** 429 Too Many Requests - **Header:** Retry-After (seconds) - **Scope:** Per-key + global protection - **Strategy:** Exponential backoff with jitter ## Rate Limiting The LLMsRelay gateway applies protective limits to preserve service stability under mixed traffic patterns. Enforcement is not a fixed public tier table: it depends on request shape, concurrency, and current gateway load. When you exceed a limit, the API returns a `429` status code along with a `Retry-After` header. ## What is limited | Scope | What it means | | --- | --- | | Per API key | A key can be throttled independently of other keys. | | Per-key concurrency | Too many simultaneous long-running requests can trigger 429. | | Global gateway protection | The gateway can shed load to protect service stability. | | Attachment-heavy traffic | Large multimodal turns may be treated more strictly than plain text traffic. | We do not publish a stable numeric RPM/TPM contract for every request shape. If you need sustained higher throughput, contact support with your expected traffic pattern. ## Response headers - `retry-after` — seconds to wait before retrying after a 429 Do not rely on undocumented rate-limit headers as a stable public contract. ## Handling 429 Responses - Implement exponential backoff (start at 1s, double each retry) - Respect `Retry-After` headers when present - Queue requests in your application layer and reduce parallelism when needed - Be especially conservative with long-running streams and multimodal turns - Use caching to reduce avoidable repeat traffic If you consistently hit rate limits, contact support to discuss higher sustained throughput for your use case. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousErrors](/docs/api-reference/errors/)[Next IDE Integration](/docs/guides/ide-integration/) --- ### Codex CLI & Extension Setup (OpenAI API) Canonical URL: https://llmsrelay.com/docs/guides/codex-cli/ # Codex CLI & Extension Setup (OpenAI API) Use the OpenAI API in OpenAI Codex CLI and the Codex VS Code extension through LLMsRelay. Configure the Responses API endpoint in under a minute. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Base URL:** https://api.llmsrelay.com/v1 - **Endpoint:** POST /v1/responses - **API key:** sk-cs4-* (Codex group) - **Models:** gpt-5.5, gpt-5.6-sol, gpt-5.6-terra, gpt-6-sol, gpt-6-astra - **Setup time:** ~1 minute ## Overview OpenAI's **Codex CLI** and the **Codex VS Code extension** talk to a model provider over the **Responses API** (`/v1/responses`). LLMsRelay serves that endpoint, so you can drive the **OpenAI API** through your existing LLMsRelay credits and a Codex-group key. Prefer an OpenAI-compatible chat client (Cline, Roo Code, Continue) that uses `/v1/chat/completions` instead? See the [OpenAI-Compatible IDEs](/docs/guides/openai-ides) guide. ## Step 1 — Create a Codex key In the [dashboard → API keys](/dashboard/keys), create an`sk-cs4-…` key and select the**Codex** group. It looks like`sk-cs4-…`. Copy it once — it's shown only at creation. ## Step 2 — Switch the tier to Codex Confirm that the key group is **Codex**. This routes the key to the OpenAI API. Create a separate Basic or Pro key for Claude if both providers must stay active; the account balance remains shared. Codex group serves the OpenAI API only With a Codex-group key, OpenAI models (e.g. `gpt-5.5`) use the OpenAI endpoints. Claude model ids and `/v1/messages` are not used on this tier. ## Step 3 — Configure Codex CLI Codex CLI reads `~/.codex/config.toml`. Add LLMsRelay Codex as a model provider that uses the Responses API: ~/.codex/config.tomltoml ``` model = "gpt-5.5" model_provider = "llmsrelay" [model_providers.llmsrelay] name = "LLMsRelay Codex" base_url = "https://api.llmsrelay.com/v1" env_key = "LLMSRELAY_API_KEY" wire_api = "responses" ``` Export your Codex-group key as the provider env var, then launch Codex: shellbash ``` export LLMSRELAY_API_KEY="sk-cs4-your-key" codex ``` `wire_api = "responses"` is what makes Codex CLI call `/v1/responses`. Streaming works out of the box. ## Step 4 — Codex VS Code extension The Codex VS Code extension shares the same `~/.codex/config.toml`. Once the provider above is set and `LLMSRELAY_API_KEY` is exported in your shell environment, the extension picks up LLMsRelay Codex automatically. Select the `gpt-5.5` model from the extension's model picker. ## Step 5 — Test You can verify the endpoint directly without the CLI: Responses API smoke testbash ``` curl https://api.llmsrelay.com/v1/responses \ -H "Authorization: Bearer sk-cs4-your-key" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-5.5", "input": "Say hello in 3 words." }' ``` A successful response confirms the key, tier, and endpoint are wired correctly. If you get a model error, run `GET /v1/models` to confirm the model id. ## Troubleshooting - **Unknown model** — the key is not in the Codex group, or the model id is wrong. Check the group and `GET /v1/models`. - **401 / unauthorized** — the env var isn't exported in the shell that launched Codex. - **Endpoint not found** — make sure the base URL ends in `/v1` and the provider uses `wire_api = "responses"`. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousClaude Code](/docs/guides/claude-code/)[Next OpenAI-Compatible IDEs](/docs/guides/openai-ides/) --- ### OpenAI-Compatible IDEs (Cline, Roo, Continue) Canonical URL: https://llmsrelay.com/docs/guides/openai-ides/ # OpenAI-Compatible IDEs (Cline, Roo, Continue) Configure Cline, Roo Code, Continue, and any OpenAI-compatible IDE to use the OpenAI API through LLMsRelay's /v1/chat/completions endpoint. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Base URL:** https://api.llmsrelay.com/v1 - **Endpoint:** POST /v1/chat/completions - **API key:** sk-cs4-* (Codex group) - **Models:** latest OpenAI models (e.g. gpt-5.5) ## When to use this guide Most IDE assistants speak the OpenAI `/v1/chat/completions` format and let you set a custom base URL. That's the path for **Cline**, **Roo Code**, **Continue**, and similar tools. Using OpenAI's own **Codex CLI** or the **Codex extension**? Those use the Responses API — see the [Codex CLI & Extension](/docs/guides/codex-cli) guide instead. ## Prerequisites - An `sk-cs4-…` key from the **Codex** group in the [dashboard](/dashboard/keys). - For Claude/Anthropic models, create or select a separate key in **Basic** or **Pro**. ## Cline 1. Open Cline settings → API Provider. 2. Choose **OpenAI Compatible**. 3. Base URL: `https://api.llmsrelay.com/v1` 4. API Key: your `sk-cs4-…` key. 5. Model ID: `gpt-5.5` ## Roo Code 1. Settings → Providers → **OpenAI Compatible**. 2. Base URL: `https://api.llmsrelay.com/v1` 3. API Key: your `sk-cs4-…` key. 4. Model: `gpt-5.5` ## Continue Edit `~/.continue/config.json`: ~/.continue/config.jsonjson ``` { "models": [ { "title": "OpenAI (LLMsRelay)", "provider": "openai", "model": "gpt-5.5", "apiKey": "sk-cs4-your-key", "apiBase": "https://api.llmsrelay.com/v1" } ] } ``` ## Direct SDK usage ### Python (openai SDK) pythonpython ``` from openai import OpenAI client = OpenAI( api_key="sk-cs4-your-key", base_url="https://api.llmsrelay.com/v1", ) resp = client.chat.completions.create( model="gpt-5.5", messages=[{"role": "user", "content": "Write a Python quicksort."}], ) print(resp.choices[0].message.content) ``` ### cURL bashbash ``` curl https://api.llmsrelay.com/v1/chat/completions \ -H "Authorization: Bearer sk-cs4-your-key" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-5.5", "messages": [{"role": "user", "content": "Hello in 3 words"}], "stream": true }' ``` ## Notes - Use `GET /v1/models` with your key to enumerate available OpenAI model ids. - `/v1/messages` (Anthropic format) is not used with Codex-group keys — stick to the OpenAI endpoints. - OpenAI usage is billed against your LLMsRelay credits at the published token rates. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousCodex CLI & Extension](/docs/guides/codex-cli/)[Next VS Code Extensions](/docs/guides/vscode-extensions/) --- ### Pricing Canonical URL: https://llmsrelay.com/docs/billing/pricing/ # Pricing LLMsRelay platform pricing with four one-time usage packs. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Pricing model:** One-time usage packs - **$500 usage:** $45 - **$1,000 usage:** $90 - **$6,000 usage:** $450, includes +20% - **$13,000 usage:** $900, includes +30% llmsrelay --pricingup to −91% $ Pay once. Get ~11× the Anthropic balance. Same Anthropic per-token rates. Massive discount on the top-up. | you pay | Anthropic-equivalent balance | discount | | --- | --- | --- | | $45 | $500balance | −91% | | $90popular | $1,000balance | −91% | Tokens are billed at Anthropic's exact per-token rates · Balance never expires · No subscription ## Billing Model LLMsRelay uses platform usage billing. Requests deduct usage from your balance based on actual token consumption and the rate card shown below. ## Cost Formula Request cost calculationtext ``` cost = (input_tokens × input_rate) + (output_tokens × output_rate) + (cache_write_tokens × cache_write_rate) + (cache_read_tokens × cache_read_rate) ``` ## Token Categories | Category | Description | | --- | --- | | Input (uncached) | Standard input tokens not served from cache | | Output | Generated response tokens | | Cache Write | Tokens written to the prompt cache on first request | | Cache Read | Tokens served from prompt cache on subsequent requests | ## Per-Model Rates Rates are per million tokens (MTok) and match the model pricing used for request deductions. Use `GET /v1/models` to discover the models available to your key; there is no separate pricing endpoint. | Model | Input | Output | Cache Write | Cache Read | | --- | --- | --- | --- | --- | | claude-opus-5 | $5.00 | $25.00 | $6.25 | $0.50 | | claude-fable-5 | $10.00 | $50.00 | $12.50 | $1.00 | | claude-opus-4.8 | $5.00 | $25.00 | $6.25 | $0.50 | | claude-opus-4.7 | $5.00 | $25.00 | $6.25 | $0.50 | | claude-opus-4.6 | $5.00 | $25.00 | $6.25 | $0.50 | | claude-sonnet-5 | $2.00 | $10.00 | $2.50 | $0.20 | | claude-sonnet-4.6 | $3.00 | $15.00 | $3.75 | $0.30 | | claude-haiku-4.5 | $1.00 | $5.00 | $1.25 | $0.10 | ## Important Notes - Streaming requests are billed the same as non-streaming requests - Cache read tokens are significantly cheaper than uncached input tokens - Each model has its own rate table, shown above Use `GET /v1/models` to retrieve the current model catalogue programmatically. Pricing rates are documented on this page. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousClaude X20 Max / Max 20x](/docs/learn/claude-x20-max-subscription/)[Next Credits](/docs/billing/credits/) --- ### Platform Usage Balance Canonical URL: https://llmsrelay.com/docs/billing/credits/ # Platform Usage Balance How LLMsRelay platform usage works, how to add balance, and how to track usage. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Balance:** Prepaid platform usage - **Expiration:** Never - **Empty balance error:** HTTP 402 - **Payment methods:** Cards, crypto, Tribute (RU) ## What Is Platform Usage? Platform usage is the prepaid balance for LLMsRelay API requests. It is deducted from your balance based on token consumption and the applicable model rate card. ## Usage Dashboard Your usage is visible in the LLMsRelay dashboard, where you can see: - Total credits consumed - Per-key usage breakdown - Request history and token counts - Cost per request ## Per-Key Limits You can set credit limits on individual API keys to control spending: - Set maximum credit consumption per key - Monitor approaching limits in the dashboard - Keys automatically stop working when their limit is reached Set per-key limits to prevent unexpected charges, especially for development or testing keys. ## Adding Credits Credits can be purchased through the LLMsRelay dashboard. Visit [llmsrelay.com](https://llmsrelay.com) to manage your balance. Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousPricing](/docs/billing/pricing/)[Next Changelog](/docs/changelog/) --- ### How to Buy LLMsRelay API Usage and Create a Key Canonical URL: https://llmsrelay.com/docs/learn/how-to-buy-claude-api-key/ # How to Buy LLMsRelay API Usage and Create a Key Create a LLMsRelay account, add platform usage, generate a key, and configure a compatible API route. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. ## What You Create A LLMsRelay API key authenticates your requests to the LLMsRelay developer gateway. It can be used with supported Anthropic-compatible and OpenAI-compatible request formats in custom applications, IDEs, and automation tools. The exact model IDs available to a key depend on its group and current gateway availability. Confirm them with `GET /v1/models` before configuring a production client. ## Step-by-Step: Getting Your API Key ### 1\. Create a LLMsRelay Account Visit [llmsrelay.com](https://llmsrelay.com) and sign up. Registration is quick — you only need an email address. ### 2\. Choose a Plan Select a one-time platform usage pack on the [Plans page](/plans/). The current price, usage amount, payment methods, and any fixed bonus are displayed before checkout. ### 3\. Add Credits Purchase platform usage through the dashboard. Available card and cryptocurrency options are shown at checkout. ### 4\. Generate Your API Key Navigate to **API Keys** in your dashboard and click **Create Key**. You can optionally set a name and credit limit per key. Your API key is shown only once. Copy and store it securely — it cannot be retrieved later. ### 5\. Start Making Requests Use your key with the Anthropic-compatible Messages API endpoint: Test your keybash ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: YOUR_API_KEY" \ -H "content-type: application/json" \ -H "anthropic-version: 2023-06-01" \ -d '{ "model": "claude-opus-5-5", "max_tokens": 256, "messages": [{"role": "user", "content": "Hello"}] }' ``` ## Frequently Asked Questions ### Can I use this key with Cursor, VS Code, or Claude Code? Yes. LLMsRelay API keys are compatible with all major IDE integrations. See our [IDE Integration guide](/docs/guides/ide-integration/) for setup instructions. ### Is there a pay-as-you-go starter ($5)? LLMsRelay uses pay-as-you-go pricing. You only pay for what you use — start with the smallest credit package to test the service. ### Which model IDs should I configure? Use `GET /v1/models` after creating your key. It returns the IDs currently available to that key and is the source of truth for configuration. ## Related Articles [Claude API Key Setup in 2 Minutes](/docs/learn/claude-api-quick-setup/)[Cheapest Claude API Access](/docs/learn/cheapest-claude-api/)[Claude API vs Direct](/docs/learn/claude-api-vs-direct/) Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousHow to Make Money with AI APIs](/docs/learn/ai-api-reseller/)[Next Claude API Key Setup in 2 Minutes](/docs/learn/claude-api-quick-setup/) --- ### How Billing Works Canonical URL: https://llmsrelay.com/docs/learn/how-billing-works/ # How Billing Works LLMsRelay uses a shared prepaid platform usage balance funded by one-time packs. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. LLMsRelay uses a shared credit balance funded by one-time usage packs. Every API request deducts from that balance based on actual token consumption. ## Billing Model - **One-time packs** — add usage without recurring billing - **Per-request billing** — each API call deducts based on tokens used - **Persistent balance** — purchased packs, referral rewards, and admin grants do not expire - **Real-time balance** — check your remaining credits in the dashboard ## Cost Calculation Each request is billed based on five token categories: | Token Type | Description | | --- | --- | | Input tokens | Standard uncached input tokens sent to the model | | Output tokens | Generated response tokens returned by the model | | Cache write tokens | Tokens written to the prompt cache | | Cache read tokens | Tokens read from the prompt cache (discounted) | | Thinking tokens | Extended thinking tokens (Opus and Sonnet models) | Streaming is billed the same as non-streaming requests. There is no additional cost for using streaming mode. ## Usage Packs | Option | Price | Credits | | --- | --- | --- | | $500 usage pack | $45 | $500 platform usage | | $1,000 usage pack | $90 | $1,000 platform usage | | $5,000 + 20% pack | $450 | $6,000 total usage | | $10,000 + 30% pack | $900 | $13,000 total usage | ## Payment Methods - **Credit / Debit Card** — Visa, Mastercard, and other major cards via secure checkout - **Cryptocurrency** — Bitcoin (BTC), Ethereum (ETH), USDT (TRC-20, ERC-20), USDC - **Telegram** — direct purchase through support chat All payment methods are processed quickly and credits appear in your account within minutes. ## FAQ #### Do credits expire? Purchased usage packs, referral rewards, and administrator grants remain available until used. #### Can I get an invoice? Yes. Contact support via Telegram to request an invoice for your purchase. #### What happens when credits reach zero? API requests return a 402 error. Top up credits at any time — there is no waiting period or reactivation fee. ## Related Pages - [Detailed Pricing](/docs/billing/pricing/) - [Credits System](/docs/billing/credits/) - [Pricing Explained](/docs/learn/pricing-explained/) - [Support & Assistance](/docs/learn/refund-policy/) - [View Plans](/plans/) Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousSupported Countries](/docs/learn/supported-countries/)[Next Activation Time](/docs/learn/activation-time/) --- ### How to Make Money with AI APIs: Reseller Guide Canonical URL: https://llmsrelay.com/docs/learn/ai-api-reseller/ # How to Make Money with AI APIs: Reseller Guide A practical guide to selling Claude and OpenAI-compatible API access under your own product, including reseller approval, payments, client keys, usage tracking, security, and implementation services. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Public entry:** Application plus the existing $90 pack - **What the pack adds:** $1,000 API usage to your personal balance - **Approval:** Manual review - **Client key groups:** Basic, Pro, Codex - **Client key limit:** Up to 500 active keys - **Management key:** csr_ key for backend use only - **Payments:** Card in USD, EUR, RUB and crypto checkout - **Optional build services:** $1,000 white-label site or $3,000 custom platform Read the commercial and technical terms before applying or connecting your production backend. [Open reseller dashboard](/dashboard/reseller/)[Review API pricing](/plans/) ## What an AI API reseller actually sells An AI API reseller packages model access into a product for a specific audience. You might serve developers in one country, agencies that need separate client budgets, a coding tool, an automation platform, or a community that wants one familiar checkout and support channel. LLMsRelay supplies the gateway, reseller wallet, management API, client API keys, and usage records. You decide how your own product looks, what you charge, how customers register, and which services you include. Your gross margin is customer revenue minus API usage, payment fees, support, refunds, taxes, and operating costs. This is a business infrastructure program, not a promise of passive income. It works best when you already understand your audience and can explain why customers should buy through you. ## Common reseller business models ### API storefront Sell prepaid API access from your own website. Your backend creates a client key after payment and assigns a spending limit. ### Managed client accounts Give agencies or teams separate keys, budgets, model groups, and usage pages while you manage the main reseller wallet. ### AI product with bundled usage Include model usage inside a SaaS product, bot, agent, IDE tool, or automation service instead of exposing raw API access as the main product. ### Local sales and support Offer local language onboarding, regional payment methods, setup help, and first-line support for customers who want a simpler buying process. ## How reseller access works 1. ### Submit the application Describe the project, website, sales channels, contact details, and expected monthly API volume. Use the Reseller section in your dashboard. A real project description helps the manual review. 2. ### Buy the existing $90 pack Use the regular usage\_1000 product. It has no reseller discount and adds $1,000 of API usage to your personal balance. Card checkout supports USD, EUR, and RUB. Heleket and Cryptomus provide cryptocurrency and network selection. 3. ### Wait for manual approval The application changes to Paid, awaiting approval. Payment does not guarantee reseller approval. The $1,000 usage remains on your personal API balance while the application is reviewed. 4. ### Create a management key After approval, create a csr\_ key and store it only on your backend. The secret is shown once. LLMsRelay stores a hash rather than the raw key. 5. ### Fund the reseller wallet Top up the separate reseller wallet with available usage packs. Client keys spend from this wallet. Personal usage and reseller usage are accounted for separately. 6. ### Issue client keys Create cs4\_ keys, select Basic, Pro, or Codex, and set a customer spending limit. Your backend can fund, block, unblock, update, or permanently delete each client key. 7. ### Show usage in your product Use the balance and usage endpoints to display totals, remaining balance, requests, tokens, errors, models, and key-level activity. Client-facing endpoints let each customer see only the statistics for their own key. ## Current entry and payment terms | Term | Value | Current terms | | --- | --- | --- | | Entry purchase | $90 | The existing $1,000 usage package, without a special reseller discount. | | Card checkout | USD, EUR, RUB | The final provider amount is shown before payment and may include card processing costs. | | Crypto checkout | Heleket or Cryptomus | The customer selects the available cryptocurrency and network on the provider checkout. | | Review | Manual | LLMsRelay may approve, reject, suspend, or remove reseller access. | | Reseller wallet | Separate balance | Client keys do not spend from the reseller's personal API balance. | | Public discounts | Not automatic | The current public entry and top-up catalog uses the listed package terms. | ## What the reseller platform provides The reseller API is designed for server-to-server integration. Do not place the management key in browser code, mobile apps, public repositories, or customer devices. ### Platform capabilities - A separate tenant\_id and reseller wallet. - A csr\_ management key with rotation and a default management rate limit. - Up to 500 active cs4\_ client keys. - Basic and Pro groups for Claude models, plus Codex for OpenAI models. - Per-key spending limits, funding, blocking, unblocking, and permanent deletion. - Canonical request fields: POST /reseller/v1/keys uses lifetime\_credit\_cap in USD, PATCH uses the same field, and POST /reseller/v1/keys/:id/fund uses amount in USD. - The legacy aliases balance\_limit\_usd and add\_usd remain accepted for compatibility. A key limit is a ceiling against the shared reseller wallet, not a second wallet. - Reseller-wide and per-key usage grouped by model, key group, and client key. - Client endpoints GET /v1/balance and GET /v1/usage. - Management endpoints under /reseller/v1. - Management-key rotation and per-key rate limits. - Balance and usage endpoints for reseller-side monitoring. ## What you must operate yourself ### Your responsibilities - Your website, domain, brand, customer registration, and account recovery. - Retail pricing, invoices, taxes, local compliance, refund terms, and payment disputes. - Customer support, onboarding, documentation, and abuse handling. - Secure storage of the csr\_ management key and any customer secrets. - Monitoring wallet balance and preventing sales that exceed available reseller funds. - Clear communication that your product is independently operated and is not an official Anthropic or OpenAI service. ## Optional implementation services $1,000 ### White-label reseller website A production storefront and customer dashboard connected to the reseller API, including brand setup, domain connection, checkout integration, and deployment. $3,000 ### Custom reseller platform A custom product with architecture, UI, workflows, integrations, testing, and production launch based on an agreed scope. ## A sensible launch checklist 1. Choose one customer segment and write down the problem you solve for it. 2. Decide whether you sell raw API keys, managed accounts, or usage bundled into another product. 3. Model your margin using realistic usage, payment fees, support time, refunds, and taxes. 4. Keep the csr\_ key on the backend and rotate it when a team member or integration changes. 5. Set conservative client limits and test block, unblock, balance, and usage flows. 6. Monitor the reseller balance and usage endpoints before accepting customer payments. 7. Publish clear pricing, support, privacy, acceptable-use, and refund terms. 8. Start with a small group of customers and review usage before scaling. Important **Income is not guaranteed.** LLMsRelay provides API infrastructure and account controls. It does not provide customers, traffic, guaranteed margins, or legal and tax advice. Validate demand before paying for custom development or committing to customer contracts. ## Frequently asked questions ### Can I set my own price for customers? Yes. Your product controls retail pricing. You should include API usage, payment fees, support, refunds, taxes, and operating costs when calculating your margin. ### Does the $90 entry payment go into the reseller wallet? No. The existing $90 package adds $1,000 of usage to your personal API balance. After approval, the reseller wallet is separate and must be funded through reseller top-ups. ### Can I resell both Claude and OpenAI-compatible models? Yes. Create Basic or Pro client keys for Claude models and Codex client keys for OpenAI models. Usage can be reported by model, group, and client key. ### Where do my customers see their balance? Your website or application should call GET /v1/balance and GET /v1/usage with the customer's client key, then display the returned data in your own UI. ### Does LLMsRelay collect payments from my customers? Not in the standard reseller API. You operate customer checkout and registration. A white-label or custom implementation can connect those flows to the reseller API. ### Is reseller approval automatic after payment? No. Approval is manual. LLMsRelay can reject, suspend, or remove access when a project creates security, abuse, payment, or compliance risk. ### Can I expose the csr\_ key in frontend code? No. The csr\_ key can create and manage customer keys, so it belongs only on a protected backend. Use client keys for customer-facing API access. ### Do you guarantee profit? No. Results depend on demand, pricing, usage patterns, payment costs, support load, refunds, taxes, and how well you operate the product. ## Related pages - [Open the Reseller dashboard](/dashboard/reseller/) - [How usage billing works](/docs/learn/how-billing-works/) - [Check available models](/docs/api-reference/models/) - [Review usage packs](/plans/) Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviouscURL](/docs/sdks/curl/)[Next How to Buy a Claude API Key](/docs/learn/how-to-buy-claude-api-key/) --- ### Anthropic-Compatible API Gateway Canonical URL: https://llmsrelay.com/docs/learn/claude-api-gateway/ # Anthropic-Compatible API Gateway Use the LLMsRelay developer gateway with Anthropic-compatible Messages API and OpenAI-compatible API formats. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. ## TL;DR LLMsRelay provides Anthropic-compatible Messages API and OpenAI-compatible endpoints. Use the key group required by the model IDs returned through `GET /v1/models`. ## What Is an Anthropic-Compatible Gateway? An Anthropic-compatible gateway accepts Messages API requests at a custom base URL. LLMsRelay also provides an OpenAI-compatible endpoint for clients that use Chat Completions. The platform account has one usage balance and can issue separate API keys with limits and model controls. The key group determines which model families are eligible. Confirm the available IDs and capabilities with `GET /v1/models`. ## Two Endpoints, Matching Key Groups | Endpoint | Base URL | Format | Best For | | --- | --- | --- | --- | | Anthropic Messages | https://api.llmsrelay.com/v1/messages | Anthropic SDK | Python/TS Anthropic SDK, Claude Code | | OpenAI Compatible | https://api.llmsrelay.com/v1 | OpenAI SDK | Cursor, VS Code, any OpenAI-compatible tool | ## Quick Examples ### Anthropic Format Anthropic-compatible requestbash ``` curl https://api.llmsrelay.com/v1/messages \ -H "x-api-key: YOUR_API_KEY" \ -H "content-type: application/json" \ -H "anthropic-version: 2023-06-01" \ -d '{ "model": "claude-opus-5-5", "max_tokens": 256, "messages": [{"role": "user", "content": "Hello!"}] }' ``` ### OpenAI Format OpenAI-compatible requestbash ``` curl https://api.llmsrelay.com/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-opus-5-5", "max_tokens": 256, "messages": [{"role": "user", "content": "Hello!"}] }' ``` ## Why Use a Gateway? - **Two request formats**: Anthropic-compatible Messages and OpenAI-compatible Chat Completions - **Usage controls**: per-key limits, request history, and token-level reporting - **One-time usage packs**: add balance without recurring billing - **Checkout options**: available card and cryptocurrency methods are shown before payment - **Configurable tools**: connect clients that allow a custom base URL ## LLMsRelay API Formats | Area | LLMsRelay behavior | | --- | --- | | Messages API | Use POST /v1/messages with x-api-key authentication | | Chat Completions | Use POST /v1/chat/completions with Bearer authentication | | Model discovery | Use GET /v1/models for the authenticated key | | Billing | Platform usage is deducted from the account balance | | Tool support | Works with clients that allow a compatible custom base URL | ## Supported Tools & Integrations - [Cursor IDE](/docs/learn/claude-api-key-for-cursor/) — AI-powered code editor - [VS Code Extensions](/docs/learn/claude-api-for-vscode/) — Continue, Cline, Roo Code - [Claude Code CLI](/docs/guides/claude-code/) — terminal AI assistant - [Python SDK](/docs/sdks/python/) — Anthropic & OpenAI libraries - [TypeScript SDK](/docs/sdks/typescript/) — Node.js integration - Any OpenAI-compatible tool — LangChain, LlamaIndex, etc. ## Related Articles [Claude API Instant Access](/docs/learn/claude-api-without-waitlist/)[Claude API Crypto Payment](/docs/learn/claude-api-crypto-payment/)[Claude API vs Direct](/docs/learn/claude-api-vs-direct/) Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousClaude API Crypto Payment](/docs/learn/claude-api-crypto-payment/)[Next Cheapest Claude API Access](/docs/learn/cheapest-claude-api/) --- ### OpenAI-Compatible API — Drop-in /v1/chat/completions Canonical URL: https://llmsrelay.com/docs/learn/openai-compatible-api/ # OpenAI-Compatible API — Drop-in /v1/chat/completions Use an OpenAI-compatible API through LLMsRelay. Point the OpenAI SDK, Cline, Roo Code, Continue, or any custom base URL client at /v1/chat/completions. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Endpoint:** POST /v1/chat/completions - **Base URL:** https://api.llmsrelay.com/v1 - **Models:** latest OpenAI models (e.g. gpt-5.5) - **Works with:** OpenAI SDK, Cline, Roo, Continue ## TL;DR LLMsRelay exposes a standard `/v1/chat/completions` endpoint for the OpenAI API. Change the base URL on any OpenAI client and set the model to `gpt-5.5` — no code rewrite. ## Drop-in example TypeScript (openai SDK)typescript ``` import OpenAI from "openai"; const client = new OpenAI({ apiKey: "sk-cs4-your-key", baseURL: "https://api.llmsrelay.com/v1", }); const resp = await client.chat.completions.create({ model: "gpt-5.5", messages: [{ role: "user", content: "Explain async/await" }], stream: true, }); ``` ## Supported - Streaming (`stream: true`) and final usage chunks. - System messages and multi-turn conversations. - Function / tool calling in OpenAI format. - `GET /v1/models` for the live model list. For per-IDE setup, see the [OpenAI-Compatible IDEs](/docs/guides/openai-ides) guide. ## Related Articles [OpenAI API Access](/docs/learn/openai-api/)[Codex CLI API](/docs/learn/codex-cli-api/) Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousCodex CLI API](/docs/learn/codex-cli-api/)[Next Cheapest OpenAI API](/docs/learn/cheapest-openai-api/) --- ### Codex CLI API — Run the OpenAI API in OpenAI Codex CLI Canonical URL: https://llmsrelay.com/docs/learn/codex-cli-api/ # Codex CLI API — Run the OpenAI API in OpenAI Codex CLI Connect OpenAI Codex CLI and the Codex extension to the OpenAI API through LLMsRelay's Responses API with a Codex-group key. LLMsRelay is an independently operated API gateway. It is not affiliated with or endorsed by Anthropic, PBC. Product and model names are used only to describe compatibility. > **TL;DR:** - **Endpoint:** POST /v1/responses - **Base URL:** https://api.llmsrelay.com/v1 - **wire_api:** responses - **Models:** gpt-5.5, gpt-5.6-sol, gpt-5.6-terra, gpt-6-sol, gpt-6-astra ## TL;DR OpenAI **Codex CLI** and the **Codex extension** use the Responses API. LLMsRelay Codex serves `/v1/responses` directly, so a single config block points Codex at the OpenAI API. ## Config ~/.codex/config.tomltoml ``` model = "gpt-5.5" model_provider = "llmsrelay" [model_providers.llmsrelay] name = "LLMsRelay Codex" base_url = "https://api.llmsrelay.com/v1" env_key = "LLMSRELAY_API_KEY" wire_api = "responses" ``` shellbash ``` export LLMSRELAY_API_KEY="sk-cs4-your-key" codex ``` ## Full guide For the complete walkthrough — creating a Codex-group key, configuring the VS Code extension, and testing — see the [Codex CLI & Extension](/docs/guides/codex-cli) guide and the [Codex API Reference](/docs/api-reference/codex). ## Related Articles [OpenAI API Access](/docs/learn/openai-api/)[OpenAI-Compatible API](/docs/learn/openai-compatible-api/) Ready to start? Create a key and configure a compatible API route in under 2 minutes. [View Plans](/plans/) [PreviousOpenAI API Access](/docs/learn/openai-api/)[Next OpenAI-Compatible API](/docs/learn/openai-compatible-api/) --- ## Machine-readable resources - llms.txt: https://llmsrelay.com/llms.txt - llms-full.txt: https://llmsrelay.com/llms-full.txt - OpenAPI description: https://llmsrelay.com/openapi.json - API catalog: https://llmsrelay.com/.well-known/api-catalog - Sitemap: https://llmsrelay.com/sitemap.xml - Knowledge graph: https://llmsrelay.com/knowledge-graph.jsonld ## Contact - Telegram: https://telegram.me/claudestorestore - Website: https://llmsrelay.com/