OpenAI-compatible. Claude and GPT, unified behind a single gateway. Reliable routing, transparent billing, production-ready from day one.
Technology shouldn't stand above people — it should serve them.
Built for reliability and cost efficiency, so you can focus on your product
Real-time health checks across every upstream provider. Failing channels are isolated automatically — your traffic never notices.
Live usage and billing in your dashboard. No minimums, no hidden fees — every token accounted for.
Swap the base URL and key — your existing SDKs and frameworks keep working, zero code changes required.
Models, pricing, billing, and how to get connected
Create an account and confirm the code we email you, add credit on the wallet page, then create a key on the keys page and pick a group for it. Point your base URL at https://ainzy.net/v1 and you are done. The setup guide has copy-paste configs for Claude Code, Codex, Cursor and the SDKs.
DeepSeek V4 Flash and V4 Pro, DeepSeek V4.1 Flash, GLM 5.3, 5.2 and 5.3 Flash, Kimi K3, GPT-5.4 through GPT-5.6 with the Codex models, and Grok 4.3 to 4.6. The pricing page lists every model with its price per million tokens. Claude models are not available yet.
Three, on the same key and the same host: OpenAI /v1/chat/completions, OpenAI /v1/responses, and Anthropic /v1/messages. Streaming works on all three. Whichever SDK you already use, there is a good chance it works untouched.
Yes. Point them at the Ainzy gateway and they work as-is. The easiest route is CC Switch, a free desktop tool that stores the settings for you and lets you switch providers from the tray. The setup guide also shows the manual config for each tool.
A key belongs to one group, and the group decides which models the key can reach and what you pay. The Chinese-model groups run at 15% of official list price (25% for DeepSeek V4.1 Flash) with cache hits billed at the cache rate, the Codex group at 5%, and the Grok group at 1%. Pick the group when you create the key; one account can hold keys in several groups.
Per token, on actual usage, in US dollars — no subscription and no monthly minimum. Input, output and cache-hit tokens are priced separately. Every request lands in the usage log with its token counts and cost, so the bill is auditable line by line. Failed requests are not billed.
Self-serve on the wallet page: Alipay in CNY at ¥7 per $1, or USDT 1:1. Credit is available immediately after payment. New accounts start at a zero balance, so registering on its own does not grant any usage.
Redundant upstream providers with continuous health monitoring. Failing channels are isolated and routed around automatically, and failed requests are never billed to you.
Start with the usage log: it records the exact error for every call, which usually answers the question on its own. The error table in the setup guide covers the common ones. If that does not settle it, email us with your account email and the time of the request.
Multi-provider failover, transparent billing, 24/7 uptime — this is our first product, and we're serious about teams making the switch. Two grants, ready for you.
For teams and production workloads. Show us your in-progress project or usage history on another platform — dedicated onboarding, approved on the spot.
For independent developers and early-stage projects. Show us your project or usage history elsewhere — qualifying requests are approved as a grant.