One key.
600+ models.
An OpenAI-compatible LLM gateway. Point your existing OpenAI SDK at one base URL and reach GPT, Claude, Gemini, Grok, DeepSeek, GLM, Kimi, MiniMax, Qwen, and 600+ more, at a flat 5%+ discount off list.
The routing fee is the only new cost.
Drop it into what you're running today.
Drop-in OpenAI compatibility
POST /v1/chat/completionsPoint the OpenAI SDK at the base URL: same request shape, same response shape, everything else unchanged. View in the docs
Claude over the same endpoint
model: "anthropic/..."Reach Claude the same way as everything else: the OpenAI-compatible endpoint, not a separate Anthropic dialect. View in the docs
Smart auto-routing
model: "auto"model: "auto" scores each prompt's difficulty and routes to the best-fit model automatically. View in the docs
Automatic failover
If a provider errors or rate-limits mid-request, Singular retries the next healthy provider. View in the docs
Fallback model list
models: ["gpt-5.1", "claude-sonnet-5"]Pass an ordered models array instead of one model: Singular tries each in order until one succeeds. View in the docs
Routing profiles
routing_profile: "eco"Bias auto-routing toward cost (eco), balance (auto), quality (premium), or tool use (agentic). View in the docs
Provider preferences
Pin, exclude, or sort providers; require zero data retention; cap price per call. View in the docs
Routing suffixes
Append :nitro for throughput, :floor for price, or :exacto for strict tool compatibility. View in the docs
Streaming
Token-by-token SSE on every chat model. View in the docs
Tools & structured output
Function calling and JSON-schema outputs pass through on a single turn. Multi-turn tool history isn't preserved yet, and structured-output repair is best effort: validate JSON on your side. View in the docs
Machine-readable pricing feed
GET /pricing.jsonFetch live retail rates as timestamped JSON: build cost-aware routing or a budget dashboard without scraping a page. View in the docs
Real-cost billing
tokens × rate ÷ 1MEvery call decrements your balance by the model's actual cost: nothing rounded up. View in the docs
We'll tell you what's actually ready.
/v1/chat/completions (stable, verified)/v1/embeddings (registered, not yet verified end-to-end)/v1/rerank (registered, not yet verified end-to-end)/v1/images/generations · /v1/images/edits (registered, not yet verified end-to-end)/v1/audio/speech (registered, not yet verified end-to-end)/v1/audio/transcriptions · /v1/audio/translations (registered, not yet verified end-to-end)/v1/moderations (registered, not yet verified end-to-end)On the roadmap: ask us for early access.Chat Completions is Singular's only verified, stable surface today. The other routes exist on the same key, but don't yet have a verified, stable, end-to-end path in production. Treat them as unverified until you've tested your specific case, and check /docs for the current status before depending on one.
Know what to retry before you ship.
A 401 includes WWW-Authenticate: Bearer; a 429 usually includes Retry-After. Branch on HTTP status first.
| Status | Meaning | Retry? |
|---|---|---|
| 400 / 422 | Invalid request or unsupported shape | No: fix the request |
| 401 | Missing or invalid key | No: replace credentials |
| 402 | Payment or balance action required | No blind retry: fund the key |
| 403 | Revoked key or not permitted | No: replace key or change policy |
| 404 | Unknown resource or unavailable surface | No |
| 408 | Request timeout | Usually, with a bound |
| 429 | Rate limited | Yes: honor Retry-After |
| 5xx | Singular or upstream failure | Usually: backoff with jitter |
Below official rates on every model.
A snapshot of flagship models: frontier labs at 5% off list, open-weight cost-leaders deeper still.
| Model | Official (in / out) | Singular (in / out) | You save |
|---|---|---|---|
| OpenAI GPT-5.5 | $5.00 / $30.00 | $4.75 / $28.50 | 5% |
| OpenAI GPT-5.1 | $1.25 / $10.00 | $1.19 / $9.50 | 5% |
| OpenAI GPT-5 mini | $0.25 / $2.00 | $0.238 / $1.90 | 5% |
| Anthropic Claude Fable 5 | $10.00 / $50.00 | $9.50 / $47.50 | 5% |
| Anthropic Claude Sonnet 5* | $2.00 / $10.00 | $1.90 / $9.50 | 5% |
| Anthropic Claude Haiku 4.5 | $1.00 / $5.00 | $0.95 / $4.75 | 5% |
| Google Gemini 3.1 Pro | $2.00 / $12.00 | $1.90 / $11.40 | 5% |
| xAI Grok 4.3 | $1.25 / $2.50 | $1.19 / $2.38 | 5% |
| Z.AI GLM-5.2 | $1.40 / $4.40 | $0.855 / $2.57 | 39–42% |
| Moonshot Kimi K2.6 | $0.95 / $4.00 | $0.475 / $2.47 | 38–50% |
| DeepSeek V3.1 | $0.21 / $0.79 | $0.19 / $0.665 | 10–16% |
| Alibaba Qwen3.5-Plus | $0.40 / $2.40 | $0.38 / $2.28 | 5% |
Exact cost per call: (prompt_tokens × prompt_rate + completion_tokens × completion_rate) / 1,000,000, in the model's listed USD-per-million rate. Pull the same rates Singular uses as timestamped JSON at GET /pricing.json to build your own cost-aware routing or budget alerts.
* Claude Sonnet 5 shown at Anthropic's current intro rate ($2 / $10); their standard list ($3 / $15) resumes August 31, 2026. Point-in-time snapshot. Rates live-update at /pricing. 600+ models total, all priced below their provider's official rate.
View full pricingSame key, different workloads.
Drop-in SDK replacement
Swap the base URL in an existing OpenAI integration: nothing else changes.
Multi-model apps & agents
Confirmed with OpenWebUI and n8n; any OpenAI-compatible client or coding agent that supports Chat Completions and a custom base URL should work too.
Cost-optimized production routing
Let auto pick the cheapest model that can do the job, request by request.
Privacy-sensitive workloads
Require zero-data-retention providers per request via the provider preference object.
What a security review will ask for anyway.
From sign-up to first response.
- 1Create a Singular account.Starts with $1.50 in free credit. No card required.
- 2Add credit in Settings → Billing.Pick a preset ($5 / $25 / $50 / $100) or a custom amount, then check out.
- 3Pay by card or PayPal.Funds apply to your balance once checkout settles. Docs is explicit that crypto payment doesn't top up a balance: it pays for one call at a time, separately.
- 4Issue a key in Settings → API keys.Shown once: copy and save it immediately. A lost key can't be recovered.
curl https://api.impossi.build/v1/chat/completions \
-H "Authorization: Bearer $SINGULAR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"auto","messages":[{"role":"user","content":"Hello!"}]}'Questions, answered plainly.
What is Singular's API?
How does pricing/billing work? Is it a subscription?
What happens if I lose my API key?
Which clients and SDKs work out of the box?
How does auto routing decide which model to use?
Is my data used to train models?
Can I bring my own provider keys?
What's the difference between the Singular chat app and the API?
Are the docs readable by AI agents and crawlers?
What will you build with 600+ models behind one key?
curl https://api.impossi.build/v1/chat/completions \
-H "Authorization: Bearer $SINGULAR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"auto","messages":[{"role":"user","content":"Hello!"}]}'