SingularSingular
Menu
Singular API: Beta

One key.
600+ models.

An OpenAI-compatible LLM gateway. Point your existing OpenAI SDK at one base URL and reach GPT, Claude, Gemini, Grok, DeepSeek, GLM, Kimi, MiniMax, Qwen, and 600+ more, at a flat 5%+ discount off list.

quickstart.sh
model: "auto". One request, scored and routed live across 600+ models
OpenAI
Anthropic
Google
DeepSeek
xAI
Meta
Z.AI
Moonshot
Xiaomi MiMo
routed to Anthropic · streaming response
Why Singular's API

The routing fee is the only new cost.

0+
models, one integration
0–0%
below official rates: 5% on frontier labs, deeper on open-weight models
$0
platform, seat, and minimum fees
Built for how you already build

Drop it into what you're running today.

Drop-in OpenAI compatibility

POST /v1/chat/completions

Point the OpenAI SDK at the base URL: same request shape, same response shape, everything else unchanged. View in the docs

Claude over the same endpoint

model: "anthropic/..."

Reach Claude the same way as everything else: the OpenAI-compatible endpoint, not a separate Anthropic dialect. View in the docs

Smart auto-routing

model: "auto"

model: "auto" scores each prompt's difficulty and routes to the best-fit model automatically. View in the docs

Automatic failover

If a provider errors or rate-limits mid-request, Singular retries the next healthy provider. View in the docs

Fallback model list

models: ["gpt-5.1", "claude-sonnet-5"]

Pass an ordered models array instead of one model: Singular tries each in order until one succeeds. View in the docs

Routing profiles

routing_profile: "eco"

Bias auto-routing toward cost (eco), balance (auto), quality (premium), or tool use (agentic). View in the docs

Provider preferences

Pin, exclude, or sort providers; require zero data retention; cap price per call. View in the docs

Routing suffixes

Append :nitro for throughput, :floor for price, or :exacto for strict tool compatibility. View in the docs

Streaming

Token-by-token SSE on every chat model. View in the docs

Tools & structured output

Function calling and JSON-schema outputs pass through on a single turn. Multi-turn tool history isn't preserved yet, and structured-output repair is best effort: validate JSON on your side. View in the docs

Machine-readable pricing feed

GET /pricing.json

Fetch live retail rates as timestamped JSON: build cost-aware routing or a budget dashboard without scraping a page. View in the docs

Real-cost billing

tokens × rate ÷ 1M

Every call decrements your balance by the model's actual cost: nothing rounded up. View in the docs

Beyond chat

We'll tell you what's actually ready.

Text & chat
/v1/chat/completions (stable, verified)
Embeddings
/v1/embeddings (registered, not yet verified end-to-end)
Rerank
/v1/rerank (registered, not yet verified end-to-end)
Image generation
/v1/images/generations · /v1/images/edits (registered, not yet verified end-to-end)
Speech (TTS)
/v1/audio/speech (registered, not yet verified end-to-end)
Transcription / translation
/v1/audio/transcriptions · /v1/audio/translations (registered, not yet verified end-to-end)
Moderation
/v1/moderations (registered, not yet verified end-to-end)
Video
On the roadmap: ask us for early access.

Chat Completions is Singular's only verified, stable surface today. The other routes exist on the same key, but don't yet have a verified, stable, end-to-end path in production. Treat them as unverified until you've tested your specific case, and check /docs for the current status before depending on one.

Errors & retries

Know what to retry before you ship.

A 401 includes WWW-Authenticate: Bearer; a 429 usually includes Retry-After. Branch on HTTP status first.

StatusMeaningRetry?
400 / 422Invalid request or unsupported shapeNo: fix the request
401Missing or invalid keyNo: replace credentials
402Payment or balance action requiredNo blind retry: fund the key
403Revoked key or not permittedNo: replace key or change policy
404Unknown resource or unavailable surfaceNo
408Request timeoutUsually, with a bound
429Rate limitedYes: honor Retry-After
5xxSingular or upstream failureUsually: backoff with jitter
Pricing

Below official rates on every model.

A snapshot of flagship models: frontier labs at 5% off list, open-weight cost-leaders deeper still.

ModelOfficial (in / out)Singular (in / out)You save
OpenAI GPT-5.5$5.00 / $30.00$4.75 / $28.505%
OpenAI GPT-5.1$1.25 / $10.00$1.19 / $9.505%
OpenAI GPT-5 mini$0.25 / $2.00$0.238 / $1.905%
Anthropic Claude Fable 5$10.00 / $50.00$9.50 / $47.505%
Anthropic Claude Sonnet 5*$2.00 / $10.00$1.90 / $9.505%
Anthropic Claude Haiku 4.5$1.00 / $5.00$0.95 / $4.755%
Google Gemini 3.1 Pro$2.00 / $12.00$1.90 / $11.405%
xAI Grok 4.3$1.25 / $2.50$1.19 / $2.385%
Z.AI GLM-5.2$1.40 / $4.40$0.855 / $2.5739–42%
Moonshot Kimi K2.6$0.95 / $4.00$0.475 / $2.4738–50%
DeepSeek V3.1$0.21 / $0.79$0.19 / $0.66510–16%
Alibaba Qwen3.5-Plus$0.40 / $2.40$0.38 / $2.285%
OpenAI GPT-5.55%
Official$5.00 / $30.00
Singular$4.75 / $28.50
OpenAI GPT-5.15%
Official$1.25 / $10.00
Singular$1.19 / $9.50
OpenAI GPT-5 mini5%
Official$0.25 / $2.00
Singular$0.238 / $1.90
Anthropic Claude Fable 55%
Official$10.00 / $50.00
Singular$9.50 / $47.50
Anthropic Claude Sonnet 5*5%
Official$2.00 / $10.00
Singular$1.90 / $9.50
Anthropic Claude Haiku 4.55%
Official$1.00 / $5.00
Singular$0.95 / $4.75
Google Gemini 3.1 Pro5%
Official$2.00 / $12.00
Singular$1.90 / $11.40
xAI Grok 4.35%
Official$1.25 / $2.50
Singular$1.19 / $2.38
Z.AI GLM-5.239–42%
Official$1.40 / $4.40
Singular$0.855 / $2.57
Moonshot Kimi K2.638–50%
Official$0.95 / $4.00
Singular$0.475 / $2.47
DeepSeek V3.110–16%
Official$0.21 / $0.79
Singular$0.19 / $0.665
Alibaba Qwen3.5-Plus5%
Official$0.40 / $2.40
Singular$0.38 / $2.28

Exact cost per call: (prompt_tokens × prompt_rate + completion_tokens × completion_rate) / 1,000,000, in the model's listed USD-per-million rate. Pull the same rates Singular uses as timestamped JSON at GET /pricing.json to build your own cost-aware routing or budget alerts.

* Claude Sonnet 5 shown at Anthropic's current intro rate ($2 / $10); their standard list ($3 / $15) resumes August 31, 2026. Point-in-time snapshot. Rates live-update at /pricing. 600+ models total, all priced below their provider's official rate.

View full pricing
What people build with it

Same key, different workloads.

Drop-in SDK replacement

Swap the base URL in an existing OpenAI integration: nothing else changes.

Multi-model apps & agents

Confirmed with OpenWebUI and n8n; any OpenAI-compatible client or coding agent that supports Chat Completions and a custom base URL should work too.

Cost-optimized production routing

Let auto pick the cheapest model that can do the job, request by request.

Privacy-sensitive workloads

Require zero-data-retention providers per request via the provider preference object.

Security & privacy

What a security review will ask for anyway.

Zero retention (opt-in)
Require zero-data-retention providers only, per request, via the provider preference object.
BYOK
Bring your own provider keys and route through Singular; the provider bills you directly.
Server-side only
Load keys from a secret manager or server-only environment variable, never client-side code.
Separate keys per environment
Create independent keys for development, staging, production, and each service; revoke one without touching the rest.
Revoke on exposure
A key that leaks into a client bundle, log, paste, or repository should be revoked and replaced immediately.
Bounded retries
Cap retries so an application can't multiply spend during an upstream incident.
Get your key

From sign-up to first response.

  1. 1
    Create a Singular account.
    Starts with $1.50 in free credit. No card required.
  2. 2
    Add credit in Settings → Billing.
    Pick a preset ($5 / $25 / $50 / $100) or a custom amount, then check out.
  3. 3
    Pay by card or PayPal.
    Funds apply to your balance once checkout settles. Docs is explicit that crypto payment doesn't top up a balance: it pays for one call at a time, separately.
  4. 4
    Issue a key in Settings → API keys.
    Shown once: copy and save it immediately. A lost key can't be recovered.
bash
curl https://api.impossi.build/v1/chat/completions \
  -H "Authorization: Bearer $SINGULAR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"auto","messages":[{"role":"user","content":"Hello!"}]}'
FAQ

Questions, answered plainly.

What is Singular's API?
An OpenAI-compatible LLM API gateway: one base URL, one key, one prepaid balance, and access to 600+ models from every major lab, including Claude.
How does pricing/billing work? Is it a subscription?
No subscription. It's prepaid, pay-as-you-go: top up a balance and every call decrements it by the model's real cost, at a flat 5%+ discount off official rates. No seats, no platform fee, no minimum.
What happens if I lose my API key?
Keys are shown once, at creation, and can't be recovered if lost: revoke it and issue a new one. The platform is in beta; full dashboard/account-recovery tooling is coming.
Which clients and SDKs work out of the box?
Anything built for the OpenAI SDK: point the client at Singular's base URL and it works unmodified, including for Claude models, which run over the same Chat Completions endpoint rather than the Anthropic SDK's own dialect. Confirmed with OpenWebUI and n8n; a coding agent or automation tool works if it can use Chat Completions, override the base URL, accept a custom model ID, and tolerate an opaque /v1/models result. Verify against your specific client, since not every tool marketed as "OpenAI compatible" only needs Chat Completions.
How does auto routing decide which model to use?
Set model: "auto" and Singular scores each prompt's difficulty and routes it to the best-fit model automatically, with automatic failover to the next healthy provider if one errors or rate-limits.
Is my data used to train models?
There's no documented blanket guarantee either way, so don't take silence as a yes or a no. What is documented: you can require zero-data-retention providers only, per request, by setting the provider preference object's zdr flag. Singular won't route that request to a provider that doesn't meet it.
Can I bring my own provider keys?
Yes. BYOK routes through Singular on your own provider keys and the provider bills you directly for that usage; Singular still meters and bills its own routing fee regardless.
What's the difference between the Singular chat app and the API?
The chat app is the consumer product: pick a model, auto-route, or compare side-by-side in a thread. The API is the same routing engine as a programmatic gateway, for wiring into your own app, agent, or workflow.
Are the docs readable by AI agents and crawlers?
Yes. /llms.txt is a curated Markdown entrypoint, /docs/llms.txt indexes the full technical reference, and /llms-full.txt bundles every documentation page as one Markdown file.

What will you build with 600+ models behind one key?

quickstart.sh
curl https://api.impossi.build/v1/chat/completions \
  -H "Authorization: Bearer $SINGULAR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"auto","messages":[{"role":"user","content":"Hello!"}]}'