OpenAI-compatible · one endpoint · automatic failover

Every model,
one router.

Infero — from the Latin inferre, to bring in, to conclude — routes your requests across dozens of frontier models through a single OpenAI-compatible surface. One key, one bill, zero rewrites. If a route goes down, the next one picks up mid-sentence.

40+curated models
1base URL to remember
0SDK rewrites
24/7route health checks
§ 01 — Catalogue

One endpoint, every family.

Frontier proprietary models, open-weight reasoning, fast coders and cheap workhorses — curated, versioned, and reachable through the same /v1/chat/completions you already call.

Proprietary

GPT · o-series

GPT-5.5, 5.6 Sol/Luna/Terra, GPT-6 Astra — flagship reasoning and vision, served straight from upstream.

Proprietary

Claude

Sonnet and Opus class models for long-context agentic work, code review and careful writing.

Proprietary

Gemini

Gemini Flash tier for high-volume, low-latency pipelines — billed per token like everything else.

Open weights

DeepSeek V4

Pro and Flash, tools-enabled, million-token context at a fraction of closed-model pricing.

Open weights

GLM · Kimi · Qwen

GLM-5 series, Kimi K2.7-Code, Qwen 3.x — the strongest open coders and reasoners, always current.

Open weights

MiniMax · Step · LongCat

Specialist and experimental models cycled in as they earn their place on the router.

§ 02 — Mechanics

A line of code,
then you're routing.

i.

Point the base URL

Swap api.openai.com for api.infero.sbs/v1. OpenAI, OpenRouter and every compatible client just works.

ii.

Call any model

Use one key across the whole catalogue. Combos route to the best live provider automatically.

iii.

Failover is silent

Dead routes are pulled and retried upstream-side. Your app sees one steady endpoint, not a zoo of providers.

~/infero · cURL
$ curl https://api.infero.sbs/v1/chat/completions \
  -H "Authorization: Bearer inf_…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "messages": [
      { "role": "user", "content": "Route me somewhere fast." }
    ],
    "stream": true
  }'

> data: {"id":"chatcmpl-…","provider":"infero"}
> data:   "choices":[{"delta":{"content":"Routed."}}]}
> data: [DONE]
§ 03 — Pricing

Pay for tokens.
Nothing else.

Transparent per-million-token pricing, prepaid credit, no top-up fee, no per-request surcharge. Route intelligently: cheap models for volume, flagships for the hard turns.

ModelContextInput / MTokOutput / MTok
GPT-5.6 Sol gpt-5.6-sol1M$5.00$30.00
GPT-5.6 Luna gpt-5.6-luna400k$1.00$6.00
Claude Sonnet class claude-sonnet200k$3.00$15.00
DeepSeek V4 Pro deepseek-v4-pro1M$1.74$3.48
DeepSeek V4 Flash deepseek-v4-flash1M$0.19$0.51
Kimi K2.7 Code kimi-k2.7-code262k$0.95$4.00
GLM-5 Flash glm-5-flash128k$0.10$0.40
Gemini Flash class gemini-flash1M$0.30$2.50
$0–$500 / moPay-as-you-go · listed price
$500+ / moScale · 5% off, auto-applied
$2,000+ / moVolume · 10% off
$10,000+ / moEnterprise · 15% off
§ 04 — Start

Claim your key.

Point any OpenAI-compatible client at https://api.infero.sbs/v1 with your bearer token and go. Works with OpenAI SDKs, LangChain, Cline, OpenCode, anything that speaks the protocol.

python · openai sdk
from openai import OpenAI

client = OpenAI(
    base_url="https://api.infero.sbs/v1",
    api_key="inf_…",
)

r = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "Hello, router."}],
)
print(r.choices[0].message.content)