OpenAI compatible API · Attested · Public status

DeepInfra performance

Measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for DeepInfra.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

deepinfra

50 samples

Provider overview

Continuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.

p50 TTFT1645 ms
p95 TTFT8778 ms
p50 TTFB1873 ms
Effective throughput43 tok/s n=15
Uptime96.00%

Measured model routes

Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
qwen/qwen3.5-27b 369 ms 368 ms 100.00% 1
gryphe/mythomax-l2-13b 393 ms 393 ms 100.00% 2
z-ai/glm-4.6 504 ms 504 ms 100.00% 2
qwen/qwen3.5-122b-a10b 664 ms 664 ms 100.00% 1
z-ai/glm-5 742 ms 742 ms 100.00% 2
qwen/qwen3-235b-a22b-thinking-2507 764 ms 764 ms 100.00% 2
openai/gpt-oss-120b 923 ms 923 ms 43 tok/s n=2 100.00% 2
tencent/hy3 1187 ms 1187 ms 53 tok/s n=1 100.00% 1
moonshotai/kimi-k2.5 1188 ms 1188 ms 100.00% 2
qwen/qwen3.5-397b-a17b 1338 ms 1338 ms 100.00% 2
thinkingmachines/inkling-small 1557 ms 1557 ms 63 tok/s n=1 100.00% 3
qwen/qwen3.5-35b-a3b 1571 ms 1571 ms 100.00% 1
google/gemma-3-27b-it 1645 ms 1645 ms 100.00% 1
deepseek/deepseek-v4-pro 1762 ms 1762 ms 92 tok/s n=1 100.00% 1
z-ai/glm-4.7 1874 ms 1873 ms 100.00% 1
deepseek/deepseek-v4-flash 1887 ms 1887 ms 28 tok/s n=2 100.00% 1
google/gemini-2.5-flash 2584 ms 2583 ms 100.00% 1
qwen/qwen3.6-27b 2627 ms 2627 ms 100.00% 1
qwen/qwen3-14b 2696 ms 2696 ms 100.00% 4
minimax/minimax-m2.7 2766 ms 2765 ms 100.00% 2
openai/gpt-oss-20b 2819 ms 2819 ms 100.00% 3
google/gemini-3.1-flash-lite 2872 ms 2872 ms 100.00% 3
nousresearch/hermes-3-llama-3.1-405b 3107 ms 3107 ms 100.00% 3
qwen/qwen3.6-35b-a3b 3189 ms 3189 ms 100.00% 1
deepseek/deepseek-v3.1-terminus 3346 ms 3345 ms 100.00% 1
deepseek/deepseek-v3.2 13567 ms 13567 ms 100.00% 1
qwen/qwen3.5-9b 476 ms 476 ms 66.67% 3
meta-llama/llama-3.1-70b-instruct 1653 ms 1653 ms 50.00% 2
deepseek/deepseek-v4-flash-0731 63 tok/s n=1 0
google/gemma-4-31b-it 41 tok/s n=1 0
minimax/minimax-m3 5 tok/s n=1 0
moonshotai/kimi-k2.6 35 tok/s n=2 0
thinkingmachines/inkling 63 tok/s n=2 0
z-ai/glm-5.2 42 tok/s n=1 0
Workspace access

Sign in

Choose a sign in method. New email and OAuth accounts include $0.10 in starter credit; wallet-only accounts start at $0.

By signing in you agree to the terms of service and privacy policy.