OpenAI compatible API · Attested · Public status

Venice performance

Review measured TTFT, effective throughput, uptime, and sampled model routes for Venice on TrustedRouter using metadata-only production probes.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

Venicevenice

30 samples

Provider overview

Continuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.

p50 TTFT1878 ms
p95 TTFT16644 ms
Effective throughput51 tok/s n=4
Uptime96.67%

Measured model routes

Modelp50 TTFTEffective throughputUptimeConfig excludedAvailability samples
google/gemma-4-uncensored 724 ms 100.00% 2
qwen/qwen3-coder-480b-a35b-instruct-turbo 928 ms 100.00% 2
qwen/qwen3-235b-a22b-instruct-2507 1051 ms 100.00% 1
deepseek/deepseek-v4-1-flash 1310 ms 100.00% 1
z-ai/glm-4.7-flash 1384 ms 100.00% 1
qwen/qwen-3-8-2-4t-a95b 1398 ms 100.00% 1
qwen/qwen3.5-9b 1578 ms 100.00% 2
qwen/qwen3.5-397b-a17b 1705 ms 100.00% 2
z-ai/glm-4.6 1798 ms 100.00% 1
deepseek/deepseek-v4-flash-0731-fast 1876 ms 100.00% 1
deepseek/deepseek-v4-pro-0423 1878 ms 25 tok/s n=1 100.00% 1
deepseek/deepseek-v4-flash-0731 1929 ms 100.00% 1
moonshotai/kimi-k2-7-code 1972 ms 50.00% 2
qwen/qwen-3-6-plus 2382 ms 100.00% 2
z-ai/glm-5v-turbo 2791 ms 100.00% 2
z-ai/glm-4.7 2865 ms 100.00% 1
qwen/qwen-3-7-max 2914 ms 100.00% 2
z-ai/glm-5 3899 ms 100.00% 1
minimax/minimax-m25 4763 ms 100.00% 1
qwen/qwen-3-8-27b 5079 ms 100.00% 1
minimax/minimax-m27 5303 ms 100.00% 1
meta-llama/llama-3.2-3b 16644 ms 100.00% 1
moonshotai/kimi-k3 76 tok/s n=1 0
qwen/qwen3-235b-a22b-thinking-2507 9 tok/s n=1 0
z-ai/glm-5.2 86 tok/s n=1 0
Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.