OpenAI compatible API · Attested · Public status

Cloudflare Workers AI performance

Measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for Cloudflare Workers AI.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

cloudflare-workers-ai

51 samples

Provider overview

Continuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.

p50 TTFT1502 ms
p95 TTFT13273 ms
p50 TTFB1665 ms
Effective throughput76 tok/s n=3
Uptime96.08%

Measured model routes

Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
aisingapore/gemma-sea-lion-v4-27b-it 879 ms 879 ms 100.00% 4
meta-llama/llama-3.2-3b-instruct 1100 ms 1100 ms 100.00% 4
openai/gpt-oss-120b 1240 ms 1240 ms 76 tok/s n=2 100.00% 3
qwen/qwen2.5-coder-32b-instruct 1260 ms 1260 ms 100.00% 1
qwen/qwq-32b 1314 ms 1314 ms 100.00% 2
meta-llama/llama-3.2-1b-instruct 1459 ms 1459 ms 100.00% 5
ibm-granite/granite-4.0-h-micro 1501 ms 1501 ms 100.00% 2
meta-llama/llama-4-scout-17b-16e-instruct 1502 ms 1502 ms 100.00% 5
mistralai/mistral-small-3.1-24b-instruct 1511 ms 1511 ms 100.00% 2
google/gemma-4-26b-a4b-it 1756 ms 1756 ms 100.00% 1 probe_config_error 2
meta-llama/llama-3.1-8b-instruct-fp8 2145 ms 2145 ms 100.00% 3
moonshotai/kimi-k3 2269 ms 2269 ms 25 tok/s n=1 100.00% 2
z-ai/glm-4.7-flash 2494 ms 2494 ms 100.00% 3
deepseek/deepseek-r1-distill-qwen-32b 2715 ms 2715 ms 100.00% 3
openai/gpt-oss-20b 2730 ms 2730 ms 100.00% 1
qwen/qwen3-30b-a3b-fp8 2863 ms 2863 ms 100.00% 3
meta-llama/llama-3.3-70b-instruct-fp8-fast 816 ms 816 ms 75.00% 4
nvidia/nemotron-3-120b-a12b 13273 ms 13272 ms 50.00% 2
Workspace access

Sign in

Choose a sign in method. New email and OAuth accounts include $0.10 in starter credit; wallet-only accounts start at $0.

By signing in you agree to the terms of service and privacy policy.