OpenAI compatible API · Attested · Public status
Cloudflare Workers AI performance
Measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for Cloudflare Workers AI.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
cloudflare-workers-ai
51 samples
Continuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.
| p50 TTFT | 1502 ms |
|---|---|
| p95 TTFT | 13273 ms |
| p50 TTFB | 1665 ms |
| Effective throughput | 76 tok/s n=3 |
| Uptime | 96.08% |
Measured model routes
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| aisingapore/gemma-sea-lion-v4-27b-it | 879 ms | 879 ms | — | 100.00% | — | 4 |
| meta-llama/llama-3.2-3b-instruct | 1100 ms | 1100 ms | — | 100.00% | — | 4 |
| openai/gpt-oss-120b | 1240 ms | 1240 ms | 76 tok/s n=2 | 100.00% | — | 3 |
| qwen/qwen2.5-coder-32b-instruct | 1260 ms | 1260 ms | — | 100.00% | — | 1 |
| qwen/qwq-32b | 1314 ms | 1314 ms | — | 100.00% | — | 2 |
| meta-llama/llama-3.2-1b-instruct | 1459 ms | 1459 ms | — | 100.00% | — | 5 |
| ibm-granite/granite-4.0-h-micro | 1501 ms | 1501 ms | — | 100.00% | — | 2 |
| meta-llama/llama-4-scout-17b-16e-instruct | 1502 ms | 1502 ms | — | 100.00% | — | 5 |
| mistralai/mistral-small-3.1-24b-instruct | 1511 ms | 1511 ms | — | 100.00% | — | 2 |
| google/gemma-4-26b-a4b-it | 1756 ms | 1756 ms | — | 100.00% | 1 probe_config_error |
2 |
| meta-llama/llama-3.1-8b-instruct-fp8 | 2145 ms | 2145 ms | — | 100.00% | — | 3 |
| moonshotai/kimi-k3 | 2269 ms | 2269 ms | 25 tok/s n=1 | 100.00% | — | 2 |
| z-ai/glm-4.7-flash | 2494 ms | 2494 ms | — | 100.00% | — | 3 |
| deepseek/deepseek-r1-distill-qwen-32b | 2715 ms | 2715 ms | — | 100.00% | — | 3 |
| openai/gpt-oss-20b | 2730 ms | 2730 ms | — | 100.00% | — | 1 |
| qwen/qwen3-30b-a3b-fp8 | 2863 ms | 2863 ms | — | 100.00% | — | 3 |
| meta-llama/llama-3.3-70b-instruct-fp8-fast | 816 ms | 816 ms | — | 75.00% | — | 4 |
| nvidia/nemotron-3-120b-a12b | 13273 ms | 13272 ms | — | 50.00% | — | 2 |