DeepSeek: DeepSeek V4 Flash 0731 Performance
Compare measured TTFT, throughput, uptime, and route health for DeepSeek V4 Flash 0731 across TrustedRouter providers using metadata-only production probes.
deepseek/deepseek-v4-flash-0731
Measured performance
Continuously sampled p50/p95 time-to-first-token (TTFT), effective throughput, and success rate for DeepSeek: DeepSeek V4 Flash 0731. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime, and no prompt or output content is stored.
| Provider | p50 TTFT | p95 TTFT | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| relace | 1344 ms | 3984 ms | — | 100.00% | — | 12 |
| confidential-ai | 2088 ms | 4668 ms | 244 tok/s n=2 | 100.00% | — | 14 |
| deepinfra | 2117 ms | 2899 ms | 115 tok/s n=1 | 100.00% | — | 134 |
| baseten | 2190 ms | 4783 ms | — | 99.62% | — | 260 |
| nebius | 4438 ms | 8685 ms | — | 99.26% | — | 269 |
| mancer | 1293 ms | 3698 ms | — | 100.00% | — | 5 |
| sail-research | 1406 ms | 2140 ms | — | 100.00% | — | 4 |
| io-net | 1512 ms | 1512 ms | — | 100.00% | — | 1 |
| fireworks | 1541 ms | 1541 ms | 12 tok/s n=1 | 100.00% | — | 1 |
| pearl | 1610 ms | 4072 ms | 147 tok/s n=1 | 100.00% | — | 8 |
| scaleway | 1685 ms | 4073 ms | — | 100.00% | — | 3 |
| alibaba | 2212 ms | 2611 ms | — | 100.00% | — | 2 |
| engy | 2287 ms | 5648 ms | — | 100.00% | — | 6 |
| nextbit | 2345 ms | 2415 ms | — | 100.00% | — | 3 |
| siliconflow | 2456 ms | 4365 ms | — | 100.00% | — | 3 |
| featherless | 2593 ms | 3647 ms | — | 100.00% | — | 2 |
| nvidia-nim | 3035 ms | 11204 ms | — | 50.00% | — | 4 |
| inceptron | 10217 ms | 12536 ms | — | 87.50% | — | 8 |
Full provider & model leaderboard.
24 routes.
More routes give the auto router more room to fail over around provider 429 and 5xx responses.
Gateway overhead is measured separately.
Public status separates TLS/health overhead from full model latency so slow LLMs do not inflate the router metric.
Metadata rollups.
Status samples store latency, outcome, provider, model, route, cost, and region metadata only.
View public status or inspect provider routes.