MoonshotAI: Kimi K2.6 Performance
Compare measured TTFT, throughput, uptime, and route health for MoonshotAI: Kimi K2.6 across TrustedRouter providers using metadata-only production probes.
moonshotai/kimi-k2.6
Measured performance
Continuously sampled p50/p95 time-to-first-token (TTFT), effective throughput, and success rate for MoonshotAI: Kimi K2.6. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime, and no prompt or output content is stored.
| Provider | p50 TTFT | p95 TTFT | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| kimi | 2783 ms | 10688 ms | 31 tok/s n=2 | 100.00% | — | 14 |
| novita | 1426 ms | 1426 ms | 32 tok/s n=2 | 100.00% | — | 1 |
| parasail | 1474 ms | 3661 ms | — | 66.67% | — | 3 |
| sail-research | 1499 ms | 3767 ms | 46 tok/s n=1 | 100.00% | — | 4 |
| fireworks | 1586 ms | 1726 ms | 48 tok/s n=1 | 100.00% | — | 2 |
| digitalocean | 1587 ms | 3601 ms | 16 tok/s n=1 | 100.00% | — | 2 |
| atlas-cloud | 1737 ms | 1737 ms | 32 tok/s n=2 | 100.00% | — | 1 |
| baseten | 1813 ms | 1813 ms | — | 100.00% | — | 2 |
| inceptron | 1813 ms | 4549 ms | — | 100.00% | — | 2 |
| telnyx | 1838 ms | 1856 ms | — | 100.00% | — | 2 |
| wafer | 1913 ms | 3663 ms | 140 tok/s n=1 | 100.00% | — | 7 |
| azure | 1960 ms | 3695 ms | 101 tok/s n=2 | 100.00% | — | 3 |
| wandb | 2659 ms | 2659 ms | — | 100.00% | — | 1 |
| deepinfra | 3050 ms | 3158 ms | 23 tok/s n=1 | 100.00% | — | 3 |
| featherless | 4849 ms | 6028 ms | — | 100.00% | — | 3 |
| chutes | 7173 ms | 12782 ms | 18 tok/s n=2 | 100.00% | — | 4 |
| io-net | — | — | 47 tok/s n=1 | — | — | 0 |
| siliconflow | — | — | 27 tok/s n=2 | — | — | 0 |
Full provider & model leaderboard.
20 routes.
More routes give the auto router more room to fail over around provider 429 and 5xx responses.
Gateway overhead is measured separately.
Public status separates TLS/health overhead from full model latency so slow LLMs do not inflate the router metric.
Metadata rollups.
Status samples store latency, outcome, provider, model, route, cost, and region metadata only.
View public status or inspect provider routes.