OpenAI compatible API · Attested · Public status
NVIDIA NIM performance
Review measured TTFT, effective throughput, uptime, and sampled model routes for NVIDIA NIM on TrustedRouter using metadata-only production probes.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
NVIDIA NIMnvidia-nim
30 samplesContinuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.
| p50 TTFT | 3056 ms |
|---|---|
| p95 TTFT | 18949 ms |
| Effective throughput | — |
| Uptime | 40.00% |
Measured model routes
| Model | p50 TTFT | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|
| poolside/laguna-xs-2.1 | 711 ms | — | 33.33% | — | 3 |
| meta-llama/llama-3.2-11b-vision-instruct | 745 ms | — | 100.00% | — | 2 |
| nvidia/nemotron-3.5-lightning-30b-a3b | 819 ms | — | 100.00% | — | 1 |
| nvidia/nemotron-3-super-120b-a12b | 1415 ms | — | 100.00% | — | 1 |
| mistralai/mistral-nemotron | 1698 ms | — | 50.00% | — | 2 |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning | 3056 ms | — | 33.33% | — | 3 |
| openai/gpt-oss-20b | 3119 ms | — | 100.00% | — | 3 |
| deepseek/deepseek-v4-flash-0731 | 11204 ms | — | 33.33% | — | 3 |
| nvidia/nemotron-3-ultra-550b-a55b | 14912 ms | — | 100.00% | — | 1 |
| google/gemma-4-31b-it | — | — | 0.00% | — | 2 |
| meta-llama/llama-3.2-90b-vision-instruct | — | — | 0.00% | — | 4 |
| moonshotai/kimi-k3 | — | — | 0.00% | — | 3 |
| z-ai/glm-5.3-flash | — | — | 0.00% | — | 2 |