OpenAI compatible API · Attested · Public status

Baseten

Baseten models on TrustedRouter with prices, routes, policy notes, and source links.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

baseten

No provider claim

All providers

ProviderBaseten
Models12 public models
Prepaid routes12
BYOK routes12
Zero data retentionnot claimed
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteNo provider-ZDR claim is tracked here. Baseten's inference and security documentation are linked for users who need to review API data handling.
Policy source

Measured performance

64 samples

Continuously sampled across Baseten's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT2016 ms
Effective throughput73 tok/s n=16
Uptime95.31%
Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
thinkingmachines/inkling-1m 847 ms 847 ms 64 tok/s n=1 100.00% 1 router_error 4
z-ai/glm-4.7 1476 ms 1475 ms 100.00% 8
thinkingmachines/inkling-small 1484 ms 1484 ms 69 tok/s n=1 100.00% 6
openai/gpt-oss-120b 1590 ms 1590 ms 91 tok/s n=2 100.00% 1
nvidia/nemotron-3-ultra-550b-a55b 1796 ms 1796 ms 130 tok/s n=2 100.00% 6
deepseek/deepseek-v4-pro 2016 ms 2016 ms 46 tok/s n=1 100.00% 8
moonshotai/kimi-k3 2117 ms 2117 ms 81 tok/s n=2 100.00% 1 router_error 8
z-ai/glm-5.2 2912 ms 2912 ms 56 tok/s n=1 100.00% 6
z-ai/glm-5.2-fast 3179 ms 3179 ms 70 tok/s n=2 87.50% 8
moonshotai/kimi-k2.6 3440 ms 3440 ms 76 tok/s n=2 83.33% 6
moonshotai/kimi-k2.7-code 567 ms 567 ms 40 tok/s n=2 66.67% 3

Baseten performance history · Full provider & model leaderboard.

Provider models

Models served by Baseten.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Endpoints Prompt Completion Routes
deepseek/deepseek-v4-flash-0731
DeepSeek: DeepSeek V4 Flash 0731
1,048,576 2 $0.1365/1M $0.273/1M prepaid BYOK
deepseek/deepseek-v4-pro
DeepSeek: DeepSeek V4 Pro
IQ 115#30 1,048,576 2 $1.827/1M $3.654/1M prepaid BYOK
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#19 262,144 2 $0.9975/1M $4.2/1M prepaid BYOK
moonshotai/kimi-k2.7-code
MoonshotAI: Kimi K2.7 Code
IQ 118#22 262,144 2 $0.9975/1M $4.2/1M prepaid BYOK
moonshotai/kimi-k3
MoonshotAI: Kimi K3
IQ 122#15 1,048,576 2 $3.15/1M $15.75/1M prepaid BYOK
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA: Nemotron 3 Ultra
512,288 2 $0.63/1M $2.52/1M prepaid BYOK
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 105#52 131,072 2 $0.105/1M $0.525/1M prepaid BYOK
thinkingmachines/inkling-1m
Inkling
1,048,576 2 $1.05/1M $4.2525/1M prepaid BYOK
thinkingmachines/inkling-small
Thinking Machines: Inkling Small
524,288 2 $0.525/1M $1.26/1M prepaid BYOK
z-ai/glm-4.7
Z.ai: GLM 4.7
IQ 103#58 204,800 2 $0.63/1M $2.31/1M prepaid BYOK
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#16 1,048,576 2 $1.47/1M $4.62/1M prepaid BYOK
z-ai/glm-5.2-fast
GLM 5.2 Fast on Fireworks
1,048,576 2 $2.205/1M $6.93/1M prepaid BYOK
Workspace access

Sign in

Choose a sign in method. New email and OAuth accounts include $0.10 in starter credit; wallet-only accounts start at $0.

By signing in you agree to the terms of service and privacy policy.