Together
Explore Together models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.
Togethertogether
ZDRThese privacy labels describe Together, the upstream model provider. ZDR is a retention policy; verified confidential inference additionally requires attested provider compute and end-to-end encryption.
| Provider | Together |
|---|---|
| Routing status | Active |
| Provider website | https://www.together.ai/ |
| Models | 12 public models |
| Credits routes | 12 |
| Zero data retention | yes All routes |
| Verified confidential inference | Not verified |
| Policy note | Tracked as provider ZDR. Together documents that inference inputs and outputs are not stored by default; temporary prompt caching may be used for performance, and sharing content for training is opt-in. Policy source |
Measured performance
125 samplesContinuously sampled across Together's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 2834 ms |
|---|---|
| Effective throughput | 89 tok/s n=4 |
| Uptime | 100.00% |
| Model | p50 TTFT | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|
| minimax/minimax-m3 | 1599 ms | 234 tok/s n=1 | 100.00% | — | 30 |
| prism-ml/ternary-bonsai-27b | 2834 ms | — | 100.00% | — | 77 |
| meta-llama/llama-3.3-70b-instruct | 797 ms | — | 100.00% | — | 2 |
| z-ai/glm-5.2 | 1094 ms | 77 tok/s n=1 | 100.00% | — | 4 |
| qwen/qwen3.8-2.4t-a95b | 1227 ms | — | 100.00% | — | 2 |
| moonshotai/kimi-k3 | 1585 ms | 101 tok/s n=1 | 100.00% | — | 2 |
| openai/gpt-oss-120b | 1826 ms | 47 tok/s n=1 | 100.00% | — | 1 |
| thinkingmachines/inkling | 2302 ms | — | 100.00% | — | 2 |
| qwen/qwen3.5-9b | 2594 ms | — | 100.00% | — | 2 |
| z-ai/glm-5.3-flash | 2707 ms | — | 100.00% | — | 3 |
Together performance history · Full provider & model leaderboard.
Models served by Together.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Input | Cached input | Output |
|---|---|---|---|---|---|
deepseek/deepseek-v4.1-flashDeepSeek: DeepSeek V4.1 Flash |
IQ 116#38 | 1,048,576 | $0.3165/1M | $0.01/1M | $1.266/1M |
meta-llama/llama-3.3-70b-instructMeta: Llama 3.3 70B Instruct |
— | 131,072 | $1.0972/1M | Not published | $1.0972/1M |
meta-models/muse-glimmer-30bMuse Glimmer 30B on Fireworks |
— | 131,072 | $0.36925/1M | $0.0422/1M | $1.5825/1M |
minimax/minimax-m3MiniMax: MiniMax M3 |
IQ 115#45 | 524,288 | $0.3165/1M | $0.0633/1M | $1.266/1M |
moonshotai/kimi-k3MoonshotAI: Kimi K3 |
IQ 121#27 | 1,048,576 | $3.165/1M | $0.3165/1M | $15.825/1M |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 98#99 | 131,072 | $0.15825/1M | Not published | $0.633/1M |
prism-ml/ternary-bonsai-27bTernary Bonsai 27B |
— | 262,144 | $0.01/1M | Not published | $0.01/1M |
qwen/qwen3.5-9bQwen: Qwen3.5-9B |
IQ 95#108 | 262,144 | $0.17935/1M | Not published | $0.26375/1M |
qwen/qwen3.8-2.4t-a95bQwen: Qwen3.8 2.4T A95B |
IQ 122#24 | 262,144 | $2.11/1M | $0.26375/1M | $6.33/1M |
thinkingmachines/inklingThinking Machines: Inkling |
IQ 107#70 | 524,288 | $1.055/1M | $0.17935/1M | $4.27275/1M |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#28 | 1,048,576 | $1.477/1M | $0.2743/1M | $4.642/1M |
z-ai/glm-5.3-flashZ.ai: GLM 5.3 Flash |
IQ 116#40 | 1,048,576 | $0.15825/1M | $0.03165/1M | $0.5275/1M |
Questions
Does Together have zero data retention?
TrustedRouter records Together as supporting provider-level zero data retention based on the policy source linked on this page. This is a provider policy claim, separate from TrustedRouter's content-stateless real-time gateway and from end-to-end confidential compute.
Is Together end-to-end encrypted?
TrustedRouter does not currently mark Together as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.
Which Together models are available through TrustedRouter?
This page currently lists 12 public Together models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.