Together AI Pricing Sheet & Dedicated Endpoints 2026
Over 100+ open-source foundation models hosted on high-performance cloud clusters. Verified daily against official cloud provider documentation.
Together AI Production Pricing Matrix
Real-time unit costs, context sizes, operational throughput, and optimal use-case recommendations.
| Service / Model Tier | Primary Rate | Secondary / Output Rate | Context / Capacity | Throughput / Latency | Optimal Workload |
|---|---|---|---|---|---|
| DeepSeek V3 (671B MoE) | $0.60 | $0.60 | 64,000 | ~70 tok/s | Cost-effective open-weights MoE |
| DeepSeek R1 (Full 671B) | $1.50 | $1.50 | 64,000 | ~45 tok/s | Full-scale open-weights reasoning |
| Llama 3.3 70B Instruct Turbo | $0.60 | $0.60 | 128,000 | ~110 tok/s | Ultra-low-latency 70B inference |
| Qwen 2.5 72B Instruct | $0.60 | $0.60 | 32,000 | ~90 tok/s | Multilingual and mathematical tasks |
| FLUX.1 Schnell (Image) | $0.003 / image | N/A | N/A | ~1.2s / img | High-speed 4-step diffusion generation |
Architectural & Financial Billing Nuances
Committed Use & Volume Discounts
Enterprise accounts spending over $2,000/mo typically qualify for 20% to 45% volume concessions or reserved throughput capacity agreements.
SLA & Latency Guarantees
Standard tiers offer 99.9% availability. Dedicated instances provide private VPC endpoints, zero noisy-neighbor degradation, and sub-100ms TTFT guarantees.
Frequently Asked Questions: Together AI
Developer guidance on API compatibility, rate limit increases, and billing optimization.
Does Together AI support fine-tuning?
Yes, Together AI supports both LoRA and Full Parameter fine-tuning on custom datasets with instant one-click deployment.
What is Together Turbo architecture?
Together Turbo uses optimized FlashAttention kernels and speculative decoding to deliver 2-3x higher throughput compared to stock vLLM.