59 Models →
Official 2026 Rate Card

Together AI Pricing Sheet & Dedicated Endpoints 2026

Over 100+ open-source foundation models hosted on high-performance cloud clusters. Verified daily against official cloud provider documentation.

Together AI Production Pricing Matrix

Real-time unit costs, context sizes, operational throughput, and optimal use-case recommendations.

Service / Model Tier Primary Rate Secondary / Output Rate Context / Capacity Throughput / Latency Optimal Workload
DeepSeek V3 (671B MoE)$0.60$0.6064,000~70 tok/sCost-effective open-weights MoE
DeepSeek R1 (Full 671B)$1.50$1.5064,000~45 tok/sFull-scale open-weights reasoning
Llama 3.3 70B Instruct Turbo$0.60$0.60128,000~110 tok/sUltra-low-latency 70B inference
Qwen 2.5 72B Instruct$0.60$0.6032,000~90 tok/sMultilingual and mathematical tasks
FLUX.1 Schnell (Image)$0.003 / imageN/AN/A~1.2s / imgHigh-speed 4-step diffusion generation

Architectural & Financial Billing Nuances

Committed Use & Volume Discounts

Enterprise accounts spending over $2,000/mo typically qualify for 20% to 45% volume concessions or reserved throughput capacity agreements.

SLA & Latency Guarantees

Standard tiers offer 99.9% availability. Dedicated instances provide private VPC endpoints, zero noisy-neighbor degradation, and sub-100ms TTFT guarantees.

Frequently Asked Questions: Together AI

Developer guidance on API compatibility, rate limit increases, and billing optimization.

Does Together AI support fine-tuning?

Yes, Together AI supports both LoRA and Full Parameter fine-tuning on custom datasets with instant one-click deployment.

What is Together Turbo architecture?

Together Turbo uses optimized FlashAttention kernels and speculative decoding to deliver 2-3x higher throughput compared to stock vLLM.