
Together AI
LLM inference cloud infrastructure with 200+ open-source models, token-based pricing, fast inference for production
Overview
Together AI is a cloud LLM inference provider offering a massive catalog of 200+ open-source models (Llama 3.1, Mistral, DeepSeek, Code Llama, Qwen, etc.) and proprietary options via a single OpenAI-compatible API. Pay-as-you-go pricing based solely on tokens consumed: chat/text from $0.03 (Llama-3.1 8B) to $4.50 (Llama-3.1 405B) per million tokens input, output $0.12-$15/M tokens depending on model. Provisioned Throughput (PTUs) offer capacity reservation for high-volume users with volume discounts. Fine-tuning supported ($0.48-$8/M tokens). Code sandbox, on-demand GPUs (H100, B200, GPU clusters rented hourly).
Together AI offers no French-language interface and no French support; the tool is exclusively API/SDK for developers (Python, JavaScript, REST). Rich and active technical documentation for integration. Key distinctive advantage: massive model choice with no peer (200+), prices displayed transparently per token (no hidden 5.5% markup like OpenRouter), zero minimum contracts or long-term commitments. Weakness: high technical learning curve required; needs ML/dev expertise or dedicated technical team to exploit infrastructure fully. Critical TCO point: token pricing varies drastically by model size and type (Llama-3.1 8B at $0.03/M vs 405B at $4.50/M = 150x spread); precise understanding of use case and model selection before engaging is vital to avoid cost drift.
Our verdict
Best for dev/ML teams seeking maximum choice of open-source models with transparent pricing and stable infrastructure. Not for you if you are non-technical or seek generic ChatGPT-like UI: Together is developer infrastructure, not an end-user product.