Together AI pricing guide details
Pricing Overview
Together AI publishes usage pricing across several products. For serverless inference, its pricing table lists prices per 1M tokens that vary by model and by input versus output tokens. For example, Ternary Bonsai 27B is listed at $0.00 for both input and output, while Kimi K3 is listed at $3.00 input and $15.00 output; some models also have separately listed cached-input prices.
The pricing page also lists GPU capacity on an hourly, per-GPU basis. GPU Cluster on-demand rates shown include $3.99 for NVIDIA HGX H100, $5.99 for NVIDIA HGX H200, and $8.19 for NVIDIA HGX B200. Reserved rates are shown for selected terms, while some hardware and longer-term options are marked “Contact us.”
Other disclosed usage prices include Code Sandbox at $0.0446 per vCPU-hour and $0.0149 per GiB RAM-hour, Code Interpreter at $0.03 per 60-minute session, and Shared Filesystem storage at $0.16 per GiB/month. Fine-tuning pricing is based on tokens processed, with a stated minimum charge of $4.00 per job for the standard pricing table.
Plan Comparison
Together AI’s pricing page is organized by service type and model or hardware rather than by traditional subscription tiers. Serverless inference is listed with model-specific usage rates, while capacity-oriented services use throughput-unit, GPU-hour, or resource-based pricing.
| Service | Official pricing basis | Buying consideration |
|---|---|---|
| Serverless Inference | Price per 1M tokens, with separate input and output rates for listed chat models | Confirm the exact model and whether cached-input pricing applies. |
| Provisioned Throughput | Throughput units (PTUs); the page’s estimate assumes continuous 24/7 provisioning | Confirm the model, traffic profile, required PTUs, and deployment duration. |
| Dedicated Inference | Per GPU per hour for published on-demand hardware rates | Some hardware and reserved-capacity entries are listed as “Contact us” or “Contact sales”; confirm availability and the applicable reservation terms. |
| GPU Clusters | Per GPU per hour, with published on-demand and selected reserved-duration rates | Confirm the hardware and reservation duration; some entries do not show a public rate. |
| Fine-Tuning | Per 1M tokens processed, with pricing varying by training method and model size | Confirm the training method, model size, token volume, and any minimum charge. |
For serverless use, the official Chat table includes LFM2.5-8B-A1B at $0.03 per 1M input tokens and $0.12 per 1M output tokens. Other listed models vary materially, so the table does not establish a single universal serverless rate. The pricing page also explains that fine-tuning charges are based on tokens processed in the training dataset plus applicable evaluation-dataset tokens, and that standard fine-tuning jobs have a $4.00 minimum charge. Buyers should price the specific service, model, and capacity commitment they intend to use rather than extrapolating from one model’s token rate.
Hidden Costs and Considerations
While Together AI's per-token and per-hour rates are transparent, several factors can impact your actual bill. Larger context windows consume more tokens per request, driving up costs on long-document tasks. Dedicated endpoints bill by the hour regardless of utilization, so underused GPUs become expensive idle capacity. Fine-tuning costs compound with dataset size and the number of training epochs -- a multi-pass training run on a large dataset will multiply the base $3/M token rate accordingly. There are no egress fees or platform surcharges listed, but teams should monitor token consumption closely to avoid budget surprises.
Cost Estimates by Team Size
The following estimates assume serverless inference on a mid-range model at approximately $0.50 per million tokens, which is representative of popular open-source models in the 7B-13B parameter range.
| Team Size | Estimated Monthly Usage | Estimated Monthly Cost |
|---|---|---|
| Solo developer / Prototype | ~5M tokens | $2.50 (covered by free credits initially) |
| Small team (3-5 developers) | ~50M tokens | $25 |
| Mid-size team (10-20 developers) | ~500M tokens | $250 |
| Production workload (dedicated A100) | 1 GPU, 24/7 | ~$576/month ($0.80/hr x 720 hrs) |
| Enterprise (multi-GPU cluster) | 4 GPUs, 24/7 | ~$2,304/month |
These figures scale linearly with token volume. Teams running sizable models at $2.50/M tokens should multiply the serverless estimates by 5x. Adding fine-tuning to the mix introduces a one-time training cost that varies with dataset size.
How Together AI Pricing Compares
Together AI’s official pricing page presents serverless inference as a usage-based service with model-specific rates rather than a single published serverless starting price. The Chat table lists separate input and output prices per 1M tokens; some models also show a lower cached-input rate.
| Example model | Input price per 1M tokens | Output price per 1M tokens |
|---|---|---|
| LFM2.5-8B-A1B | $0.03 | $0.12 |
| gpt-oss-20B | $0.05 | $0.20 |
| Qwen3.7-Max | $1.25 | $3.75 |
| Kimi K3 | $3.00 | $15.00 |
The same page also lists pricing for provisioned throughput, dedicated inference, GPU clusters, sandbox resources, storage, and fine-tuning. This means a useful comparison should match the specific Together AI service and model to the workload rather than treating serverless inference as a single flat rate. Buyers should confirm the selected model’s input, output, and any cached-input rates, as well as whether their workload needs serverless inference or reserved capacity.