Fireworks AI pricing guide details
Pricing last verified: September 2026. Plans and pricing may change -- check the vendor site for current details.
Pricing Overview
Fireworks AI uses a usage-based, pay-per-token pricing model for serverless inference. There are no fixed monthly subscriptions or seat-based fees -- you pay only for the tokens you consume. New accounts receive $1 in free credits to start experimenting immediately. Pricing scales by model size, with smaller models costing significantly less than larger ones, making Fireworks AI accessible for both prototyping and production workloads.
Beyond serverless inference, Fireworks AI offers on-demand GPU rentals for custom workloads, fine-tuning services for LoRA SFT, batch inference at discounted rates, and image generation endpoints. Cached input tokens receive a 50% discount, rewarding applications that reuse context windows.
Plan Comparison
Fireworks offers usage-based services rather than a traditional subscription-plan table. Serverless Inference uses per-token pricing with high rate limits and postpaid billing; new users receive $1 in free credits. The supplied pricing page directs readers to its documentation for current prices for popular models in the Standard, Priority, and Fast serverless tiers.
| Service | Published pricing basis | Disclosed pricing |
|---|---|---|
| Serverless Inference | Per token | Current popular-model prices are linked in Fireworks documentation; the supplied page does not reproduce that model-price table. |
| Embeddings | Per 1M input tokens | Up to 150M parameters: $0.008; 150M–350M: $0.016; Qwen3 8B: $0.1. |
| On-demand deployments | Per GPU second, shown as hourly rates | H100 80 GB: $8.00/hour; H200 141 GB: $8.00/hour; B200 180 GB: $13.00/hour; B300 288 GB: $15.00/hour; GB300 288 GB: $20.00/hour. |
Region-restricted on-demand deployments in the US or Europe are priced at 1.5× the published rates. Enterprise deployments are available through Fireworks, with faster speeds, lower costs, and higher rate limits described on the supplied pricing page.
Hidden Costs and Considerations
Fine-tuning costs: Managed supervised and preference fine-tuning is priced per 1M training tokens. Rates vary by base-model size and method, with separate published columns for LoRA SFT, LoRA DPO, full-parameter SFT, and full-parameter DPO. Buyers should confirm the base-model size, training method, dataset token count, number of epochs, and any reasoning traces when budgeting.
On-demand GPU rental: On-demand deployments are billed per GPU second, while the pricing page displays hourly rates. From Sep. 1, an H100 80 GB GPU is listed at $8.00 per hour and a B200 180 GB GPU at $13.00 per hour. Region-restricted deployments carry a 1.5× premium, so buyers should confirm the deployment region and expected runtime when estimating dedicated-capacity costs.
Serverless-training usage: The Serverless Training API charges for tokens used to prefill, use cached prefill, sample, and train. The supplied evidence lists separate rates by base model and operation, and says checkpoint storage for serverless models is included during private preview; buyers should confirm the applicable model and token mix.
Initial credits: Fireworks advertises $1 in free credits for getting started. Its supplied pricing evidence describes serverless inference as postpaid and per token, but does not state how long the credit will cover a particular workload or provide a broader free-tier policy.
Cost Estimates by Team Size
Solo developer: Fireworks offers serverless inference with per-token pricing, postpaid billing, and $1 in free credits. The supplied pricing evidence directs readers to Fireworks documentation for current prices for Standard, Priority, and Fast serverless tiers; it does not provide a model-specific token rate for estimating a monthly workload.
Small team (5 engineers): A team evaluating serverless inference should confirm the current per-token price for its selected model and serverless tier before projecting usage. Managed training is priced per 1M training tokens for supervised and preference fine-tuning, with rates varying by model-size band and training method. The supplied evidence does not provide a team-size monthly estimate.
Mid-size team (20 engineers): For dedicated capacity, Fireworks bills on-demand deployments per GPU second and displays hourly rates. The listed rates from Sep. 1 are $8.00 per hour for an H100 80 GB GPU and $13.00 per hour for a B200 180 GB GPU. Buyers should confirm the GPU type, expected runtime, and whether a region-restricted deployment applies; region-restricted deployments carry a 1.5× premium.
How Fireworks AI Pricing Compares
Fireworks AI sits competitively in the serverless LLM inference market. For small models, Groq offers Llama 8B at $0.05/$0.08 per 1M input/output tokens -- roughly half the cost of Fireworks AI's $0.10/1M tier for comparable model sizes. However, Fireworks AI's model selection is broader.
Together AI prices range from $0.10 to $2.50 per 1M tokens depending on model size, closely matching Fireworks AI's tiered structure. Mistral's Small model costs $0.1/$0.3 per 1M input/output tokens, competitive with Fireworks AI's medium tier.
Fireworks AI's strongest value proposition is the 50% cached input discount and 50% batch inference discount, which neither Groq nor Together AI match at the same level. For workloads with high cache hit rates or tolerance for batch latency, effective per-token costs drop to $0.05/1M for small models -- matching Groq's base rate.