Pricing hub
Llama API Pricing
Meta's public documentation does not publish a standard first-party per-token API price in the same way the other providers do, so this page focuses on open-weight deployment tradeoffs.
Live pricing table
Use this table to compare the published input, cached input, and output rates before you run the calculator.
Input 1,000 / Output 250
| Model | Input / 1M | Cached input / 1M | Output / 1M | Example request | Source |
|---|---|---|---|---|---|
Llama 4 Maverick Latest released Llama flagship for broad multimodal workloads. Meta does not publish a standard first-party per-token API price in its public docs. | No public rate | No public rate | No public rate | No public rate | Meta pricing docs |
Llama 4 Scout Open-weight Scout tier for lightweight routing, extraction, and experimentation. Open-weight model. Meta does not publish a standard first-party per-token API price in its public docs. | No public rate | No public rate | No public rate | No public rate | Meta pricing docs |
Llama 3.3 70B Open-weight 70B-class Llama model for self-hosted or third-party-hosted deployments. Open-weight model. Meta does not publish a standard first-party per-token API price in its public docs. | No public rate | No public rate | No public rate | No public rate | Meta pricing docs |
Llama 3.1 8B Open-weight 8B Llama model for low-cost, self-hosted inference and routing. Open-weight model. Meta does not publish a standard first-party per-token API price in its public docs. | No public rate | No public rate | No public rate | No public rate | Meta pricing docs |
How to use this pricing page
- 1. Start with real prompts. Estimate the input and output token counts from the workflow you actually run.
- 2. Check the example cost. The table shows a request-level estimate so you can compare the published rate card with your usage pattern.
- 3. Validate the savings. Open the LLM cost calculator and test a few prompt sizes before you make a budget call.
- 4. Compare the model choice. If you need a second opinion, use one of the linked comparison pages below.
Highlights
- • Useful when you want to benchmark open-weight economics.
- • Helps explain why a model may be absent from a strict token-based comparison.
- • Pairs well with the calculator and pillar page for deployment planning.
Source: Meta AI docs
Compare this provider
FAQ
Why is there no numeric rate here?
Meta's public docs do not publish a standard first-party per-token API price for these open-weight models, so the pricing table shows the absence of a public rate instead of guessing.
Can I still compare Meta to paid APIs?
Yes, but you should compare the full deployment economics, including hosting, inference, and operational overhead, not just token price.
When does Meta make sense?
Meta is often worth considering when you value open-weight control, deployment flexibility, or self-hosted economics over an immediately published API rate.