Pricing hub

Llama API Pricing

Meta's public documentation does not publish a standard first-party per-token API price in the same way the other providers do, so this page focuses on open-weight deployment tradeoffs.

Live pricing table

Use this table to compare the published input, cached input, and output rates before you run the calculator.

Input 1,000 / Output 250

ModelInput / 1MCached input / 1MOutput / 1MExample requestSource
Llama 4 Maverick

Latest released Llama flagship for broad multimodal workloads.

Meta does not publish a standard first-party per-token API price in its public docs.

No public rateNo public rateNo public rateNo public rateMeta pricing docs
Llama 4 Scout

Open-weight Scout tier for lightweight routing, extraction, and experimentation.

Open-weight model. Meta does not publish a standard first-party per-token API price in its public docs.

No public rateNo public rateNo public rateNo public rateMeta pricing docs
Llama 3.3 70B

Open-weight 70B-class Llama model for self-hosted or third-party-hosted deployments.

Open-weight model. Meta does not publish a standard first-party per-token API price in its public docs.

No public rateNo public rateNo public rateNo public rateMeta pricing docs
Llama 3.1 8B

Open-weight 8B Llama model for low-cost, self-hosted inference and routing.

Open-weight model. Meta does not publish a standard first-party per-token API price in its public docs.

No public rateNo public rateNo public rateNo public rateMeta pricing docs

How to use this pricing page

  1. 1. Start with real prompts. Estimate the input and output token counts from the workflow you actually run.
  2. 2. Check the example cost. The table shows a request-level estimate so you can compare the published rate card with your usage pattern.
  3. 3. Validate the savings. Open the LLM cost calculator and test a few prompt sizes before you make a budget call.
  4. 4. Compare the model choice. If you need a second opinion, use one of the linked comparison pages below.

Highlights

  • • Useful when you want to benchmark open-weight economics.
  • • Helps explain why a model may be absent from a strict token-based comparison.
  • • Pairs well with the calculator and pillar page for deployment planning.

Source: Meta AI docs

Compare this provider

FAQ

Why is there no numeric rate here?

Meta's public docs do not publish a standard first-party per-token API price for these open-weight models, so the pricing table shows the absence of a public rate instead of guessing.

Can I still compare Meta to paid APIs?

Yes, but you should compare the full deployment economics, including hosting, inference, and operational overhead, not just token price.

When does Meta make sense?

Meta is often worth considering when you value open-weight control, deployment flexibility, or self-hosted economics over an immediately published API rate.