PRICING EXPLAINED

Know which price you are looking at

There are three different numbers to separate: the model provider's list price, the current LLMFly AI price or billing multiplier, and the amount recorded for a real request. Mixing them produces an inaccurate budget.

Budget formulaInput + output + retriesConfirm against a real usage record
01

Model rates

Filter 24 models across 3 groups. All amounts are USD per 1M tokens.

Open live Model Plaza ↗
Model group

Showing 6 models

openai

ChatGPT(Codex)

70% off · billing multiplier 0.3x

6 / 6 models
Model / tierLLMFly AIOfficial priceSavings vs. official / 1M
codex-auto-review≤272K / >272K · whole request
≤272K
Input$0.06Output$0.36Cache write$0.075Cache read$0.006
>272K
Input$0.12Output$0.54Cache write$0.15Cache read$0.012
≤272K
Input$0.2Output$1.2Cache write$0.25Cache read$0.02
>272K
Input$0.4Output$1.8Cache write$0.5Cache read$0.04
70% OFF
≤272KIn save $0.14Out save $0.84Write save $0.175Read save $0.014
>272KIn save $0.28Out save $1.26Write save $0.35Read save $0.028
gpt-5.4Standard
Input$0.75Output$4.5Cache writeCache read$0.075
Input$2.5Output$15Cache writeCache read$0.25
70% OFF
In save $1.75Out save $10.5Read save $0.175
gpt-5.4-miniStandard
Input$0.225Output$1.35Cache writeCache read$0.0225
Input$0.75Output$4.5Cache writeCache read$0.075
70% OFF
In save $0.525Out save $3.15Read save $0.0525
gpt-5.5Standard
Input$1.5Output$9Cache writeCache read$0.15
Input$5Output$30Cache writeCache read$0.5
70% OFF
In save $3.5Out save $21Read save $0.35
gpt-5.6-sol≤272K / >272K · whole request
≤272K
Input$1.5Output$9Cache write$1.875Cache read$0.15
>272K
Input$3Output$13.5Cache write$3.75Cache read$0.3
≤272K
Input$5Output$30Cache write$6.25Cache read$0.5
>272K
Input$10Output$45Cache write$12.5Cache read$1
70% OFF
≤272KIn save $3.5Out save $21Write save $4.375Read save $0.35
>272KIn save $7Out save $31.5Write save $8.75Read save $0.7
gpt-5.6-terra≤272K / >272K · whole request
≤272K
Input$0.6Output$3.6Cache write$0.75Cache read$0.06
>272K
Input$1.2Output$5.4Cache write$1.5Cache read$0.12
≤272K
Input$2Output$12Cache write$2.5Cache read$0.2
>272K
Input$4Output$18Cache write$5Cache read$0.4
70% OFF
≤272KIn save $1.4Out save $8.4Write save $1.75Read save $0.14
>272KIn save $2.8Out save $12.6Write save $3.5Read save $0.28
02

Do not stop at token price

The number that matters when buying credits

For LLMFly AI credits, use the current price shown in Model Plaza—not the provider reference table above. Then send a small request and check the usage record before projecting monthly spend.

A usable monthly estimate

request.exampleCopy-ready
monthly_cost =
  (input_tokens / 1_000_000 * input_rate) +
  (output_tokens / 1_000_000 * output_rate) +
  cache_and_tool_charges + retry_cost

Costs teams often miss

  • Conversation history resent on every turn.
  • Reasoning tokens counted as output.
  • AI agent tool loops and validation calls.
  • Retries after 429 or transient upstream failures.
  • Cache writes, cache reads, and storage when the model supports them.

Frequently asked questions

Do provider list prices equal LLMFly AI charges?

No. They are reference figures. Use Model Plaza and the usage record for actual charges.

Why estimate with complete tasks?

A cheaper token can still produce a more expensive result when it needs more retries, longer context, or manual correction.

Where do I find the current model price?

Open Model Plaza and check the models available to your API key.

Check the current model price

Use Model Plaza for the platform price, then verify one request in the usage log.

Open Model Plaza