Back to blog

From provider tokens to Modalis credits: understand cost before sending a request

Modalis AI Editorial Team ·

Reviewed by Modalis AI Lab — Pricing and Model Governance · Verified 2026-08-30

See how input, output, cache, thinking, and tools shape charges, while separating provider cost from Modalis credits, plan access, and commercial floors.

A provider invoices tokens and paid components. A SaaS customer consumes credits under a commercial plan. Those two statements describe connected but different layers. Confusing them leads to bad comparisons, especially when a short visible answer includes a large prompt, hidden reasoning, cached content or a web-search charge.

Two ledgers, one request

The provider ledger records technical usage: uncached input, cached input, cache writes when applicable, output tokens, reasoning or thinking tokens and separately priced tools. The Modalis ledger applies the selected versioned pricing rule, a commercial margin, model floor, plan entitlement, credit reservation and final settlement.

Provider list rates used in this snapshot

Standard token list price per one million tokens — verified August 30, 2026
ModelInputOutputSource
GPT-5.6 LunaUS$0.20US$1.20OpenAI
GPT-5.6 TerraUS$2.00US$12.00OpenAI
GPT-5.6 SolUS$4.00US$20.00OpenAI
Claude Sonnet 5US$2.00US$10.00Anthropic
Claude Opus 5US$5.00US$25.00Anthropic

These rates are an educational snapshot, not a complete provider invoice. Cached input can have a lower rate, cache writes can have another rate, long-context requests may trigger multipliers, batch processing can be discounted and tools such as web search may be billed separately. Provider pricing can change after publication.

A comparable token-only simulation

For a request with 10,000 uncached input tokens and 1,000 output tokens, the approximate provider token cost is shown below. The example excludes cache, thinking details, tools, long-context rules, retries, taxes and Modalis pricing. It exists to illustrate the relative list-rate shape, not to predict a customer charge.

Illustrative raw provider cost: 10k input + 1k output
ModelInput costOutput costApprox. total
GPT-5.6 LunaUS$0.0020US$0.0012US$0.0032
GPT-5.6 TerraUS$0.0200US$0.0120US$0.0320
GPT-5.6 SolUS$0.0400US$0.0200US$0.0600
Claude Sonnet 5US$0.0200US$0.0100US$0.0300
Claude Opus 5US$0.0500US$0.0250US$0.0750

Why a short request can hit the floor

Very small token requests can cost fractions of a cent at the provider. The SaaS still has payment, observability, support, retry, security and platform costs. A per-model floor creates a predictable minimum that preserves the commercial economics of serving the request. When calculated component consumption is higher than the floor, actual usage takes precedence.

Current Modalis entitlement and minimum request charge
ModelMinimum planMinimum creditsWhat it means
GPT-5.6 LunaFree1A successful request consumes at least 1 credit
GPT-5.6 TerraPremium5A successful request consumes at least 5 credits
Claude Sonnet 5Premium6A successful request consumes at least 6 credits
GPT-5.6 SolScale12A successful request consumes at least 12 credits
Claude Opus 5Scale18A successful request consumes at least 18 credits

Minimum plan is an entitlement rule, not a monthly quota statement. The public pricing page is the source of truth for the credits included in a current subscription. A superadmin account can bypass normal entitlement for operations and therefore should not be used to prove what a Free, Premium or Scale customer can see.

Input and output need separate controls

Output is more expensive than input on all five models in this snapshot. The final visible answer can also be only part of output usage when reasoning is enabled. On the input side, conversation history, system instructions, attachment text and tool results all count even if the user typed one short sentence.

  • Limit history deliberately. Keep what the model needs; summarize or remove stale turns.
  • Set realistic maximum output. A huge ceiling makes reservation and worst-case cost unnecessarily high.
  • Treat attachments as input. A PDF or long source file can dominate a request.
  • Measure thinking-enabled scenarios. Visible answer length alone is not a reliable cost proxy.
  • Cap paid tools. Web search and similar components need their own pre-request ceiling.

How to validate a real request

  1. Record the user balance before the request and the selected model.
  2. Send a uniquely identifiable prompt and wait for a completed response.
  3. Confirm input and output usage, provider model, status and provider request ID in the usage dashboard.
  4. Compare the balance delta with actual consumed credits, not with the initial reservation.
  5. Where available, compare provider-console tokens and cost after its reporting delay.
  6. Investigate duplicated records, missing IDs, unexpected refunds or margin below the commercial policy.

A single request is a smoke test, not a pricing certification. More complex validation should cover long history, attachments, cache hits and writes, thinking, maximum output, web search, provider errors, refusals and retries. Each case changes a different component of the price.

Budget forecasts should therefore use a distribution, not one average request. Model a normal case, a high-input case and a high-output or tool-heavy case, then compare each with the available balance. This makes the estimate useful for both product limits and customer communication without pretending that stochastic workloads have one fixed price.

See where each price point fits in the five-model launch guide, and apply the Modalis Lab evaluation method before changing a production default.

Official sources and further reading

Check access before estimating volume

Review the current model catalog and pricing page, then run a small measured test with your real prompt shape.

Review model access