Back to blog

The Modalis model evaluation method: reproducible prompts, measured cost and honest limits

Modalis AI Editorial Team ·

Reviewed by Modalis AI Lab — Pricing and Model Governance · Verified 2026-08-30

See how Modalis separates provider claims from observations, controls prompts and settings, repeats runs, and publishes evidence limits in model comparisons.

Model comparisons become misleading when one polished answer is treated as scientific proof. Language-model output varies, providers change behavior and a benchmark can reward something unrelated to the customer workflow. The Modalis AI Lab therefore uses evaluation as a product decision process, not as a race to declare one permanent winner.

Five principles before the first prompt

  1. Test a real decision. Every run should inform routing, access, pricing or product communication.
  2. Separate evidence layers. Provider specifications, provider benchmarks and Modalis observations receive different labels.
  3. Control the input. Prompt, history, attachments, tools, effort and output ceiling remain stable inside a comparison.
  4. Measure more than prose. Tokens, cost, latency, tool calls, refusals, retries and reviewer intervention matter.
  5. Publish limitations. A result applies to the tested version, date, language, prompt set and environment.

Step 1: freeze the source snapshot

Before writing or testing, the reviewer records the official model page, pricing page or release note and a verification date. This matters because even primary sources can change. Sonnet 5, for example, launched with an announced temporary price, and Anthropic later made that price permanent. An old announcement remains historically correct but is no longer the best source for a current price.

  • Prefer the current model reference and release notes for specifications and price.
  • Use the launch article for historical positioning and dated benchmark context.
  • Record exact model IDs rather than relying only on a family nickname.
  • Re-verify volatile claims immediately before publication or a commercial change.

Step 2: build a representative prompt suite

A useful suite samples the workload distribution rather than showcasing only the hardest prompt. The initial Modalis matrix includes concise instruction following, multilingual drafting, structured extraction, grounded document analysis, coding, long-context synthesis and tool-oriented tasks. Safety refusals and provider errors are separate outcome classes.

Minimum evaluation dimensions
DimensionExample evidenceWhy it matters
CorrectnessReference answer, tests or expert rubricFluent text can still be wrong
Instruction complianceSchema validation and explicit constraintsProduction output often has a contract
RobustnessRepeated runs and paraphrased inputsOne lucky completion is not a policy
EfficiencyInput/output tokens, latency and retriesAccepted-result cost is more useful than list price
OperationsRequest ID, status, refunds and telemetryA model must work inside the product
SafetyRefusal reason and user-visible handlingA refusal is not the same as an outage

Step 3: control settings and repeat runs

A comparison keeps system prompt, user prompt, history, attachments, available tools and maximum output stable. Reasoning or thinking effort is explicitly recorded. When a model does not support the same parameter, the difference is documented instead of silently substituting a setting.

Non-deterministic work should be repeated. The number of runs depends on risk and cost, but every published conclusion needs enough observations to distinguish a pattern from an anecdote. Results should include failures and refusals; discarding them after the fact inflates the apparent success rate.

Step 4: calculate accepted-result economics

Raw token cost is necessary but incomplete. Suppose a cheaper model needs three attempts and ten minutes of expert repair while a larger model passes once. The larger provider invoice can still produce the lower business cost. Conversely, a flagship model adds no value when the smaller model already passes an automatic validator.

  • Provider cost by component and the pricing version selected for the request.
  • Modalis credits reserved, consumed and refunded.
  • Success rate under the predeclared rubric.
  • Median and tail latency, not only the fastest response.
  • Retries, tool failures and manual interventions required to accept the output.

Step 5: publish a result that can be audited

An editorial comparison states the date, models, prompt family, settings, run count, rubric and known limitations. It links official sources and distinguishes direct measurement from interpretation. If the raw prompt contains customer data, the public report uses a sanitized equivalent and explains that substitution.

Originality is part of the method

Google’s people-first content guidance explicitly rejects writing to a preferred word count and asks whether a page adds original analysis rather than merely rewriting sources. For Modalis, length is a consequence of evidence: a comparison is extensive when the method, data and decision deserve the space.

Translations follow the same rule. Portuguese, English and Spanish pages carry the full argument, tables, limitations and sources. They are reviewed as readable articles in each language, not generated as thin doorway pages that only swap keywords.

A living evaluation, not a permanent verdict

Every result has a shelf life. A model revision, provider price update, prompt change or new Modalis pricing rule can justify a re-run. The article’s verification date tells the reader when its volatile facts were checked; the modification date changes only when the content changes substantially.

This discipline creates a more useful promise than “best AI model.” Modalis can explain which model passed which workflow, under which constraints, at what observed cost—and change that recommendation when the evidence changes.

Start from the five-model launch map and pair the evaluation with the token, credit and plan guide so quality and economics remain connected.

Official sources and further reading

Use the method on your own workflow

Choose a representative prompt, define the pass condition and compare eligible models with the settings held constant.

Open the model catalog