The Modalis model evaluation method: reproducible prompts, measured cost and honest limits
Modalis AI Editorial Team ·
Reviewed by Modalis AI Lab — Pricing and Model Governance · Verified 2026-08-30
See how Modalis separates provider claims from observations, controls prompts and settings, repeats runs, and publishes evidence limits in model comparisons.
Model comparisons become misleading when one polished answer is treated as scientific proof. Language-model output varies, providers change behavior and a benchmark can reward something unrelated to the customer workflow. The Modalis AI Lab therefore uses evaluation as a product decision process, not as a race to declare one permanent winner.
Five principles before the first prompt
- Test a real decision. Every run should inform routing, access, pricing or product communication.
- Separate evidence layers. Provider specifications, provider benchmarks and Modalis observations receive different labels.
- Control the input. Prompt, history, attachments, tools, effort and output ceiling remain stable inside a comparison.
- Measure more than prose. Tokens, cost, latency, tool calls, refusals, retries and reviewer intervention matter.
- Publish limitations. A result applies to the tested version, date, language, prompt set and environment.
Step 1: freeze the source snapshot
Before writing or testing, the reviewer records the official model page, pricing page or release note and a verification date. This matters because even primary sources can change. Sonnet 5, for example, launched with an announced temporary price, and Anthropic later made that price permanent. An old announcement remains historically correct but is no longer the best source for a current price.
- Prefer the current model reference and release notes for specifications and price.
- Use the launch article for historical positioning and dated benchmark context.
- Record exact model IDs rather than relying only on a family nickname.
- Re-verify volatile claims immediately before publication or a commercial change.
Step 2: build a representative prompt suite
A useful suite samples the workload distribution rather than showcasing only the hardest prompt. The initial Modalis matrix includes concise instruction following, multilingual drafting, structured extraction, grounded document analysis, coding, long-context synthesis and tool-oriented tasks. Safety refusals and provider errors are separate outcome classes.
| Dimension | Example evidence | Why it matters |
|---|---|---|
| Correctness | Reference answer, tests or expert rubric | Fluent text can still be wrong |
| Instruction compliance | Schema validation and explicit constraints | Production output often has a contract |
| Robustness | Repeated runs and paraphrased inputs | One lucky completion is not a policy |
| Efficiency | Input/output tokens, latency and retries | Accepted-result cost is more useful than list price |
| Operations | Request ID, status, refunds and telemetry | A model must work inside the product |
| Safety | Refusal reason and user-visible handling | A refusal is not the same as an outage |
Step 3: control settings and repeat runs
A comparison keeps system prompt, user prompt, history, attachments, available tools and maximum output stable. Reasoning or thinking effort is explicitly recorded. When a model does not support the same parameter, the difference is documented instead of silently substituting a setting.
Non-deterministic work should be repeated. The number of runs depends on risk and cost, but every published conclusion needs enough observations to distinguish a pattern from an anecdote. Results should include failures and refusals; discarding them after the fact inflates the apparent success rate.
Step 4: calculate accepted-result economics
Raw token cost is necessary but incomplete. Suppose a cheaper model needs three attempts and ten minutes of expert repair while a larger model passes once. The larger provider invoice can still produce the lower business cost. Conversely, a flagship model adds no value when the smaller model already passes an automatic validator.
- Provider cost by component and the pricing version selected for the request.
- Modalis credits reserved, consumed and refunded.
- Success rate under the predeclared rubric.
- Median and tail latency, not only the fastest response.
- Retries, tool failures and manual interventions required to accept the output.
Step 5: publish a result that can be audited
An editorial comparison states the date, models, prompt family, settings, run count, rubric and known limitations. It links official sources and distinguishes direct measurement from interpretation. If the raw prompt contains customer data, the public report uses a sanitized equivalent and explains that substitution.
Originality is part of the method
Google’s people-first content guidance explicitly rejects writing to a preferred word count and asks whether a page adds original analysis rather than merely rewriting sources. For Modalis, length is a consequence of evidence: a comparison is extensive when the method, data and decision deserve the space.
Translations follow the same rule. Portuguese, English and Spanish pages carry the full argument, tables, limitations and sources. They are reviewed as readable articles in each language, not generated as thin doorway pages that only swap keywords.
A living evaluation, not a permanent verdict
Every result has a shelf life. A model revision, provider price update, prompt change or new Modalis pricing rule can justify a re-run. The article’s verification date tells the reader when its volatile facts were checked; the modification date changes only when the content changes substantially.
This discipline creates a more useful promise than “best AI model.” Modalis can explain which model passed which workflow, under which constraints, at what observed cost—and change that recommendation when the evidence changes.
Start from the five-model launch map and pair the evaluation with the token, credit and plan guide so quality and economics remain connected.
Official sources and further reading
- OpenAI API model catalog — OpenAI
- Using GPT-5.6 — OpenAI
- What is new in Claude Sonnet 5 — Anthropic
- Introducing Claude Opus 5 — Anthropic
- Claude Platform release notes — Anthropic
- Creating helpful, reliable, people-first content — Google Search Central
Use the method on your own workflow
Choose a representative prompt, define the pass condition and compare eligible models with the settings held constant.