Gate:SelectionLens:Conflicted AssumptionsSeat:CIO & Technical LeadershipType:Buy-Side Decision Audit

Three grades of evidence

Vendor capability claims arrive at one of three evidence grades. Most procurement processes never ask which, and the omission propagates: an ungraded claim enters a business case, the business case reaches a board, and by then nobody remembers what the claim was originally evidence of.

A claim may not be cited above the grade its evidence reached.

Grade information is what gets lost first as a claim travels from a sales conversation to a business case to a board paper.

FIGURE 01: THE CAPABILITY GRADE LADDERGATE 04 · SELECTION
01 · Demonstrated
Core Question: “Can it do this at all?”
Test EnvironmentVendor data · Controlled environment · Vendor operators
Variance StatusExcludes enterprise data quality, internal edge cases, integration friction
Permitted Citation Boundary:

Exploratory evaluation only. Never cited in business cases as production proof.

02 · Piloted
Core Question: “Can it do this here?”
Test EnvironmentBounded trial · Cooperative business unit · Embedded vendor engineers
Variance StatusExcludes unattended volume, system incidents, organizational turnover
Permitted Citation Boundary:

Pilot sign-off only. Evidence that the pilot worked, not that production will hold.

03 · Production-Survived
Core Question: “Does it keep doing this?”
Test EnvironmentUnattended operations · Real peak volume · Model upgrades · Incident response
Variance StatusIncludes full institutional friction and edge-case failure modes
Permitted Citation Boundary:

Board packs, business cases, and binding multi-year contractual commitments.

“The three grades answer different qualitative questions: can it do this at all, can it do this here, and does it keep doing this. A high-confidence Demonstrated claim does not become a low-confidence Piloted one.”

The Framework

The Three Qualitative Grades

GRADE 01

Demonstrated

It worked in a demonstration, on data the vendor selected, under conditions the vendor arranged, with the vendor’s own people operating it.

This is evidence, and dismissing it would be wrong. It establishes that the capability is not fictional — that the system can do the thing at least once, in favourable circumstances. That is genuinely more than nothing.

What it is not is evidence about your institution. A demonstration is designed, and what it is designed to exclude is variance: your data quality, your edge cases, your integration surface, your users doing things nobody anticipated.

GRADE 02

Piloted

It worked in a bounded trial, usually within a cooperative business unit, usually with vendor engineers available, usually over weeks rather than quarters.

A pilot introduces some of the variance a demonstration excludes, which is why it is a higher grade. But pilots are also selected, staffed and supported in ways production is not. The unit that volunteers for a pilot is not a representative unit. The period during which vendor engineers are present is not a representative period.

The most common grading error in enterprise AI is treating a successful pilot as evidence that production will work. It is evidence that the pilot worked.
GRADE 03

Production-survived

It has run unattended, at real volume, through conditions nobody arranged — a peak period, an incident, a model upgrade, staff turnover on both sides — for long enough that the failure modes have had an opportunity to appear.

This is the only grade that answers the question a buyer is actually asking. It is also the grade vendors can least often supply, particularly for capabilities that are genuinely new, which is precisely when institutions most want reassurance.

Analytical Distinction

Why these are not three levels of confidence

The distinction matters because it is qualitative, not quantitative. The three grades answer different questions: can it do this at all, can it do this here, and does it keep doing this. A high-confidence Demonstrated claim does not become a low-confidence Piloted one. It remains an answer to a different question.

The rule that follows is simple and is the practical output of the whole framework: a claim may not be cited above the grade its evidence reached. Not because overstatement is dishonest — usually nobody intends it — but because grade information is what gets lost first as a claim travels from a sales conversation to a business case to a board paper, and the grade is the part that determines what the claim can bear.

Closing Action

Run it once on your strongest claim:

Take the single strongest claim in your current vendor’s material — the headline figure, the flagship reference, the case study everyone repeats. Assign it one of the three grades using only what you have actually been shown, not what you have been told.

Most claims grade Demonstrated. Some grade Piloted. Production-survived claims exist, and vendors who have them will usually let you speak to the reference client — ask to do so without the vendor on the call, and ask specifically what broke.

A Demonstrated grade is not a reason to walk away from a vendor. It is a reason not to write the claim into a business case as though it were Production-survived, and to size the commitment to the evidence rather than to the enthusiasm.
Runnable Instrument

The Measurement Clause Checklist

The full ten-question instrument for auditing capability claims and contract terms.

View Checklist

“The Capability Grade Ladder: a claim may not be cited above the grade its evidence reached.”