Grading Criteria v1.0
Issued 2026-09-10 · Effective 2026-09-10
This document is the standard AI Dad applies when testing an AI tool and assigning a grade. A note like v1.0 §4.2 in a report's "Method" field points to a clause here.
How this came to exist — stated up front
Grading Criteria v1.0 was first written down after reviewing the results and the grading problems in seven records we had already published. This standard was not fixed in advance of those tests. Their existing grades came from case-by-case judgement and from public wording that differs from what is here.
Tests run after v1.0 takes effect have their test plan fixed before execution. The earlier records keep their originals and their original verdicts; re-evaluation under v1.0 is published separately. A requirement we cannot confirm from the evidence of the time is not treated as met.
1. Scope
1.1 This standard evaluates one combination of version, feature, purpose, environment and settings. It does not evaluate a tool as a whole.
1.2 A report's results and grade hold only for the item, version, task and conditions named in it. They do not warrant the same outcome under other conditions or in later runs.
1.3 Every "all" and "zero" verdict in this standard is a result within the samples tested. It does not mean the tool is error-free on all inputs.
1.4 A verdict under this standard is an evaluation against criteria AI Dad has published. It is not an accredited certification and not a warranty of the tool's overall quality or safety.
2. Grade structure
2.1 Grades run from 5 up to 1. Grade 1 is the highest.
2.2 A grade is recognised only when every grade below it has passed. No skipping. If grade 3 was not tested or did not pass, grade 2 is not recognised even where its conditions were met.
2.3 The recognised grade is the highest grade proven in an unbroken run of passes.
2.4 The recognised grade and the point where testing stopped are recorded separately. A note like "grade 2, stopped at grade 3" does not exist under this standard.
2.5 Failure at grade 5 is recorded as "No grade — failed to run." Since grade 5 means it runs, a crash or a failed connection cannot be written as grade 5.
2.6 The point where testing stopped is the lowest grade at which the unbroken run of passes first breaks. If a grade above it was actually tested and did not pass, that verdict is recorded too — §2.2 bars the *recognition* of a higher grade, not the *recording* of what was observed there. There is only ever one stopping point, so a failure above it is marked distinctly from the stop itself.
3. Verdict states
3.1 Each grade takes one of five states.
| State | Meaning |
|---|---|
| Met | Tested, and the passing condition was met |
| Not met | Tested, and it fell short |
| Not run | It could have been tested; we did not test it |
| Not applicable | The grade does not arise for this item |
| Insufficient evidence | A record of testing exists, but no evidence remains to judge on |
3.2 "Not run" is never converted into met or not met. On a report, a not-run field is written as "not run" rather than left blank.
3.3 "Not run" and "not applicable" are different things.
- Not run — it could be measured and we did not measure it. Test it later and the field fills in.
- Not applicable — there is nothing there to measure. A search tool has no defect samples to inject and no clean control group.
Put both in one field and "we were lazy" becomes indistinguishable from "that question does not arise for this tool".
3.4 "Not applicable" does not break the ladder. It is skipped and the climb continues. The report must say why the grade does not arise, and the reason has to be explained by the target type in §4.8. Written without a reason, "not applicable" becomes a way around §2.2 (no skipping).
3.5 A grade marked not applicable is left out of the recognised grade. The report shows what was left out, as in grade 4 (grade 3 not applicable).
3.6 Every report uses the same form, the one built for grade 1. Tests we did not run are not trimmed away — a trimmed report reads as a complete one.
4. Passing conditions by grade
4.0 Common — The evaluator fixes the test plan before execution. Changes made after seeing results do not apply to that test; they are issued as a separate test.
The plan states:
| Item | Content |
|---|---|
| Scope | Target version, feature, purpose, environment, settings |
| Required test types | Defect types and normal types, each |
| Expected result per sample | The answer key. It must exist before the run |
| Runs per item | May differ item to item. §4.7 |
| Excluded stretches | Warm-up and the like. §5.10 |
| Performance thresholds | Named before the run wherever a performance item exists. §11 |
| Time limit · retry policy | |
| Success signal | Which of status value, exit code or body counts as success |
4.0.1 Method is written so it can be repeated. The "Method" field of a report is not prose but these four steps, so that someone else can run the same thing again.
> (1) Prepare — where the test data and the answer key must be > (2) Run — the actual command or operating steps. The filename the result is saved to > (3) Match — on what basis the result is set against the answer key > (4) Read — which value in the output to take
4.1 · Grade 5, Runs — Every run in the test plan must finish inside the stated environment and time limit and return a final response or artefact in the specified form. A single run lost to a crash, timeout, failed connection or an unplanned wait on input fails this grade.
4.2 · Grade 4, Records — Each counted unit of input must correspond to a result record, and the number of requests must equal completed plus failed plus unprocessed. Zero cases of a failure, omission or partial result reported as full success; zero contradictions between body, status value and exit code. Zero cases of a pre-specified gate-defeat test being reported as a normal pass.
4.3 · Grade 3, Detects — Every required defect type in the plan must be tested, and all valid in-scope defect samples must be detected. Misses and downgrades must each be zero. A detection counts only when it reaches the failure or blocking signal named in advance.
4.4 · Grade 2, Precise — Every required normal type in the plan must be tested, with zero false positives on valid normal samples. Every normal control must reach a clean pass with no warning. Grading is done by someone other than the sample author, and at least one independent reviewer beyond the grader cross-checks the answer key, the raw output and the classifications.
4.5 · Grade 1, Proven — For three or more incidents that actually occurred, the occurrence and its conditions must be evidenced, and the inputs and key conditions at the time reconstructed and tested. Each case must show the warning or blocking step acting before the harmful action went through. The responsible party and the action in the real workflow must be named, and the accountable person must sign the result and its scope.
4.6 What grade 1 establishes is "blocking was confirmed on real incident cases." Unless a block was observed in live operation, we do not write "prevented a real incident." Reports distinguish incident reconstruction from blocking observed in operation.
4.7 Runs are set per item. Not everything is run the same number of times. How many times something runs is decided by how much that item drifts and what one run costs, and it is written into the plan before execution. A report shows the count for each item as it was.
A report built on a smaller sample is not evidence of the same strength. 49 of 50 and 1 of 2 do not say the same thing.
4.8 The question each grade asks stays the same; what is measured changes with the item.
The passing conditions in §4.1–4.5 are written for a tool you inject defects into and ask to find them. For some items that premise does not hold. A search tool has no defect to inject and no clean control group.
So the question a grade asks is left alone, and what measures it is set to match the target type.
| Grade | What it asks (every type) |
|---|---|
| 5 · Runs | Does it get to the end without dying |
| 4 · Records | Does it show what it failed to do |
| 3 · Detects | Does it notice what is wrong in what it was given |
| 2 · Precise | Does it avoid calling a clean thing wrong |
| 1 · Proven | Does it stop a real incident |
4.8.1 Target type — decided by what the tool mainly produces.
| Type | What comes out | What grade 3 measures |
|---|---|---|
| Inspect | A verdict | Does it find every defect sample |
| Collect | Something that was out there | Does it refuse a target that does not exist or an address that is wrong |
| Generate | Content that did not exist | Does it avoid stating the unsupported as fact |
| Execute | The result of doing something | Does it refuse bad arguments and conditions |
4.8.2 The type is written into the test plan and shown on the report. Change the type and what gets measured changes, so it is a different test.
4.8.3 Where a tool spans several types, the task set in that test decides which one applies. What is evaluated is that task, not the tool as a whole (§1.1).
4.8.4 A grade that does not arise even under its type is "not applicable" (§3.4). Grade 2 asks whether a clean thing is wrongly called wrong — and a collect-type tool, which issues no verdict at all, has no such thing as a false positive.
🚨 Pick a type and then skip its items, and that is "not run", not "not applicable." The type decides what to measure; it is not a way out.
5. How results are counted
5.1 Defect verdicts and exclusions are counted in six categories only.
| Category | Meaning |
|---|---|
| Detected | Raised the defect on the designated failure or blocking signal |
| Missed | Failed to find the defect |
| Downgraded | Noted it only as a hint or an aside, never on the exit code or blocking signal |
| False positive | Flagged a clean sample as defective |
| Out of scope | Excluded as outside the evaluation scope |
| Bad input | Excluded because the test material itself was corrupt |
5.2 Downgrades are recorded separately in the raw data but counted as misses in the detection rate. Detection rate = detected ÷ (detected + missed + downgraded)
5.3 An out-of-scope exclusion must be evidenced by a scope document fixed before results were seen. The count and the reasons are published.
5.4 Bad input is excluded from the score but its count is always recorded. However, a sample deliberately malformed to test whether the tool rejects it is not bad input. Its expected result is a proper rejection.
5.5 The total number of normal samples and the number that passed cleanly are separate required fields. A normal sample that drew no warning belongs to none of the six categories above.
5.6 Execution failures — crash, timeout, failed connection — are recorded separately as run status, not among the six. A crash is a failure to run, not a miss.
5.7 Where one document holds several defects, detection is counted per defect and the document count is stated separately. Flagging the same defect ten times is not ten detections.
5.8 Results are written as "k of N" and never flattened into a percentage. The unit is the real one measured — runs, questions, keywords, URLs, items.
5.9 Failed runs are not deleted.
5.10 Any stretch excluded from measurement is named before the run. Some tools are slow on the first pass — loading a model, opening a connection, warming a cache. To leave such a stretch out of the measurement, how many runs are excluded must be in the plan before execution.
- The count and the reason are recorded on the report. Nothing is dropped quietly.
- We do not decide after the fact that a run "was warm-up". That is post-hoc justification.
- A crash inside an excluded stretch cannot be excluded. Grade 5 is judged across every run, warm-up included (§4.1).
6. Immediate failure
6.1 Gate defeat — a required check or validation step did not execute, or lost its effect. Reporting a normal pass in that state is grade 4, not met.
6.2 Constant false alarm — where clean input necessarily produces a warning by construction, that is grade 2, not met. A light that is always red teaches people to ignore it, which we treat as no less dangerous than a gate that lets things through.
7. Grader independence
7.1 A grader may not modify the tool or the test material.
7.2 A grader may not grade samples they authored.
7.3 From grade 2 upward, at least one independent reviewer beyond the grader cross-checks the work.
7.4 Grading duty rotates. Anyone with a conflict of interest steps off that case.
7.5 Samples whose answer is fixed outside the tester — Some samples have an answer that is settled by an external source rather than by the tester's judgement: a video ID that does not exist, an address that does not respond, an input that violates the format. For these, the risk §7.2 guards against — reading an ambiguous result the way you intended it — does not arise, even when the person who built the sample scores it. §7.2 is therefore not applied to them, subject to the following.
- The source that fixed the answer is written on the report.
- The rule used to select the samples is written into the test plan before the run (§4.0). Selection bias is not resolved by this exception.
- A sample whose answer depends on the tester's interpretation does not qualify — anything turning on quality, appropriateness, or domain judgement.
8. Evidence
8.1 We retain the raw inputs and outputs, logs, run times, target version, environment and any record of human intervention.
8.2 Without evidence a grade is "insufficient evidence." It is never assumed to have been met.
9. Validity and retesting
9.1 A recognised grade is bound to the evaluated combination of code, model, prompt, settings, dependencies and runtime environment. When that combination changes, the grade does not carry over to the changed target.
9.2 Earlier reports are kept as a record of the target as evaluated. A new test is issued under a new number stating that it supersedes the earlier report.
9.3 For a hosted tool whose version cannot be pinned or confirmed, we state the date and whatever identifying information exists, and do not warrant that the current version still behaves the same.
9.4 Where a verdict was wrong even under the criteria applied at the time, we correct or withdraw the report without waiting for a revision of this standard. The original, the reason, the date and the replacing report are published and linked to one another.
10. Records that carry no grade
10.1 Head-to-head records — two tools put to the same task to see how far their results diverge. No grade is assigned, and neither tool is called better. Without an answer key, detections and false positives cannot be judged.
10.2 Reproducibility — how closely repeated runs agree under stated input, environment and timing conditions. The comparison target and the tolerance are recorded with it. A reproducibility result alone never raises the recognised grade, because a tool that returns the same wrong answer every time is reproducible without being correct. Where reproducibility is essential to a particular job, it is set as an additional passing condition in that test plan in advance.
11. Performance items and functional items
11.1 Test items come in two kinds.
| Kind | What it asks | How it is judged |
|---|---|---|
| Functional | Does it work or not | All / zero |
| Performance | How well does it work | Against a threshold named in advance |
11.2 The five rungs of this standard are functional by default. They count as "all defects detected", "zero false positives", "zero contradictions".
11.3 Performance items may be measured alongside. Accuracy, recall, agreement rate, latency — anything that comes out as a number. The threshold must be named in the plan before the run.
🚨 A threshold is never set after seeing the result. "It came out at 98%, so let us call 95% the line" is post-hoc justification, and a number arrived at that way cannot ground a verdict.
11.4 A threshold needs a reason. There is no universal industry pass mark. Write why that line for this job — as in "below this, a person has to look again."
11.5 Clearing a performance threshold does not cover a functional item that was not met. However high the accuracy, a tool that reports a failure as success stops at grade 4. A large number of successes does not offset a defect it failed to report.
11.6 A performance item is reported with measured value, threshold and verdict side by side. A number without its threshold leaves the reader to invent a pass mark of their own.
12. Versioning of this standard
12.1
| Change | Version |
|---|---|
| Grade definitions, passing lines, required tests, exclusion rules or independence requirements change | Major v1 → v2 |
| Explanation or examples added without affecting verdicts | Minor v1.0 → v1.1 |
| Typos and links | Patch v1.0.0 → v1.0.1 |
12.2 Even where a change is labelled clarification, if the same evidence would now pass or fail differently, the major version goes up.
12.3 Each version records its issue date, effective date, reason for change and the differences from the previous version. Earlier documents are not overwritten.
12.4 A v1.0 report does not become void when v2.0 is issued. It stands as a record of what was confirmed under v1.0. It may not, however, be presented as having passed v2.0.
> Grade 4 · Criteria v1.0 · tool version X > Re-evaluation under v2.0: not run
Questions — If you disagree with this standard or believe a verdict is wrong, tell us through [contact](/en/contact). We will check it again against the original evidence and, if we were wrong, publish the correction with its history.