AI Dad

Grading Criteria v1.0

Issued 2026-09-10 · Effective 2026-09-10

This document is the standard AI Dad applies when testing an AI tool and assigning a grade. A note like v1.0 §4.2 in a report's "Method" field points to a clause here.

How this came to exist — stated up front

Grading Criteria v1.0 was first written down after reviewing the results and the grading problems in seven records we had already published. This standard was not fixed in advance of those tests. Their existing grades came from case-by-case judgement and from public wording that differs from what is here.

Tests run after v1.0 takes effect have their test plan fixed before execution. The earlier records keep their originals and their original verdicts; re-evaluation under v1.0 is published separately. A requirement we cannot confirm from the evidence of the time is not treated as met.


1. Scope

1.1 This standard evaluates one combination of version, feature, purpose, environment and settings. It does not evaluate a tool as a whole.

1.2 A report's results and grade hold only for the item, version, task and conditions named in it. They do not warrant the same outcome under other conditions or in later runs.

1.3 Every "all" and "zero" verdict in this standard is a result within the samples tested. It does not mean the tool is error-free on all inputs.

1.4 A verdict under this standard is an evaluation against criteria AI Dad has published. It is not an accredited certification and not a warranty of the tool's overall quality or safety.

2. Grade structure

2.1 Grades run from 5 up to 1. Grade 1 is the highest.

2.2 A grade is recognised only when every grade below it has passed. No skipping. If grade 3 was not tested or did not pass, grade 2 is not recognised even where its conditions were met.

2.3 The recognised grade is the highest grade proven in an unbroken run of passes.

2.4 The recognised grade and the point where testing stopped are recorded separately. A note like "grade 2, stopped at grade 3" does not exist under this standard.

2.5 Failure at grade 5 is recorded as "No grade — failed to run." Since grade 5 means it runs, a crash or a failed connection cannot be written as grade 5.

2.6 The point where testing stopped is the lowest grade at which the unbroken run of passes first breaks. If a grade above it was actually tested and did not pass, that verdict is recorded too — §2.2 bars the *recognition* of a higher grade, not the *recording* of what was observed there. There is only ever one stopping point, so a failure above it is marked distinctly from the stop itself.

3. Verdict states

3.1 Each grade takes one of five states.

StateMeaning
MetTested, and the passing condition was met
Not metTested, and it fell short
Not runIt could have been tested; we did not test it
Not applicableThe grade does not arise for this item
Insufficient evidenceA record of testing exists, but no evidence remains to judge on

3.2 "Not run" is never converted into met or not met. On a report, a not-run field is written as "not run" rather than left blank.

3.3 "Not run" and "not applicable" are different things.

Put both in one field and "we were lazy" becomes indistinguishable from "that question does not arise for this tool".

3.4 "Not applicable" does not break the ladder. It is skipped and the climb continues. The report must say why the grade does not arise, and the reason has to be explained by the target type in §4.8. Written without a reason, "not applicable" becomes a way around §2.2 (no skipping).

3.5 A grade marked not applicable is left out of the recognised grade. The report shows what was left out, as in grade 4 (grade 3 not applicable).

3.6 Every report uses the same form, the one built for grade 1. Tests we did not run are not trimmed away — a trimmed report reads as a complete one.

4. Passing conditions by grade

4.0 Common — The evaluator fixes the test plan before execution. Changes made after seeing results do not apply to that test; they are issued as a separate test.

The plan states:

ItemContent
ScopeTarget version, feature, purpose, environment, settings
Required test typesDefect types and normal types, each
Expected result per sampleThe answer key. It must exist before the run
Runs per itemMay differ item to item. §4.7
Excluded stretchesWarm-up and the like. §5.10
Performance thresholdsNamed before the run wherever a performance item exists. §11
Time limit · retry policy
Success signalWhich of status value, exit code or body counts as success

4.0.1 Method is written so it can be repeated. The "Method" field of a report is not prose but these four steps, so that someone else can run the same thing again.

> (1) Prepare — where the test data and the answer key must be > (2) Run — the actual command or operating steps. The filename the result is saved to > (3) Match — on what basis the result is set against the answer key > (4) Read — which value in the output to take

4.1 · Grade 5, Runs — Every run in the test plan must finish inside the stated environment and time limit and return a final response or artefact in the specified form. A single run lost to a crash, timeout, failed connection or an unplanned wait on input fails this grade.

4.2 · Grade 4, Records — Each counted unit of input must correspond to a result record, and the number of requests must equal completed plus failed plus unprocessed. Zero cases of a failure, omission or partial result reported as full success; zero contradictions between body, status value and exit code. Zero cases of a pre-specified gate-defeat test being reported as a normal pass.

4.3 · Grade 3, Detects — Every required defect type in the plan must be tested, and all valid in-scope defect samples must be detected. Misses and downgrades must each be zero. A detection counts only when it reaches the failure or blocking signal named in advance.

4.4 · Grade 2, Precise — Every required normal type in the plan must be tested, with zero false positives on valid normal samples. Every normal control must reach a clean pass with no warning. Grading is done by someone other than the sample author, and at least one independent reviewer beyond the grader cross-checks the answer key, the raw output and the classifications.

4.5 · Grade 1, Proven — For three or more incidents that actually occurred, the occurrence and its conditions must be evidenced, and the inputs and key conditions at the time reconstructed and tested. Each case must show the warning or blocking step acting before the harmful action went through. The responsible party and the action in the real workflow must be named, and the accountable person must sign the result and its scope.

4.6 What grade 1 establishes is "blocking was confirmed on real incident cases." Unless a block was observed in live operation, we do not write "prevented a real incident." Reports distinguish incident reconstruction from blocking observed in operation.

4.7 Runs are set per item. Not everything is run the same number of times. How many times something runs is decided by how much that item drifts and what one run costs, and it is written into the plan before execution. A report shows the count for each item as it was.

A report built on a smaller sample is not evidence of the same strength. 49 of 50 and 1 of 2 do not say the same thing.

4.8 The question each grade asks stays the same; what is measured changes with the item.

The passing conditions in §4.1–4.5 are written for a tool you inject defects into and ask to find them. For some items that premise does not hold. A search tool has no defect to inject and no clean control group.

So the question a grade asks is left alone, and what measures it is set to match the target type.

GradeWhat it asks (every type)
5 · RunsDoes it get to the end without dying
4 · RecordsDoes it show what it failed to do
3 · DetectsDoes it notice what is wrong in what it was given
2 · PreciseDoes it avoid calling a clean thing wrong
1 · ProvenDoes it stop a real incident

4.8.1 Target type — decided by what the tool mainly produces.

TypeWhat comes outWhat grade 3 measures
InspectA verdictDoes it find every defect sample
CollectSomething that was out thereDoes it refuse a target that does not exist or an address that is wrong
GenerateContent that did not existDoes it avoid stating the unsupported as fact
ExecuteThe result of doing somethingDoes it refuse bad arguments and conditions

4.8.2 The type is written into the test plan and shown on the report. Change the type and what gets measured changes, so it is a different test.

4.8.3 Where a tool spans several types, the task set in that test decides which one applies. What is evaluated is that task, not the tool as a whole (§1.1).

4.8.4 A grade that does not arise even under its type is "not applicable" (§3.4). Grade 2 asks whether a clean thing is wrongly called wrong — and a collect-type tool, which issues no verdict at all, has no such thing as a false positive.

🚨 Pick a type and then skip its items, and that is "not run", not "not applicable." The type decides what to measure; it is not a way out.

5. How results are counted

5.1 Defect verdicts and exclusions are counted in six categories only.

CategoryMeaning
DetectedRaised the defect on the designated failure or blocking signal
MissedFailed to find the defect
DowngradedNoted it only as a hint or an aside, never on the exit code or blocking signal
False positiveFlagged a clean sample as defective
Out of scopeExcluded as outside the evaluation scope
Bad inputExcluded because the test material itself was corrupt

5.2 Downgrades are recorded separately in the raw data but counted as misses in the detection rate. Detection rate = detected ÷ (detected + missed + downgraded)

5.3 An out-of-scope exclusion must be evidenced by a scope document fixed before results were seen. The count and the reasons are published.

5.4 Bad input is excluded from the score but its count is always recorded. However, a sample deliberately malformed to test whether the tool rejects it is not bad input. Its expected result is a proper rejection.

5.5 The total number of normal samples and the number that passed cleanly are separate required fields. A normal sample that drew no warning belongs to none of the six categories above.

5.6 Execution failures — crash, timeout, failed connection — are recorded separately as run status, not among the six. A crash is a failure to run, not a miss.

5.7 Where one document holds several defects, detection is counted per defect and the document count is stated separately. Flagging the same defect ten times is not ten detections.

5.8 Results are written as "k of N" and never flattened into a percentage. The unit is the real one measured — runs, questions, keywords, URLs, items.

5.9 Failed runs are not deleted.

5.10 Any stretch excluded from measurement is named before the run. Some tools are slow on the first pass — loading a model, opening a connection, warming a cache. To leave such a stretch out of the measurement, how many runs are excluded must be in the plan before execution.

6. Immediate failure

6.1 Gate defeat — a required check or validation step did not execute, or lost its effect. Reporting a normal pass in that state is grade 4, not met.

6.2 Constant false alarm — where clean input necessarily produces a warning by construction, that is grade 2, not met. A light that is always red teaches people to ignore it, which we treat as no less dangerous than a gate that lets things through.

7. Grader independence

7.1 A grader may not modify the tool or the test material.

7.2 A grader may not grade samples they authored.

7.3 From grade 2 upward, at least one independent reviewer beyond the grader cross-checks the work.

7.4 Grading duty rotates. Anyone with a conflict of interest steps off that case.

7.5 Samples whose answer is fixed outside the tester — Some samples have an answer that is settled by an external source rather than by the tester's judgement: a video ID that does not exist, an address that does not respond, an input that violates the format. For these, the risk §7.2 guards against — reading an ambiguous result the way you intended it — does not arise, even when the person who built the sample scores it. §7.2 is therefore not applied to them, subject to the following.

8. Evidence

8.1 We retain the raw inputs and outputs, logs, run times, target version, environment and any record of human intervention.

8.2 Without evidence a grade is "insufficient evidence." It is never assumed to have been met.

9. Validity and retesting

9.1 A recognised grade is bound to the evaluated combination of code, model, prompt, settings, dependencies and runtime environment. When that combination changes, the grade does not carry over to the changed target.

9.2 Earlier reports are kept as a record of the target as evaluated. A new test is issued under a new number stating that it supersedes the earlier report.

9.3 For a hosted tool whose version cannot be pinned or confirmed, we state the date and whatever identifying information exists, and do not warrant that the current version still behaves the same.

9.4 Where a verdict was wrong even under the criteria applied at the time, we correct or withdraw the report without waiting for a revision of this standard. The original, the reason, the date and the replacing report are published and linked to one another.

10. Records that carry no grade

10.1 Head-to-head records — two tools put to the same task to see how far their results diverge. No grade is assigned, and neither tool is called better. Without an answer key, detections and false positives cannot be judged.

10.2 Reproducibility — how closely repeated runs agree under stated input, environment and timing conditions. The comparison target and the tolerance are recorded with it. A reproducibility result alone never raises the recognised grade, because a tool that returns the same wrong answer every time is reproducible without being correct. Where reproducibility is essential to a particular job, it is set as an additional passing condition in that test plan in advance.

11. Performance items and functional items

11.1 Test items come in two kinds.

KindWhat it asksHow it is judged
FunctionalDoes it work or notAll / zero
PerformanceHow well does it workAgainst a threshold named in advance

11.2 The five rungs of this standard are functional by default. They count as "all defects detected", "zero false positives", "zero contradictions".

11.3 Performance items may be measured alongside. Accuracy, recall, agreement rate, latency — anything that comes out as a number. The threshold must be named in the plan before the run.

🚨 A threshold is never set after seeing the result. "It came out at 98%, so let us call 95% the line" is post-hoc justification, and a number arrived at that way cannot ground a verdict.

11.4 A threshold needs a reason. There is no universal industry pass mark. Write why that line for this job — as in "below this, a person has to look again."

11.5 Clearing a performance threshold does not cover a functional item that was not met. However high the accuracy, a tool that reports a failure as success stops at grade 4. A large number of successes does not offset a defect it failed to report.

11.6 A performance item is reported with measured value, threshold and verdict side by side. A number without its threshold leaves the reader to invent a pass mark of their own.

12. Versioning of this standard

12.1

ChangeVersion
Grade definitions, passing lines, required tests, exclusion rules or independence requirements changeMajor v1 → v2
Explanation or examples added without affecting verdictsMinor v1.0 → v1.1
Typos and linksPatch v1.0.0 → v1.0.1

12.2 Even where a change is labelled clarification, if the same evidence would now pass or fail differently, the major version goes up.

12.3 Each version records its issue date, effective date, reason for change and the differences from the previous version. Earlier documents are not overwritten.

12.4 A v1.0 report does not become void when v2.0 is issued. It stands as a record of what was confirmed under v1.0. It may not, however, be presented as having passed v2.0.

> Grade 4 · Criteria v1.0 · tool version X > Re-evaluation under v2.0: not run


Questions — If you disagree with this standard or believe a verdict is wrong, tell us through [contact](/en/contact). We will check it again against the original evidence and, if we were wrong, publish the correction with its history.