Testing Standard v1.0
Adopted 2026-09-10 · In force 2026-09-10
This document sets the conditions under which we test an AI tool submitted to us: what we take, what we do not, what you supply, and what comes back.
How results are judged is not here — that is [Grading Criteria v1.0](/en/criteria). This standard governs preparation and procedure; the criteria govern the verdict. Where the two conflict, we do not pick whichever suits us; we stop the test and fix the document.
Stated up front
AI Dad is not an accredited testing body. The "field-test report" issued here is neither an accredited test certificate nor a certification — it is a private evaluation record made against a published standard of our own. Results hold only for the combination, samples, environment and date written on the report.
1. What we can take right now
1.1 The ceiling is grade 3. We do not take grade 2 or grade 1.
| Grade | Accepted | Why |
|---|---|---|
| 5 · Runs | Yes | The verdict is mechanical — did it get to the end |
| 4 · Records | Yes | Comparing body, status and exit code settles it |
| 3 · Detects | Conditionally | Limited to the defects in §2.3 |
| 2 · Precise | No | Criteria §4.4 requires at least one independent reviewer |
| 1 · Proven | No | Criteria §4.5 requires three documented real incidents and a sign-off |
1.2 Grades 2 and 1 are closed for want of people, not technique. AI Dad is run by one person. Reading your own work again is not independent review, and having another AI model look it over does not secure a reviewer either. Writing it down and then doing it alone would be a lie.
1.3 Grade 2 opens when an independent reviewer is actually in place. Until then we do not use the phrase "grade 2 test."
1.4 The ceiling is low, but below it nothing is skipped. Apply for grade 3 and grades 5 and 4 are tested too, because a lower grade cannot be skipped (criteria §2.2).
2. We build the samples
2.1 Samples supplied by the client are not used to decide a grade. They can be drawn from the cases that work, and even with no such intent the blind spots of whoever built them become the boundary of the test.
2.2 What you supply is used to understand the tool and confirm it runs. Where a run on your material appears on the report, its origin is stated — it is never written up as though it came from our own samples.
2.3 The trade-off is that grade 3 covers fewer defects. Since we both build and score the samples, criteria §7.5 limits us to defects whose answer is fixed outside the tester.
| Accepted | Not accepted |
|---|---|
| Targets that do not exist (dead video IDs, deleted documents) | "Is this answer appropriate?" |
| Addresses that do not respond, broken connections | "Does this summary represent the source?" |
| Inputs that violate a format or schema | "Does this call match industry practice?" |
| Unauthorised requests, out-of-range arguments | Anything where domain experts would disagree |
The right-hand column puts our interpretation into the answer. Building the sample, setting the answer and scoring it would not be a test; it would be talking to ourselves.
2.4 The rule used to select samples is written before the run and published on the report. The §2.3 exception removes bias in the answer, not bias in the selection.
2.5 A sample found to be invalid is not quietly dropped. The reason, when it was found, and its effect on the result are recorded — so that a poor score cannot be improved by swapping samples.
3. What is tested
3.1 What gets evaluated is not the tool but one combination of code, model, prompt, settings, dependencies and runtime (criteria §1.1, §9.1).
3.2 The target type is fixed at intake — inspect, collect, generate or execute. Change the type and what is measured changes, making it a different test (criteria §4.8).
3.3 Where a tool spans several types, the task set for this test decides which one applies.
3.4 Hosted tools whose version cannot be pinned are accepted. The report then carries only the observation date and whatever identifying information exists, and does not warrant that the same result holds now (criteria §9.3). An observation date is not treated as a pinned version.
4. What you supply
Without these the test does not stand up.
| Required | |
|---|---|
| ① | A way to run it — package, container or executable, or a service URL and a test account |
| ② | The combination — version and build identifiers, model identifier, prompts and settings in force, connected tools, dependencies, runtime |
| ③ | Operating instructions — how to install and call it, how to confirm it is working, how to reset it, what resources it needs |
| ④ | Scope of application — the ceiling grade applied for, and what the tool claims it can do |
| ⑤ | Input and output specification — what it accepts, what it returns, what is in scope and what is out |
| ⑥ | Declared constraints — known limits, forbidden operations, outbound traffic, paid calls, usage caps |
| ⑦ | Confirmation of rights — the right to test this tool, to retain evidence, and to publish results within the agreed scope |
4.1 We do not ask for full source or model weights. Identifying the target, pinning the conditions and keeping evidence is enough.
4.2 We do not ask for your final samples or answer key (§2.1). Format examples are welcome.
4.3 Items that do not apply to your type are written as "not applicable" with a reason. They are not left blank.
5. Procedure
5.1 The test plan is fixed before the run (criteria §4.0). Items, sample composition, repeat counts and any excluded warm-up are settled before any result is seen.
5.2 The method is written in four steps, so that someone else can run it again (criteria §4.0.1).
1. Prepare — what was placed in what environment 2. Run — what was asked, how many times, where the output was saved 3. Match — what it was compared against 4. Read — what was looked at to reach the verdict
5.3 The target cannot change mid-test. If it does, the test ends there and the changed target is taken as a new one.
5.4 Repeat counts are set per item (criteria §4.7), from how much the item moves and what one run costs, and appear on the report item by item.
5.5 Failed runs are not deleted (criteria §5.9).
6. Not testable
6.1 "Not testable" is an intake state, not a result. It is not a sixth member of the five verdict states in criteria §3.1.
6.2 After a chance to complete the submission, any of the following ends it.
- Function, type or scope cannot be pinned down, so no test items can be selected
- No way to run it, or access is not granted
- The combination cannot be identified or held fixed, or changes are demanded mid-test
- No valid defect samples can be built within §2.3 (grade 3 applications only)
- No right to use the material for testing, or no lawful and controllable environment to run it in
- A request to see the answers in advance, to use only favourable samples, to delete failed runs, or to guarantee a pass
6.3 A tool that fails to run during the test is not "not testable." If the environment was sound and the instructions were followed and the tool died, that is an observation, recorded as "no grade — failed to run" (criteria §2.5).
6.4 A test is not rejected after the fact because the result was poor. Once accepted, it stays on the record whatever it says.
7. What you get
7.1 One document is issued: the AI Tool Field-Test Report. It carries a page per grade passed, and tests that were not run are not trimmed away (criteria §3.6).
7.2 Fields the report always carries
- Item tested, version, target type
- The task, and the four-step method
- Sample size and the result as "k of N" — never flattened into a percentage (criteria §5.8)
- The recognised grade and the point where testing stopped
- Not run — an empty field here is itself the signal
- Not applicable, with the reason
- Insufficient evidence — grades tested with no record left to judge on
- Conditions of validity
7.3 Empty fields are not filled in. What was not measured stays "not run."
7.4 How to read a grade — a grade is not a score but how far up the ladder it got. Whether the tool is usable is decided against what the job requires. For work a person checks afterwards, grade 4 is enough; for a place with no person in the loop, even grade 3 is not.
8. Using the result
8.1 Any citation states that this is an evaluation against a standard of our own, together with the standard's name, version and the report number.
8.2 Phrases that may not be used — "certified", "passed accredited testing", "performance guaranteed". What we issue is not a certification.
8.3 The grade belongs to the combination tested. It cannot be carried across to the product as a whole or to later versions (criteria §9.1).
8.4 Grade 1 being named "Proven" does not mean effectiveness is guaranteed in the field. What was proven, and how far, is written alongside it (criteria §4.6).
9. Publication
9.1 Whether a test is published is settled at intake. Once settled, it is not revisited because the result was poor.
9.2 Tests kept private are still counted publicly. The item and result stay confidential, but "N private tests" is published alongside. If only good results are visible and the rest disappear, the visible record misrepresents the whole. It is the same reason failed runs are not deleted (criteria §5.9).
9.3 Before publication you get a chance to check the facts. Facts that are wrong get corrected. Disliking the verdict is not a factual error.
10. Corrections and retesting
10.1 Factual errors are corrected. The original, the reason, the date and the superseding report are linked and published (criteria §9.4).
10.2 A verdict can be disputed on stated grounds. If it was wrong under the standard as it then stood, it is corrected or withdrawn without waiting for the standard to be revised.
10.3 A changed target is a new test. The old grade is not carried over. A new report is issued, marked as superseding the previous one.
11. Versioning
11.1 At intake, the versions of this standard and of the grading criteria in force are recorded together.
11.2 A later revision does not retroactively rewrite reports already issued. Where the verdict was wrong under the standard as it then stood, §10.1 applies.
11.3 A test this standard cannot carry is not taken by stretching the standard. It is declined.