About
What AI Dad does
AI Dad runs AI tools and publishes what happened, including the failures. We test coding agents, meeting assistants, search tools, video analysis tools, and other AI software.
Jeong Min-hyeong runs the site independently from South Korea. “We” refers to the publication and its editorial standards, not a separate research team.
This domain, daddynote.co.kr, previously hosted parenting content. It now documents AI tools and the evidence behind their ratings.
Our five-level ladder
| Level | Meaning | What we check |
|---|---|---|
| Runs · 구동 | The tool works | Does it run in the stated environment and produce a result for the assigned task? |
| Recorded · 기록 | The run leaves a trace | Can we trace the result through the inputs, settings, outputs, and execution record? |
| Detects errors · 검출 | The tool recognizes mistakes | Does it identify incorrect results in tasks where we can establish what is wrong? |
| Repeatable · 정밀 | The same input produces the same answer | Do repeated runs under the same conditions meet a predefined consistency rule? |
| Verified in practice · 실증 | We check the actual result | Can we confirm the claim against the actual artifact or real-world subject? |
A rating applies to the test described. It does not certify the entire product or every possible use.
We define the requirements for each level before testing. We do not award a higher level while leaving its prerequisites unverified. When a finding covers only part of a task, we state that limit.
We also explain what counts as “the same answer.” Some tests require identical text; others compare calculated values or execution outcomes. We do not mark a result as verified in practice unless we inspect the actual artifact or subject.
How we count
We report “k out of N runs.” We do not replace the counts with a percentage.
If seven of ten runs pass, we write “7 out of 10 runs passed” and explain what happened in the other three. These numbers illustrate the format; they are not a test result.
Each test records:
- The task, inputs, environment, and relevant settings.
- The pass and fail criteria and number of runs.
- The outcome of each run and the available evidence.
- Interruptions, timeouts, errors, retries, and human intervention.
- Untested items and limits on what the results establish.
Failed runs stay in the record. A retry counts as a new run. If we exclude a run from a particular count, we retain its record and explain why.
Anything we have not tested stays “Not tested.” We do not fill the gap with an estimate or another product’s results.
Firsthand results, dated
We do not present another organization’s benchmark as our own measurement. When we refer to product documentation or outside research, we identify the source and distinguish it from our test results.
Every results card includes the test date and tool version. We also identify the model, plan, and settings when they affect the outcome. If the provider does not disclose a version, we say so and record the available model information and test date.
An AI tool can change substantially in six months. We do not present an old test as a current one. A retest receives a new date and a link to the earlier record.
Corrections and evidence
When we find an error, we record the correction, its date, and any effect on the rating.
We redact personal information, confidential material, and content we cannot lawfully publish. If a restriction prevents readers from checking part of the evidence, we disclose that limitation. Keeping a record of a failed run does not require keeping someone’s personal information public.
Advertising and disclosures
AI Dad runs Google AdSense. Ads are not being served yet while the account is under review. We distinguish ads from editorial content. Advertising revenue does not determine ratings or whether a failure appears in the record.
If a company provides access, a product, payment, or other support for a test, we disclose it in the relevant article. If a link earns us a commission, we identify it near the link.
Use the Contact page to report an error or request a retest.