Generate
- Runs
- Records
- Detects
- Precise
- Proven
Genspark · batch_understand_videos
Read 100 YouTube transcripts and answer the same six questions per video
496 / 600
2026-09-09
G5stopped at Records
Records from running tools ourselves for real work, written down as observed. These are not the places where the tools do well — they are about how a tool reports its own failure.
A low grade does not mean a bad tool. Most of these do their job. They stopped at the rung where a bad input still comes back as “success” — the place that turns into a silent failure once a person is no longer looking.
7 records · graded 6 (with a grade 4 · no grade 2) · head-to-head 1
All 7
Collect 3 · Generate 2 · Execute 1 · head-to-head 1
Generate
Read 100 YouTube transcripts and answer the same six questions per video
496 / 600
2026-09-09
G5stopped at Records
Collect
Pull the first page of search results for 25 keywords
25 / 25
2026-09-09
G5stopped at Records
Generate
Write an ad-revenue research report using only Naver and Google official docs
1 / 2
2026-09-09
no grade
Execute
Connect for video understanding and cross-checking
not recorded
2026-09-09
no grade
Collect
Fed it two video IDs that do not exist, to see whether it reports the failure
0 / 2
2026-09-09
G5stopped at Records
Collect
Fetch five different kinds of page: an English help doc, a JavaScript-rendered help page, Wikipedia, a government site, and a URL that does not exist
2 / 4
2026-09-09
G5stopped at Records
Head-to-head
No grade — two tools on one task
Put the same three questions (two English, one Korean) to both tools and compared which domains came back on page one
10 / 30
2026-09-09
no grade
One page: the task we gave it, how we measured, how many runs, and how many of those worked. Tests we did not run stay on the page under “not run”. Below is one we actually issued.
Test record · genspark-web-search-serp
Item tested
Genspark · web_search
gsk CLI, as of 2026-09-09
Task given
Pull the first page of search results for 25 keywords
Method
How we measured precision: three keywords (two English, one Korean), three runs each, back to back on the same day, comparing the ranked list. All nine comparisons matched exactly, order included. Every run returned nine results.
This test predates Grading Criteria v1.0, so its method was not recorded as the four repeatable steps.
Sample size
25keywords answered
Result
25 / 25
Works in both Korean and English. All 25 returned results, and the same keyword gives the same answer on a repeat run.
Grade
Grade 5 · stopped at Records
Observed failures
It does not return the ten results asked for — between six and ten depending on the keyword (stable within a keyword). It never says why the rest are missing.
Not run
Detects · Proven2 rungs not run
We only repeated within a single day. How far results drift across days is still unknown. We also did not open the ranked pages to confirm they are what they claim (the Proven rung).
Not applicable
Precise1 rungs do not arise for this item
Grade 2 asks whether clean samples are wrongly flagged. This tool returns search results and issues no verdict, so a false positive does not arise for it (§4.8.4).
Measured on
2026-09-09
Valid while
Valid only for the version and the task named above. When the tool updates, this record lapses and measurement starts over. A different task gives a different result.
The ladder was not built to measure other people's tools. It was built to sign off the tools we use for our own work. Separately from the records above, 24 in-house tools sit on the same ladder.
These are in-house tools, so no reports are published — only the counts. They are here to show the standard is not reserved for other people's software: most of what we built fell short of the grade its job required.
What each grade rests on is in Grading Criteria v1.0, and the conditions we accept work under are in Testing Standard v1.0. Tell us which clause is wrong and we will fix it.