The same text-only model answered 31 verified questions about Lex Fridman argues with ThePrimeagen about optimal number of monitors (8:18, a podcast interview) twice: once given only the YouTube auto-caption transcript, once given the Screensay document. A second model scored each answer against a reference written by a model that had watched the video.
Accuracy is the mean judge score divided by 2, as a percentage: each answer scored 0 (wrong), 1 (partial) or 2 (right). This is one video. Results for the other six are not estimated here; they will be added as rows when they are scored.
Difficulty 1 questions are the sanity floor: a transcript alone answers them, and the document must not lose any. It did not.
| Difficulty | n | Transcript only | Screensay document |
|---|---|---|---|
| 1 (a transcript answers it) | 8 | 100.0% | 100.0% |
| 2 (needs the picture or the delivery) | 15 | 46.7% | 96.7% |
| 3 (needs both, or a judgement) | 8 | 56.2% | 87.5% |
All 12 categories, in the report's order. On the 14 questions that need the picture (literal on-screen text, visual only, gesture, production, temporal order, cross-modal) the transcript scored 6/28 points, 21.4%, and the document 27/28, 96.4%. On the 17 transcript-flavoured questions the transcript scored 33/34, 97.1%, and the document 32/34, 94.1%: the document lost one question on humour, where a whispered callback was split across two caption cues and the document quoted only half of it.
| Category | n | Transcript only | Screensay document | Delta |
|---|---|---|---|---|
| Literal on-screen text button labels, a CSS line read from a code-editor inset | 2 | 0.0% | 100.0% | +100.0 pts |
| Cross-modal what was on screen while something was said | 3 | 0.0% | 100.0% | +100.0 pts |
| Production and editing a picture-in-picture inset, a search-result overlay | 2 | 0.0% | 100.0% | +100.0 pts |
| Gesture, pointing, deixis the miss is a gesture direction | 2 | 0.0% | 75.0% | +75.0 pts |
| Visual only the colour of a tie | 2 | 50.0% | 100.0% | +50.0 pts |
| Temporal order a bottle before a sheet of paper | 3 | 66.7% | 100.0% | +33.3 pts |
| Spoken content tie | 3 | 100.0% | 100.0% | 0 |
| Counting tie; all three were spoken counts | 3 | 100.0% | 100.0% | 0 |
| Speaker attribution tie; the document's speaker label was wrong and the reader recovered from the picture lines | 3 | 100.0% | 100.0% | 0 |
| Narrative arc tie | 3 | 100.0% | 100.0% | 0 |
| Emotion and tone tie; the words alone got both readers there | 2 | 100.0% | 100.0% | 0 |
| Humour and subtlety the transcript wins: one whispered callback was split across two caption cues and the document quoted half of it | 3 | 83.3% | 66.7% | -16.6 pts |
The test set is one video per genre. Verified reference questions exist for all seven; these six have no scored result yet and nothing above is extrapolated to them.
The scripts and the full report, including the failure analysis for every question the document got wrong, are in the repository: eval/gen_reference.py, eval/run_eval.py, eval/QUESTION_CATEGORIES.md, eval/results/REPORT.md, docs/EVAL_REPORT.md. Report generated 2026-09-06. github.com/screensay