Verification record
Independent checks run by the orchestrator on claims that other lanes produced. A lane's own
assertion that a document is verified: true is not evidence; this file records what was actually
re-checked, how, and what the check returned. Nothing enters the submitted application until it
appears here as PASS.
1. AFOSI file 8017D93-0/29 — PASS
Claim under test: L7_documents.json DOC-0001 through DOC-0006 cite an authenticated AFOSI FOIA
release concerning Paul Bennewitz, identifier AFOSI file 8017D93-0/29, FOIA case
2013-03291-F, hosted at openminds.tv/wp-content/uploads/DOTY-FOIA-ATR-1.pdf, with quoted
passages dated September–November 1980.
Why it needed testing: a fabricated archival identifier is the one unrecoverable error in this field, and these documents are the spine of the work sample.
Checks performed, 2026-08-04:
| Check | Result |
|---|---|
| Does the cited PDF exist and resolve? | Yes — HTTP 200, 3,508,560 bytes, application/pdf v1.7 |
| Does the file number appear in the document itself? | Yes — "we identified AFOSI file number 8017D93-0/29 as being responsive to your request", plus "Cy of AFOSI File 8017D93-0/29" and a COMPLAINT FORM header reading "AFOSI Det 1700, Kirtland AFB, NM" |
| Does the identifier appear in independent sources not derived from that PDF? | Yes, three of them — a separate Scribd copy of the release; an academia.edu paper citing "Ernest Edwards' Unredacted AFOSI Report (File 8017D93-0/29)"; and a bit.listserv.skeptic Usenet thread quoting "AFOSI is maintaining file 8017D93-0/29 identifiable with your request" |
| Is there a corroborating sibling identifier? | Yes — AFOSI case 8017D93-126 appears in separate archived material |
| Does the originating unit match what the lane stated? | Yes — AFOSI Det 1700 / District 17, Kirtland AFB |
Verdict: the document, the identifier and the quoted passage are genuine. The Usenet corroboration is the strongest single element, because it predates the PDF's web publication by years and could not have been derived from it — the identifier was circulating in skeptic correspondence long before this release was scanned and posted.
Residual caveat to carry into the sample: openminds.tv is a UFO-media host, not an archive of
record. The document is authentic, but the copy we are citing is a third-party mirror. The
application should cite the identifier and the originating body first and the mirror second, and
the FOIA requests drafted by the lane are the correct way to obtain a copy of record.
2. Story-gap absence claims — FAILED, then corrected
The catalogue lane reported zero transcript hits for Bennewitz and Doty, measured against roughly 62
transcripts because the rest were still downloading. Re-measured against the finished corpus of 128:
Bennewitz 59 mentions across 5 files, Bill Moore 14 files, Majestic 12 38 files, Jack
Parsons 100 mentions across 22 files. Full correction and the revised framing:
01_jesse/GAP_VERIFICATION_CORRECTED.md.
Rule adopted as a result: every absence claim in the final application states its own denominator — how many episodes were searched, out of how many, and by what method.
3. Y Combinator batch — contradiction resolved at source
Local project files stated S19; the optics lane stated W20. The Y Combinator company page states Winter 2020 in three places, including the founder's own bio ("I founded Legionfarm Inc. (YC W20) in 2016"). W20 is correct; the S19 reference in local files is stale and must not be reused.
4. Transcript coverage — measured, not assumed
Claimed corpus: the full back catalogue. Actual: 128 of 172 episodes (74%), which is every
publicly accessible episode. The remaining 44 are members-gated — 33 declare it in the title, and
the other 11 were established by attempting them and receiving Join this channel to get access to
members-only content. Titles, however, are visible for all 172, so any claim of the form
"never the subject of an episode" is verifiable at full coverage while any claim of the form
"mentioned N times" is measured at 74%. The application must not blur those two.
5. The host's surname — CAUGHT BEFORE IT SHIPPED
The voice lane derived a sign-off rule from the corpus: "Until next time, I'm Jesse Michaels and this is American Alchemy." The auto-captions write "Jesse Michaels" 105 times against "Jesse Michels" 12 times, so the majority reading in the corpus is the wrong one.
The correct spelling is Michels — jessemichelsmedia.com, youtube.com/@JesseMichels,
linkedin.com/in/jesse-michels-0bbb82203. The captions mishear it.
Misspelling the recipient's own surname in a writing sample submitted to him is a first-line rejection. Corrected everywhere; the sign-off reads "I'm Jesse Michels and this is American Alchemy."
Generalised rule: auto-captions are unreliable for every proper noun. Any name, place, programme or document title taken from this corpus must be re-checked against an authoritative source before it appears in writing — and for this story that means Bennewitz, Doty, Kirtland, Manzano, Parsons, Babalon, GALCIT and Aerojet all get checked individually.
6. "Zero exclamation marks" — DOWNGRADED: the test had no power. My first correction was also wrong.
The style guide asserted 0 exclamation marks in 3,053 scripted sentences and presented it as a measured fact about Jesse's register.
My first attempt to correct this was itself an error, and it is worth recording as one. I re-measured across all 128 transcripts, found 1,972,010 periods, 45,047 question marks and 23 exclamation marks, and wrote that the style guide had measured the wrong thing. But the style guide had scoped its count precisely — to 3,053 punctuated sentences from five fully-punctuated monologue documentaries — while I had counted across the whole corpus including guest speech and karaoke-format tracks. Two different populations. I had not read the scoping paragraph before contradicting it.
The claim still fails, but for a sharper reason: the observation is uninformative. The ASR
post-processor emits ! at a base rate of 23 per 2,017,080 sentence-endings. Expected exclamation
marks in a sample of 3,053 sentences is therefore 3053 × 23/2017080 ≈ 0.03. Observing zero is
exactly what you would see if Jesse shouted every second line. The measurement cannot distinguish
his register from any other register, so it is not evidence about him at all.
Punctuation here is generated by a model, not transcribed from a speaker. What survives is the direction — his register is measured rather than exclamatory — but that judgment must rest on diction ("craziest", "wild", "hardcore" carrying the emphasis) rather than on punctuation counts.
The style rule "do not write exclamations in his voice" stays. The evidence line under it changes from a false precision to an honest one.
Second generalised rule, from my own mistake: before contradicting a measurement, read its stated denominator. A number that looks wrong against a different population is not a refutation; it is a category error, and publishing it as a correction would have been worse than leaving the original claim alone.
Generalised rule: before quoting a metric derived from this corpus, ask whether it measures the speaker or measures the transcription pipeline. Word choice, phrase frequency and topic are the speaker. Punctuation, capitalisation and sentence boundaries are the pipeline.
7. Jack Parsons FBI file 65-59589 — PASS, with one URL correction
Claim under test: the Parsons lane cites FBI HQ file 65-59589 ("John Whiteside Parsons … Espionage – IS") on the FBI Vault, and quotes passages on the September 1950 removal of jet propulsion documents from Hughes Aircraft, the 1948 clearance suspension and its 1949 reversal, and the 1951 declination of prosecution.
Checks performed, 2026-08-04:
| Check | Result |
|---|---|
| Does the file exist on the official FBI Vault? | Yes — vault.fbi.gov/john-parsons-marvel-parsons lists "John Parsons (Marvel Parsons) Part 01 (Final)" |
| Does the cited download URL resolve? | No — HTTP 403. The /at_download/file path is bot-blocked. The citable URL is the /view path |
| Is the file number corroborated independently of the Vault? | Yes, twice — the scholarly history The Development of Propulsion Technology for US Space Launch Vehicles, 1926–1991 cites "FBI file on Parsons, no. 65-59589, p. 8, seen at FBI Headquarters in Washington, D.C."; Levenda's Sinister Forces cites "BUFILE 65-59589, p. 34–36" plus a McInerney-to-Hoover reply dated 18 January 1951 |
| Does the case designation match? | Yes — "Espionage – IS" appears in both the lane's citation and the independent one |
| Does the Caltech archival object resolve? | Partially — object 108252 returns HTTP 202 from ArchivesSpace rather than a clean 200; re-check in a browser before citing |
Verdict: the file, the number and the designation are genuine, and the corroborating citation from
a NASA-lineage technical history by a researcher who read the file at FBI Headquarters is the
strongest provenance in the whole dossier. Fix the URL to the /view path.
A finding worth using rather than just checking: the Vault page's own metadata gives a creation date of 26 November 2024 and publication of 27 November 2024. That dates the page, not necessarily the first release of the file — state it that way — but it means this primary material went online recently, and no episode has been built on it.
SUPERSEDED 2026-08-03 for pitch purposes. The file is genuine and check 7 stands, but the story it supports is not available: Parsons is already an episode. See check 8.
8. "Parsons and Bennewitz are uncovered stories" — FAIL. Both are made episodes.
This was the load-bearing claim under the entire work sample, and it is false. Measured against the complete corpus of 166 of 172 episodes, 3,750,413 words.
Parsons. 58% of all his mentions sit in two cuts of one episode — NASA & UFOs: Pagan Rituals, Secret Science & Time Travel and its members release NASA's Dark Secrets. "Pagan Rituals" in the title is the Parsons and Babalon material.
Bennewitz. The episode is "Aliens" Drove This Man Insane! (Ft. Greg Bishop). Bishop wrote Project Beta, the book that established the AFOSI operation against Bennewitz. Jesse did not just cover the story, he interviewed the person who documented it.
Why three separate passes missed this. The names are exactly the ones a speech recogniser gets wrong, so searching the correct spelling searches for the version that is not in the file. Captions render Bennewitz as "Benowitz" in 170 of 192 occurrences — a correct-spelling search finds 11% of the evidence, and the missing 89% reads precisely like proof that the topic was never covered. "Babalon" is always "Babylon"; Vallée is "Jacqu Valet"; Townsend Brown is "Thomas Towns".
The missing test was concentration. Counting mentions and checking titles is not enough: if one episode holds most of the mentions, that episode is the show about the topic whatever its title says. Full method and the surviving candidate in GAP_VERIFICATION_CORRECTED.md.
Generalised rule: in a machine-generated corpus, absence of evidence is usually evidence about the transcriber. Before claiming anything is uncovered, enumerate the plausible misspellings, measure concentration, and check whether the topic's definitive author is already in the guest list.
9. Kimi's Phase-2 gap re-run — FAIL, same error class, caught on arrival
The Phase-2 research lane re-ran the gap numbers at full coverage and reported "Parsons 241 mentions / 0 titles ... every earlier gap claim survives directionally", plus "Babalon 0".
Both are the check-8 failure repeating. The lane counted mentions and title-absence without the
concentration test, and searched "Babalon" — the spelling that appears zero times because the
recogniser always writes "Babylon", which appears 12 times across 2 episodes. Its 241 also
double-counts, since it summed .en and .en-orig tracks for the same videos; the deduplicated
figure is 116.
Rejected. The Phase-2 dossier's other findings stand, but its gap section is superseded by check 8 and must not reach the application. Recorded here because the same error has now occurred three times from three different directions, which is what a systematic flaw looks like rather than a slip.
10. Jesse's father and godfather — PASS, verified in his own scripted words
Claim under test: the Phase-2 lane reported that Jesse's father is Barry Michels, co-author with Phil Stutz of The Tools, the practice behind Jonah Hill's Netflix documentary Stutz.
Verified directly against the corpus, not accepted on report. The episode is
AVJEXCTAJUc, Why Jonah Hill Made A Documentary About My Family, and Jesse says it himself in
scripted narration:
"not only am I talking about Phil, I'm also talking about my dad Barry Michels [captions: "Michaels"], Phil's longtime business partner who helped develop Phil's ideas and communicate them to the world"
"please welcome this week's amazing American Alchemists Phil Stutz and my father Barry Michels"
"together Phil and my dad wrote a New York Times best-selling book called The Tools and a follow-up called Coming Alive"
One refinement the lane missed, and it is the better fact. Phil Stutz is not merely the father's business partner — he is Jesse's godfather, stated twice in the cold open: "that's my godfather Phil Stutz" and "why did Jonah Hill make a movie about my godfather". Jesse also states that he practises the method himself: "I obviously adhere to the tools".
And my own miss, recorded. Earlier the same evening I ran a title search for "Hill" while eliminating gap candidates, saw Why Jonah Hill Made A Documentary About My Family, and dismissed it as an unrelated Jonah Hill reference. It was the single most biographically important title in the index. A title search returns the string, not the meaning — anything that survives a keyword filter still has to be read.
11. Relman's "unprecedented" — PASS, and stronger in context than reported
Claim under test: the Mack lane reported that the chair of the Harvard ad hoc committee described the inquiry as unprecedented, in print, over his own name.
Fetched the source directly rather than accepting the quotation. Arnold S. Relman, "The Motivation for the Mack Inquiry," Harvard Crimson, 13 September 1995. The sentence is there verbatim:
"Mr. Dershowitz questions the justification for our unprecedented inquiry into the work of a tenured member of the faculty."
Three things the summary understated, all of which strengthen the sample:
- Relman's standing. Not merely the committee chair — Professor Emeritus of Medicine at HMS and Editor-in-Chief Emeritus of the New England Journal of Medicine. The concession comes from the most credentialed possible source on the prosecution side.
- The opponent is Alan Dershowitz, who defended Mack in the same paper on 30 June 1995 and again on the front page of the Washington Post on 4 August 1995. This reframes the story: not a silencing but a documented public argument between two Harvard heavyweights about whether the investigation was legitimate at all. Both put it in writing.
- Relman's actual objection is narrow and hard to dismiss: the committee was critical "not because he was interested in the 'abduction' phenomenon but because he wasn't doing any scientific research on the problem." That is the steelman, in the prosecution's own words, and it must be answered rather than quoted past.
12. Greer's "threat narrative" funder — CORRECTED. Not Rockefeller.
Claim under test: the Mack lane flagged, from unlabelled captions, that Greer's alleged funder of the abduction researchers may be the Crown Prince of Liechtenstein rather than Laurance Rockefeller, and asked for the tape to be checked.
Resolved directly against the transcript of areO7Mej44E, "UFO Secrets Are Held By A Global Cabal".
The lane was right, and the antecedent chain is unambiguous. Greer is discussing the Crown Prince of
Liechtenstein — captions render him "Likensstein, Hansum" — then: "he was coming to visit Lawrence
Rockefeller… and he told me, he says, 'The reason I'm funding Bud Hopkins and even back then uh John
Mack and some other people…'" The stated motive is eschatological: "so that we'll have Armageddon.
So Christ will return. I'm quoting."
Rockefeller is the person being visited, not the funder in this passage. Greer separately says Rockefeller funded his own work early on and that it "got intercepted." Two distinct claims, routinely collapsed into one — including by me, in the brief I wrote for the Mack lane, which asserted the Rockefeller version as background. The lane corrected its own instructions, which is the behaviour to reward.
Everything here remains REPORTED: it is one man's uncorroborated account of a private meeting. What is FACT is only that he said it on the show.
Generalised rule: when a claim turns on "he said", resolve the pronoun before repeating the sentence. Captions carry no speaker labels and no paragraph breaks, so an antecedent two sentences back is invisible unless you go and look for it.
13. The scoring corpus against its own rubric — FAILED IN ONE LANE, CORRECTED
Claim under test: that every one of the 801 scored claims carries a confidence band consistent with its own number, and a number consistent with its own components. This was never asserted by a lane; it is the assumption every other page on this site rests on, which is precisely why it needed a machine to check it rather than a reader to trust it.
Why it needed testing: the rubric's whole argument is that a reader can disagree with one component rather than with a vibe. A band that does not follow from its components destroys that argument silently — the label still reads documented, and nothing on the surface indicates that the arithmetic underneath disagrees.
Three mechanical checks were run over all ten lanes on 2026-08-05: does band follow from
ccs at the published thresholds; does ccs equal the sum of its eight components,
held inside the 0–100 scale; and does any total fall outside that scale.
| Check | Result |
|---|---|
| Band follows from the number | 39 of 801 failed — all 39 in the PEOPLE lane, 45.9% of that lane. The other nine lanes: zero. |
| Total equals the clamped sum of components | Passed. Nine CRAFT rows differ from a raw sum, all of them correctly clamped to the 0 or 100 boundary. |
| Total inside the 0–100 scale | 2 failed — two ANCIENT claims stored at −1, below the floor of the band scale. CRAFT clamps, ANCIENT did not. |
The failure has a direction, and that is the finding. Of the 39 mislabelled claims, 36 carried a band stronger than their own number supported and 3 carried a weaker one. Random transcription error does not land 92% of the time on the flattering side. One lane was rounding its own work up.
The cost is measurable rather than rhetorical. Before the correction the corpus reported 114 claims at documented, the band reserved for a primary document retrieved and quoted verbatim. After re-deriving every band from its own number: 93. Twenty-one claims were being presented as documented that were not, and two of them were off by two full bands.
What was corrected, and what was not: bands were re-derived from the published thresholds and the
two out-of-range totals were clamped to the floor — 41 edits, in two lanes, with the originals kept
in 05_graph/_audit_snapshots/. No component score was touched. Rewriting the components
to justify a band the lane preferred would be the same error committed deliberately; the arithmetic
was made to govern the label, not the other way round.
One detail is worth keeping. The control room computes the band client-side from the number, so the interface had been rendering these 39 claims correctly the whole time while the stored field said otherwise. Nobody would have caught this by looking at the site. It was only visible by reading the data against the rule it claims to follow.
Generalised rule: a rubric that is not machine-checked against its own output is a style guide, not a rubric. Any field derived from another field must be recomputed on read or verified on write — storing both and hoping they agree means the copy a reader sees and the copy a reviewer audits can drift apart indefinitely. The live check now runs on every deploy and is published on the method page.
14. The triage pipeline against its own denominator — TWO PARSER FAULTS, THRESHOLD INVALIDATED
Claim under test: that the daily triage ranks 155 candidates from six sources on their merits. The published denominator invited the check by stating the per-source counts, which is the only reason the fault was visible at all.
What the denominator showed: three of the six sources — 99 of 155 candidates, 64% of the daily pull — had contributed zero rows to the shortlist. A source mix where the majority of the intake never survives is either a statement about the sources or a defect in the instrument, and the two are indistinguishable from the ranking alone.
| Fault | Effect |
|---|---|
Date parser had no format for a bare ISO date. Nature publishes dc:date as 2026-08-05. | All 75 Nature items parsed to an empty date → freshness unmeasured → −15 of 100 under the strict reading. |
Identifier was scraped by regex from link and description. Nature states it as dc:identifier, one element away. | All 75 scored zero on retrievability while holding a valid DOI → −20 of 100. |
Half the daily pull was being docked 35 points of 100 by a missing format string and an unread element. Nothing in the output indicated it: an unparsed candidate and a genuinely weak one produce the same low number, and the shortlist that resulted looked entirely reasonable.
Repairing it changed the shortlist by zero rows, and that is the actual finding. Nature's best candidate gained 34 points and rose to rank 40 — still short of the cut, because it concerns colorectal cancer and the lane does not. The ranking had been right for the wrong reason. What the fault had genuinely corrupted was the threshold: the watch line of 35 was set against the depressed distribution, where it admitted 40% of candidates and read as selective. Against the repaired distribution it admitted 141 of 155 — 91%. A filter passing nine candidates in ten is not a filter, and it would have gone on reporting a healthy-looking number indefinitely.
Correction: the line was re-derived from a quantity a parser cannot move — the four machine axes can award at most 65 points, and a candidate must now earn three-quarters of them (49) before it costs a person twenty minutes. That admitted 21 of 155 at the time, against a lane that reads about fifteen a week; the agreement is recorded as corroboration and explicitly not as the reason, because selecting a threshold for the count it produces is how a rubric is fitted to its author. Finding 15 below then corrected a scoring fault and the same line admits 2 of 155 — the line did not move, which is the whole argument for deriving it from a principle rather than from a percentile.
Not corrected: the federal-register lane still scores zero on freshness, because its newest matching document is dated 2024-12-11. That is a measurement — the UAP apparatus has filed nothing in twenty months — and adjusting the axis to make those documents rank would have deleted the most interesting thing the pipeline reports.
Generalised rule: a threshold is only as trustworthy as the distribution it was fitted to, so repairing an input silently invalidates every cut line downstream of it. Publish the denominator by source, because a shortlist alone cannot show you that two-thirds of your intake is being discarded for a reason that has nothing to do with its merit.
15. The relevance axis against the lane's own published work — SCORING FAULT; RANKING POWER NOT DEMONSTRATED
Claim under test: that lane_fit — 20 of the 100 points, the axis that decides whether a
paper is about this show's subjects — measures subject overlap. Finding 14 had repaired the inputs;
this asks whether the instrument reading them is sound.
What printing the matches showed: the axis tested each of 1,012 vocabulary terms with a plain
substring, so sting matched inside existing, pear inside appearing, lore
inside explorer, craft inside spacecraft, and occult inside radio
occultation — a standard astronomy term. Each accident paid the same floor weight as a real
one-off subject, and the axis summed every match rather than the strongest, so eight of them maxed
it. Its own comment claimed to reward “three central subjects” while the code divided by a literal
3.0; only one term in the corpus carries full weight, so the ceiling was unreachable through central
subjects and reachable through generic ones. The axis was scoring abstract length.
Correction: word-boundary matching, the three strongest hits only, and a ceiling computed from
the corpus instead of a literal. Re-scored over the same cached pull, the read band falls from 21 of 155
to 2 of 155 and mean lane_fit from 0.240 to 0.173
(00_control_room/measure_lane_fit_change.py, re-runnable). Calibration against the three published
dossiers still passes. The threshold was not touched — a percentile-shaped line would have absorbed
all three corrections and gone on reporting the same healthy share.
What the repaired axis then admits, reported rather than tuned away: the three dossiers this lane
chose to publish average 0.44 on lane_fit; the three candidates the axis now ranks highest
average 0.69, and two of them beat every published dossier. They are ordinary planetary science
— Martian crater rays, ion escape from the Martian atmosphere. So the axis screens (153 of 155 fall
below the best published score, and what it excludes really is off-lane) but does not rank: among what
survives, being about Mars outscores having a decisive test, and the lane's own record says the second
is what makes a story. An axis that works as a gate is being weighted as a score, and 20 of 100 points
is the size of that error.
Not corrected, deliberately: re-weighting against three published examples is fitting to three examples, which is the failure this entire page argues against. The test that settles it is slow — score each dossier as it is opened and check whether the axis separates opened from passed-over at usable n. Until then it stands as a known defect with a stated magnitude, which is worth more than a confident re-weighting.
Generalised rule: a scoring function that never prints what it matched on cannot be reviewed, and substring containment is not word matching. Beyond that: when a corrected instrument disagrees with your own past judgement, the honest move is to publish the disagreement with its size, not to adjust the instrument until it agrees. Three examples cannot validate an axis; they can only be fitted to.
16. My own most-quoted number, against the file it came from — RETRACTED
Claim under test: the figure I have repeated about the AGI House research pipeline — “13,733 profiles at forty-plus tool calls each.” It is mine, not another lane's, which is the only reason it is worth putting on this page: a verification record that never tests its author is a marketing page with footnotes.
What the file says. Full-enriched-DB.json is 449,758,138 bytes, too large to load, so it was
counted with a brace-depth parser reading it as bytes. 13,733 is the top-level record count —
the size of the corpus, not a count of anything performed on it. The records that carry the deployed
scorer's own completion flag, step_8_2_detailed_score_prompt_completed: true, number
13,619. So the headline figure was the wrong denominator by 114 records, and wrong in kind:
it described a file, and I was reading it as work.
The second half is worse. “40+ tool calls” is a design quota written into §4.3 of the
8.1 deep-profile prompt. That prompt mandates its own audit field, run_stats_8_1, and that field
occurs zero times in 450 MB of output; the chain that actually ran capped at five calls per
agent. Per-profile call counts do not exist in any artifact. The phrase multiplied a real count
from one generation by an aspiration from another and produced a number that was never measured.
Correction, and it is the stronger sentence: 13,619 profiles scored by the deployed pipeline
between 2025-06-24 and 2025-11-21, against $1,861.47 of provider spend across nine services — summed
by merchant from costs/transactions.csv, twenty transactions to exactly the providers the prompts
name. Money is the part a composite can never fake: the spend is dated inside the build window and
went to OpenRouter, Firecrawl, Cursor, Perplexity, SerpAPI, n8n Cloud, Bright Data, SearchAPI and
OpenAI. The résumé being submitted with this application carried the old composite until
this check; it now carries the corrected figure and, on the page, a note saying what it used to say.
Also separated, not merged: applicants_working_set = 312 ·
final_2023_hack_db_objects = 2,391 · report_attendees_with_company_data = 250. An earlier
portfolio pass added three of these into one “people researched” headline. Four denominators, four
labels, and the case page keeps them apart on the page itself.
Generalised rule: the numbers most likely to be wrong are the ones you have said out loud most often, because repetition is not evidence and it feels like it. A count of records in a file is not a count of work performed on them, and a quota written in a prompt is not a measurement of what ran — if the prompt mandated an audit field, go and see whether the field exists before quoting the quota. Check your own most-quoted figure first, and publish what it cost you.
17. The retraction itself, against every page that carried the claim — INCOMPLETE FOR A DAY
Claim under test: that check 16 above had been done. It was written, published on this page, and applied to the résumé. That felt like completion and was not.
What a grep found. Searching the whole tree for the retracted phrasing turned up the old
composite still live in two places a reader is more likely to see than this page: the
homepage, in the Selected work list — “enriched 13,733 profiles, each pass
spending forty or more tool calls”, with 13,733 profiles and ≥40 tool
calls per pass as its two summary chips — and llms.txt, the
machine-readable twin of that page, which is precisely the file an AI assistant reads instead of the
HTML. So for about a day the site simultaneously published the claim and its retraction, and the
retraction was on the page fewest people open.
Corrected: both now carry 13,619 scored, $1,861.47 measured spend; the homepage says in
its own words what the entry used to say and links here; llms.txt carries the retraction inline,
because a machine reading only that file would otherwise never learn the number was withdrawn. The
case pages and the hub were already right — they had been written against the counted figures from
the start, which is why the inconsistency existed at all: the pages built from the evidence were
correct, and the pages written about it were the stale ones.
Generalised rule: a correction is not finished when it is written down, it is finished when every copy of the old claim is gone — and the copy that matters most is the one on the page with the most readers, not the one on the page where the correction was discovered. Publishing a retraction next to an uncorrected claim is worse than either alone, because it proves the author had the right number and shipped the wrong one anyway. The mechanical form of this rule: grep the tree for the exact retracted string, and treat a non-empty result as the check failing.
18. Two standing claims nobody had ever re-read — BOTH HOLD, NOW AUTOMATED
Claims under test: that the 152 distinct videos this site deep-links into are still watchable, and that the identity assertion joining ab7.ai to lfcarry.com resolves in both directions.
Citations. The prep cards and control room hand a reader links of the form “watch this claim at 1:22:14”. That promise is only as good as the upload still being public, and a video that goes private turns a citation into a dead end with no visible symptom. Checked all 152 through the oEmbed endpoint, which answers 200 for a playable public video and 401/404 for one withdrawn: 152 of 152 resolve.
Entity edge. The homepage asserts that this person founded the organisation at
https://lfcarry.com/#organization. A one-sided assertion consolidates nothing — any site can
name any identifier, and a graph that finds nothing at the far end drops the edge. Read back from
lfcarry.com: it publishes that organisation node, names the same https://ab7.ai/#person
identifier, lists ab7.ai in that person's sameAs, and carries an ordinary
<a href="https://ab7.ai/" rel="author"> on its history page for a crawler that never
executes structured data at all. 10 of 10 checks pass, in both directions.
Why they are now scripts. Both were true when checked and both rot silently — someone else's
upload changes, someone else's marketing rewrite drops a founder block, and nothing on this site changes
appearance. check_citations.py and check_entity_graph.py run on every deploy and print
their result next to the others. They are advisory rather than blocking, because a third party's outage
is not a reason to refuse to ship a typo fix — but they are loud, which is the whole difference between
a claim that is maintained and one that was true once.
Generalised rule: a claim that depends on a surface you do not control is not verified by checking it, only by re-checking it. If a check was worth running once and the thing it tests can change without telling you, the check belongs in the deploy, not in a report.
19. The deploy's own checks, against what the site was serving — THEY COULD NOT SEE A STALE PAGE
Claim under test: that “All checks green” at the end of a deploy means a reader is being served what was just uploaded.
It did not. Every assertion in the deploy script asked whether a URL answers:
200 here, 401 on the gated trees, a noindex header, a title that is not the login form. Not one
asked whether the answer was the current one. On 5 August a deploy corrected a withdrawn figure on the work
sample, uploaded it, printed all thirty-three checks green — and ab7.ai/american-alchemy/
went on serving the previous copy for about a minute while the deployment URL already had the fix. Every
status code was correct for the whole of that minute. The check had no way to be wrong, which is a different
thing from being right.
What replaced it. check_live_parity.py hashes every shipped HTML, text and XML
file, fetches the address a reader would type, and compares content rather than status — retrying for up
to three minutes, because Cloudflare reaches its locations at slightly different moments and a single early
read is indistinguishable from a broken deploy. 243 pages, and it now blocks the deploy rather than advising
it: a stale page is this site's own fault, unlike someone else's video going private.
It found something on its first real run. Two pages disagreed with their upload and kept
disagreeing past the deadline — not a stale copy, but Cloudflare's Email Obfuscation rewriting the
string person+redacted@example.test into a [email protected] blob and injecting
a decoder script. Those addresses are deliberate redaction placeholders inside an anonymised demo interface,
so the edge was hiding a mailbox that does not exist and making the mock look broken while doing it. Fixed at
source with the documented opt-out markers; the edge then strips the markers themselves, so that exact
two-line removal is now the single named exception in the parity check, spelled out literally so it cannot
launder a genuinely stale page.
The check then caught this very paragraph. Writing the placeholder address down in order to describe the rewrite was enough to trigger the rewrite, and the deploy refused to pass until this entry was opted out too — which is a better demonstration that the thing works than any sentence claiming it does.
Generalised rule: a check that reads the status line is testing the server, not the page. If the failure you are afraid of is wrong content, then content is what the check has to compare — and when it reports a difference it should print the difference, because “stale upload” and “the edge rewrote it” look identical in a hash and need opposite responses.