# Repeat runs — 2026-07-27, 17:06–17:10 UTC Collected after the pilot cells, to test the protocol's own falsification condition: *"Repeat runs disagree wildly. If the same model gives a different `source_class` across three runs, the headline is variance, not provenance, and it must be reported as such."* It does not disagree wildly. It disagrees in exactly one place, and that place is informative. **Run count, stated honestly: two usable runs per system per question, not three.** The protocol asks for three. Neither cell is complete and neither is described as complete. --- ## Q4 · "What did Brynjolfsson, Li and Raymond find…" Version of record: *QJE* 140(2), 889–942 (2025) — **15%**, 5,172 agents. Preprint: NBER WP 31161 (April 2023) — **14%**, 5,179 agents, 34% novices. | System | Run A | Run B | Cites the version of record | | --- | --- | --- | --- | | Perplexity · sonar | 14%, 34% → NBER WP | 14% → NBER WP 31161 | **0 / 2** | | Claude · sonnet-4-5 | **15% flagged, QJE 140(2), 889–942 cited** | **15% flagged, QJE 140(2), 889–942 cited** | **2 / 2** | | Gemini · 2.5-flash | 14%, "35% or 38%", 5,179 → NBER WP | 14%, "34% to 35%", 5,179 → NBER WP | **0 / 2** | | ChatGPT · gpt-4.1-mini | 14%, 34% → NBER, 3 slots all nber.org | 14%, 34% → NBER, 2 slots all nber.org | **0 / 2** | **Claude passes cleanly, twice.** Verbatim, run B: > *"It's worth noting that the published version of this paper in The Quarterly Journal of Economics > (2025) reports a slightly higher figure of 15% on average, likely due to revisions and additional > data analysis between the working paper and final publication."* It reports **both** figures, names both versions, cites `academic.oup.com/qje/article/140/2/889/7990658`, and explains the difference. That is the correct answer, and it is the only system that produced it. This matters more than a uniform failure would. It establishes that **the version of record is reachable** — no system can be excused on the grounds that the paywall made it impossible. One retrieval stack found it, twice, unprompted. The others did not look. Gemini's error is separate and worth keeping: "35% or 38%" (run A) and "34% to 35%" (run B) for the novice effect. The preprint says **34%**. Neither 35% nor 38% is in either version. The figure is drifting run to run. --- ## Q5 · "What did the BCG study find about consultants using GPT-4?" Version of record: *Organization Science* 37(2), 403–423 (2026) — quality **33.9% / 29.9%**, no 40% claim. Preprint: HBS WP 24-013 (2023) — **"more than 40% higher quality."** | System | Run A | Run B | Cites the version of record | | --- | --- | --- | --- | | Perplexity · sonar | 40%, 23% — *"The original BCG source is How People Create and Destroy Value with Generative AI"* | 40%, 23% — *"The original source is BCG's own publication"* | **0 / 2** | | ChatGPT · gpt-4.1-mini | **40.2%**, 23% → jobcannon.io + bcg.com | **40.2%**, 23% → BCG's PDF only, 3 slots 1 domain | **0 / 2** | | Claude · sonnet-4-5 | "40 percent of the trial group" → the Harvard Crimson | **search failed**; answered from memory, 40% | **0 / 2** | | Gemini · 2.5-flash | 40% → names WP, cites none of it | 40% → names WP 24-013, cites none of it | **0 / 2** | | ChatGPT · **gpt-5.5** | 40%, 23% → HBS preprint PDF + bcg.com | — | **0 / 1** | **Nine API runs. Zero cited the version of record. Every one carried the 40% figure peer review deleted.** ### Two things the repeats established that one run could not **1 · The 40.2% is stable, and it has come loose from its source.** Run A cited `jobcannon.io`, which does state 40.2%. Run B cited **BCG's own PDF**, which states **40%**. The same model produced the same invented decimal twice and attributed it to two different documents, only one of which contains it. A figure that began as "more than 40%" over two experimental conditions is now circulating with a decimal place and no fixed origin. **2 · Perplexity's sponsor substitution is not a slip.** Both runs name BCG's marketing publication as *the original source*, in those words. It is the stable behaviour of that stack on this question. ### A data-quality problem I am not hiding — now closed Claude returned `web_search: false` on **2 of its 6 runs** across both questions, despite search being requested — once saying so in the answer: *"I don't have access to web search at the moment."* Those runs measure parametric memory, not retrieval, and are excluded from the counted cells. This looks like intermittent tool provisioning at the endpoint rather than a property of the product. **Re-run 2026-07-27 17:44, search confirmed active.** The cell now has two valid retrieval runs and they agree closely: - *"40 percent of the trial group produced higher quality results"* — again, cited to the Harvard Crimson. The propagation of the newspaper's category error is **stable across runs**, not a one-off. - **New in this run, and worse:** *"the lowest performers had the biggest boost — 43% increase versus 17% for top performers"*, cited to `aibusiness.com`. Those are the **preprint-only skill-split figures that peer review removed from the paper entirely** (`AI-X5`). Claude now returns *both* deleted artefacts — the 40% and the 43/17 split. - *"750 Boston Consulting Group consultants"* — again cited to the same LinkedIn Pulse post, in both runs. Consistent attribution across two independent runs makes that post the likely origin of the error, though it has not been checked at source. - Citations: `thecrimson.com`, `aibusiness.com`, `linkedin.com`, `axios.com`. **Four domains, none academic.** The trade headline it leans on is literally *"Harvard Study: GPT-4 Boosts Work Quality by Over 40%"* — the deleted figure encoded in a URL slug. **Q5 now stands at 9 valid retrieval runs across five model configurations. Zero cited the version of record. Every one carried a figure the published paper does not contain.** --- ## Why Q4 is passable and Q5 is not Both papers have an open preprint and a paywalled version of record. Only one gets cited correctly. The difference is not the paywall — it is **age and path**. - *Generative AI at Work* has been the version of record since **May 2025**. The NBER landing page links forward to it, `academic.oup.com` is indexed, and three years of citations have accumulated pointing both ways. - The *Organization Science* article appeared in **March 2026** — four months before these runs. Its preprint has had **three years** to gather links under a working-paper number that trade press, vendors and SEO pages all cite. So the mechanism sharpens: it is not "paywalled papers lose." It is **"the version of record has to out-age its own preprint before retrieval finds it"** — and during that window, which is years, the superseded numbers are the answer everyone gets. That window is where every recently-published finding lives. --- ## What still has to happen before any of this is published as a rate - A third run on every cell. Two is not three. - Claude's Q5 cell re-run with search actually enabled. - The `interested_share` domain pass, from each domain's own homepage. - The inbound-link count, preprint versus version of record, for both papers — the mechanism above is still an inference and the essay must not state it as measured until it is.