# Retrieval-provenance audit — replication package **What provenance do you get when you ask an AI assistant for a widely-circulated commercial statistic *and explicitly ask for the original source*?** Everything needed to re-run this study is in this folder. It takes an afternoon and costs roughly the price of lunch. If you disagree with a coding decision, the untouched response is in the data file. --- ## Contents | File | What it is | | --- | --- | | `protocol.md` | The pre-registered protocol. Question-selection criteria, the ten-field coding scheme, the run design, and — written before the data — what would make the finding wrong | | `questions.json` | The six questions, the instruction appended to each, and the **ground truth established before each question was run** | | `run_audit.py` | The runner. Calls four assistants, writes every raw response to JSONL. Codes nothing | | `responses.jsonl` | The collected responses, with timestamps, model identifiers and citation lists | **On `responses.jsonl`:** these runs were collected through DataForSEO's AI-optimization endpoints, not through `run_audit.py`'s direct vendor calls. Every record says so in `collected_via`, because a replication package that hides its collection path is not one. Four records — the Q4 API cells, run before this harness existed — carry `answer_text: null` and a note; their coded values are in the protocol appendix and they are **not** reconstructed from memory here. The coding sheet is still to be built: `interested_share` requires a domain-by-domain pass that has not been run. --- ## Re-running it ```bash export OPENAI_API_KEY=... ANTHROPIC_API_KEY=... GEMINI_API_KEY=... PERPLEXITY_API_KEY=... python3 run_audit.py --dry-run # see the exact payloads, call nothing python3 run_audit.py --only Q4,Q5 --runs 1 # the two control questions, once each python3 run_audit.py --runs 3 # the full 72-response dataset ``` Responses append to `responses.jsonl`. The runner prints which pre-registered discriminator strings appeared in each answer — `[40% -> preprint (deleted from the version of record)]` — as a **prompt to a human coder, not a coding decision**. Substring matching cannot tell a figure stated as the study's result from the same figure in an unrelated aside. Every counted cell is coded by hand. **API shapes drift.** These four vendors change web-search parameters more often than they change models. A 4xx means a builder function needs updating against current vendor documentation; nothing else in the harness depends on it. --- ## Three things to know before you read the results **1 · The API surface and the consumer apps behave differently, and the difference is the finding.** The runner calls API endpoints with web search enabled. `chatgpt.com`, `claude.ai` and the Gemini app are different stacks — different retrieval, different system prompts, sometimes different models. Most people asking these questions use the product, not the API. Both were tested. See *the surface split* below; it is the most important result here and the runner alone will not reproduce it. **2 · This is not a measurement of whether the models are good.** It is a measurement of **the public evidence surface for commercial questions**, sampled through the tool most people now use to reach it. Where the sources are poor, that is a fact about what has been published as much as about what was retrieved. "AI is unreliable" is an available reading of this data, and it is the wrong one — see the Crimson case below. **3 · Two of the six questions are controls, and they carried the finding.** Q4 and Q5 were included because a clean, reachable, peer-reviewed primary source exists for each — so that if the assistants cited them correctly, the honest conclusion would narrow to "retrieval works where the literature is sound, and the problem is the literature." The protocol says so in advance, in the section headed *what would make this finding wrong*. That is not what happened. --- ## What the control cells found Both controls, all four systems, one run each, 2026-07-27. **Eight of eight returned the superseded preprint rather than the version of record.** **Q4 — Brynjolfsson, Li & Raymond.** Version of record: *QJE* 140(2), 889–942 (2025), 15% lift, 5,172 agents. On the pilot, all four returned the 2023 NBER preprint's 14% and 5,179 — Gemini even stating that the paper *"was later published in The Quarterly Journal of Economics in May 2025"*, citing `oup.com`, and reporting the preprint's numbers anyway. **Repeat runs changed this cell, and the change is the more useful result.** Across two counted runs per system, **Claude cited the version of record both times** — full pagination, `academic.oup.com`, and an unprompted explanation: *"the published version … reports a slightly higher figure of 15% on average, likely due to revisions."* Perplexity, Gemini and ChatGPT returned the preprint on every run. That matters more than a uniform failure would: it proves **the version of record is reachable**. No stack can be excused by the paywall. One found it twice without being asked. The others did not look. **Q5 — Dell'Acqua et al., the BCG/GPT-4 study.** This one discriminates cleanly, because peer review changed the headline. The preprint claims *"more than 40% higher quality"*; the published *Organization Science* paper reports **33.9%** and **29.9%** and contains no 40% claim at all. The preprint's "43% / 17%" skill-split claim was removed entirely. Across **thirteen API runs — four systems, three runs each, plus one on the strongest available model — every single one carried the 40% figure, and not one cited the version of record.** Between them the citation sets contain a student newspaper, LinkedIn, Facebook, YouTube, a legal-industry trade site, a national newspaper, a career-quiz site, a training vendor, general tech press, and the commissioning firm's own marketing page — and zero links to *Organization Science*. **Why Q4 is passable and Q5 is not.** Both papers have an open preprint and a paywalled version of record, so the paywall is not the variable. Age and path are. *Generative AI at Work* has been the version of record since May 2025, the NBER page links forward to it, and citations have accumulated both ways. The *Organization Science* article appeared in **March 2026** — four months before these runs — while its preprint had three years to gather links under a working-paper number that trade press, vendors and SEO pages all cite. So the mechanism is narrower than "paywalled papers lose." It is that **the version of record has to out-age its own preprint before retrieval finds it** — and during that window, which runs to years, the superseded numbers are what everyone gets. That window is where every recently-published finding lives. Three distinct mechanisms, each verified against primary documents: - **Version lag.** The peer-reviewed correction never surfaces. Replicated across two papers in two disciplines. - **Sponsor substitution.** BCG's own page says *"23% worse"* where the paper says 19 percentage points. Two systems reported 23% as the study's finding; Perplexity named BCG's marketing page as *"the original source"* in those words. - **False precision, acquired downstream.** `40.2%` appears in **neither** version of the paper. It appears on a career-quiz site's research-statistics page. ChatGPT returned it and cited that page. A rounded floor over two experimental conditions grew a decimal point on its way down the chain and came back as a fact with a citation attached. **The repeat run reproduced 40.2% and cited BCG's own PDF for it — a document that says 40%.** The decimal is now stable in the model's output and has come loose from any source that contains it. ### The case that decides how to read all of this Claude returned *"40 percent of the trial group produced higher quality results"* — a magnitude silently converted into a headcount, which is a claim about a different quantity entirely. It cited the Harvard Crimson. **The Crimson says exactly that, verbatim.** The model reproduced its source faithfully and cited it correctly. The error is the newspaper's. This is not a story about models inventing things. It is a story about a retrieval layer that is accurate with respect to a corpus that is wrong — with the actual paper two clicks away, and nothing in the chain having any reason to open it. --- ## The surface split Both controls were also put to the **consumer ChatGPT interface**, because a study about "the tool people use" that only tested API endpoints would be making an unstated generalisation. The consumer surface is markedly better at *citation*. On Q5 it cited the *Organization Science* DOI four times and named both versions correctly. On Q4 it volunteered, unprompted, that *"a later peer-reviewed version published in The Quarterly Journal of Economics reports a 15% average productivity increase"* — the only response in the study to flag the revision. It is no better at *content*. It still reported "more than 40% higher in quality" — and cited the *Organization Science* page for a claim that paper does not contain. **Is that just a stronger model?** No. Tested directly: the same question, the same API, `gpt-5.5` with reasoning enabled at roughly three times the cost per call. It cited the preprint PDF and BCG's marketing page, reported 40% and 23%, and **never mentioned the published version at all**. The strongest model on the API did worse on provenance than the consumer product. The retrieval stack is doing the work, not the model. `gpt-5.5`'s own reasoning summary, on the way there: > *"There's a PDF that likely contains relevant figures, like a 12.2% increase in tasks completed, > 25.1% faster performance, and over 40% improved quality. **I can cite search14 even if it's not fully > opened yet.**"* It attaches a citation to a document it has not opened, on the strength of what it expects the document to say — and what it expects is the figure peer review removed. So there are two independent failures, and only one of them is improving: | | Citation | Content | | --- | --- | --- | | API | poor — career-quiz sites, sponsor marketing | wrong figure | | Consumer | good — DOI, both versions named | **same wrong figure** | **Better citation practice made the error harder to catch, not less likely to occur.** A wrong number sourced to `jobcannon.io` can be smelled. The same number sourced to a DOI cannot. *Limits:* one system (only ChatGPT has a consumer scrape here), one run each, one day, US location. Nothing here supports a claim about consumer Claude, Gemini or Perplexity. --- ## What this does not establish Stated here rather than left for a critic to find. - **Run counts.** Both control questions have three valid retrieval runs on all four systems. One Gemini run was discarded and re-run because a length constraint had been appended to the frozen prompt; the discarded run is logged in the protocol changelog rather than quietly dropped. - **Claude's Q5 cell is effectively one retrieval run.** It returned `web_search: false` on 2 of its 5 runs despite search being requested — once saying so outright: *"I don't have access to web search at the moment."* Those runs measure parametric memory, not retrieval. This looks like intermittent tool provisioning at the endpoint rather than a property of the product, and the cell needs re-running before anything is claimed about Claude on Q5. - **The mechanism is inferred, not measured.** "The link graph rewards the free version" is the best available explanation for eight of eight, but inbound links to preprint versus version of record have not been counted. Until they are, it is an explanation. - **`interested_share` is uncoded** for the Q5 cell. The protocol requires each domain to be classified from its own homepage rather than from memory, and that pass has not been run. One classification is already settled and cuts against the thesis: `jobcannon.io` is a career-assessment site and does **not** count as interested, because it does not sell in the category the claim supports. The rule was written to bite against the coder's convenience, and it did. - **Six questions is a small set**, chosen against stated criteria, not sampled from a population. --- ## Licence and citation Data and code released for re-use and disagreement. > Isoglu, S. (2026). *Retrieval-provenance audit: what four assistants return when asked where a > number comes from.* isoglu.com.