{"question_id": "Q4", "question_class": "control", "prompt": "What did Brynjolfsson, Li and Raymond find about the effect of AI assistance on customer support agent productivity? Give the figure and cite the original source.", "system": "perplexity", "model_reported": "sonar", "run": 1, "started_utc": "2026-07-27", "surface": "api", "collected_via": "dataforseo/ai_optimization", "search_used": true, "error": null, "note": "Collected before this harness existed; verbatim text not retained. Coded values are in the protocol appendix. Not reconstructed here. Returned 14% and 34% for novices; named 'the NBER working paper Generative AI at Work'. 8 citations, none the QJE article.", "answer_text": null, "citations": null, "discriminator_hits": null} {"question_id": "Q4", "question_class": "control", "prompt": "What did Brynjolfsson, Li and Raymond find about the effect of AI assistance on customer support agent productivity? Give the figure and cite the original source.", "system": "anthropic", "model_reported": "claude-sonnet-4-5", "run": 1, "started_utc": "2026-07-27", "surface": "api", "collected_via": "dataforseo/ai_optimization", "search_used": false, "error": null, "note": "Collected before this harness existed; verbatim text not retained. Coded values are in the protocol appendix. Not reconstructed here. Returned 14% and 'about 35%'; named 'NBER Working Paper No. 31161, April 2023'. Reported web_search: false despite search being requested.", "answer_text": null, "citations": null, "discriminator_hits": null} {"question_id": "Q4", "question_class": "control", "prompt": "What did Brynjolfsson, Li and Raymond find about the effect of AI assistance on customer support agent productivity? Give the figure and cite the original source.", "system": "gemini", "model_reported": "gemini-2.5-flash", "run": 1, "started_utc": "2026-07-27", "surface": "api", "collected_via": "dataforseo/ai_optimization", "search_used": true, "error": null, "note": "Collected before this harness existed; verbatim text not retained. Coded values are in the protocol appendix. Not reconstructed here. Returned 14%, 35%, 5,179; named 'NBER Working Paper No. 31161' WHILE stating it was later published in the QJE in May 2025, and citing oup.com.", "answer_text": null, "citations": null, "discriminator_hits": null} {"question_id": "Q4", "question_class": "control", "prompt": "What did Brynjolfsson, Li and Raymond find about the effect of AI assistance on customer support agent productivity? Give the figure and cite the original source.", "system": "chat_gpt", "model_reported": "gpt-4.1-mini", "run": 1, "started_utc": "2026-07-27", "surface": "api", "collected_via": "dataforseo/ai_optimization", "search_used": true, "error": null, "note": "Collected before this harness existed; verbatim text not retained. Coded values are in the protocol appendix. Not reconstructed here. Returned 14% and 34%; named 'the working paper published by NBER in April 2023'. 4 citation slots, all nber.org.", "answer_text": null, "citations": null, "discriminator_hits": null} {"question_id": "Q5", "question_class": "control", "prompt": "What did the BCG study find about consultants using GPT-4? Give the figure and cite the original source.", "system": "perplexity", "model_reported": "sonar", "run": 1, "started_utc": "2026-07-27T16:48:15+00:00", "surface": "api", "collected_via": "dataforseo/ai_optimization", "search_used": true, "error": null, "note": "Named the commissioning sponsor's own marketing publication as 'the original source'.", "answer_text": "The BCG study found that consultants using **GPT-4** completed **12.2% more tasks**, worked **25.1% faster**, and produced output rated **40% higher in quality** than the control group on tasks within the model’s capability frontier. For tasks outside that frontier, GPT-4 users performed worse: BCG reported a **23% decline** in performance.\n\nThe original BCG source is **“How People Create and Destroy Value with Generative AI”** from Boston Consulting Group.", "citations": ["https://www.bcg.com/publications/2023/how-people-create-and-destroy-value-with-gen-ai", "https://media-publications.bcg.com/BCG-BHI-GenAI-Experimental-Findings.pdf", "https://jobcannon.io/research/stats/dell-acqua-hbs-24-013", "https://www.thecrimson.com/article/2023/10/13/jagged-edge-ai-bcg/", "https://www.facebook.com/aliansari0/posts/27234889272817212/", "https://www.bcg.com/news/25october2023-theres-time-and-place-for-generative-ai", "https://www.bcg.com/press/5september2024-generative-ai-knowledge-workers-consultants", "https://themelan.com/harvard-study-reveals-openai-gpt-4-boosts-creative-consulting-capabilities-but-fails-on-strategic-analysis/", "https://www.hbs.edu/faculty/Pages/item.aspx?num=64700", "https://venturebeat.com/ai/enterprise-workers-gain-40-percent-performance-boost-from-gpt-4-harvard-study-finds", "https://www.innovationtraining.org/bcg-chatgpt-ai-productivity-study-and-report/", "https://www.youtube.com/watch?v=HOqk7OyMbpI", "https://www.linkedin.com/pulse/double-edged-sword-ai-professional-services-bcg-gpt-4-holt-nguyen", "https://www.mi-3.com.au/20-09-2023/harvard-business-school-study-bcg-finds-knowledge-workers-using-chat-gpt-outperform", "https://medium.com/@takafumi.endo/how-ai-is-redefining-strategy-consulting-insights-from-mckinsey-bcg-and-bain-69d6d82f1bab", "https://www.linkedin.com/posts/bradsmi_harvard-business-school-partners-with-bcg-activity-7134952239290290176-_v6e", "https://www.linkedin.com/posts/marshallmatt_enterprise-workers-gain-40-percent-performance-activity-7112186312157253632-gGSB", "https://www.forbes.com/sites/danpontefract/2023/09/29/harvard-and-bcg-unveil-the-double-edged-sword-of-ai-in-the-workplace/"], "discriminator_hits": [{"match": "40%", "diagnoses": "preprint (deleted from the version of record)"}, {"match": "23%", "diagnoses": "sponsor marketing figure, not the paper's"}]} {"question_id": "Q5", "question_class": "control", "prompt": "What did the BCG study find about consultants using GPT-4? Give the figure and cite the original source.", "system": "chat_gpt", "model_reported": "gpt-4.1-mini-2025-04-14", "run": 1, "started_utc": "2026-07-27T16:48:31+00:00", "surface": "api", "collected_via": "dataforseo/ai_optimization", "search_used": true, "error": null, "note": "40.2% appears in NEITHER version of the paper. Traced to the jobcannon.io statistics page. Never names the academic paper; credits BCG as the study's conductor.", "answer_text": "A study conducted by the Boston Consulting Group (BCG) in collaboration with Harvard Business School, MIT Sloan School of Management, the Wharton School at the University of Pennsylvania, and the University of Warwick investigated the impact of OpenAI's GPT-4 on consultant performance. The research involved 758 BCG consultants who were assigned to one of three groups: Control Group (no access to GPT-4), GPT-4 Access Group, and GPT-4 with Training Group. Participants were tasked with 18 realistic consulting assignments. The findings revealed that consultants using GPT-4 completed 12.2% more tasks, worked 25.1% faster, and produced results rated 40.2% higher in quality compared to the control group. However, the study also highlighted challenges. When GPT-4 was used for tasks outside its current capabilities, participants performed 23% worse than those who did not use the tool.", "citations": ["https://jobcannon.io/research/stats/dell-acqua-hbs-24-013", "https://www.bcg.com/publications/2023/how-people-create-and-destroy-value-with-gen-ai"], "discriminator_hits": [{"match": "40.2%", "diagnoses": "downstream fabrication - appears in NEITHER version; traced to an SEO statistics page"}, {"match": "23%", "diagnoses": "sponsor marketing figure, not the paper's"}]} {"question_id": "Q5", "question_class": "control", "prompt": "What did the BCG study find about consultants using GPT-4? Give the figure and cite the original source.", "system": "anthropic", "model_reported": "claude-sonnet-4-5-20250929", "run": 1, "started_utc": "2026-07-27T16:49:09+00:00", "surface": "api", "collected_via": "dataforseo/ai_optimization", "search_used": true, "error": null, "note": "'40 percent of the trial group' converts a magnitude into a headcount. This error is the Harvard Crimson's, reproduced faithfully and cited correctly. '750' (vs 758) is cited to the LinkedIn Pulse post; origin of that error not established.", "answer_text": "Consultants who used GPT-4 completed on average 12.2 percent more tasks, 25.1 percent quicker, and 40 percent of the trial group produced higher quality results. However, the study also found a significant downside: consultants using AI for tasks considered outside of the frontier were 19 percent less likely to produce the correct solutions compared to those without AI. Original Source: The study was published in a working paper titled \"Navigating the Jagged Technological Frontier\" on September 22, 2023, by Harvard Business School researchers who conducted an experimental study of GPT-4 use by 750 Boston Consulting Group consultants. The paper was authored by researchers from Harvard, MIT, Wharton, BCG and Warwick Business School.", "citations": ["https://www.thecrimson.com/article/2023/10/13/jagged-edge-ai-bcg/", "https://www.linkedin.com/pulse/double-edged-sword-ai-professional-services-bcg-gpt-4-holt-nguyen", "https://aibusiness.com/nlp/harvard-study-gpt-4-boosts-work-quality-by-over-40-"], "discriminator_hits": [{"match": "40%", "diagnoses": "preprint (deleted from the version of record)"}, {"match": "40 percent of the trial group", "diagnoses": "Harvard Crimson misstatement - magnitude converted to headcount"}]} {"question_id": "Q5", "question_class": "control", "prompt": "What did the BCG study find about consultants using GPT-4? Give the figure and cite the original source.", "system": "gemini", "model_reported": "gemini-2.5-flash", "run": 1, "started_utc": "2026-07-27T16:49:14+00:00", "surface": "api", "collected_via": "dataforseo/ai_optimization", "search_used": true, "error": null, "note": "Citations returned as vertexaisearch.cloud.google.com redirect URLs; source domains recorded as reported by the API. Names the working paper, cites none of it.", "answer_text": "A study conducted by Harvard Business School in collaboration with Boston Consulting Group (BCG) found that consultants using GPT-4 significantly improved their performance. The study, which involved 758 BCG consultants, showed that those with access to GPT-4 completed 12.2% more tasks, finished tasks 25.1% faster, and produced results that were 40% higher in quality compared to a control group without AI access.\n\nThe original source for this study is a working paper titled \"Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality,\" published by Harvard Business School. The lead author of the paper is Fabrizio Dell'Acqua.", "citations": ["innovationtraining.org", "mi-3.com.au", "jobcannon.io", "legaldive.com", "innovationtraining.org", "hindustantimes.com", "venturebeat.com"], "discriminator_hits": [{"match": "40%", "diagnoses": "preprint (deleted from the version of record)"}]} {"question_id": "Q5", "question_class": "control", "prompt": "What did the BCG study find about consultants using GPT-4? Give the figure and cite the original source.", "system": "chat_gpt", "model_reported": "gpt-5.5-2026-04-23", "run": 1, "started_utc": "2026-07-27T17:00:32+00:00", "surface": "api", "collected_via": "dataforseo/ai_optimization", "search_used": true, "error": null, "note": "Strongest available model, reasoning enabled, ~3x cost of the gpt-4.1-mini call. NEVER mentions Organization Science or any published version. Reasoning summary recorded: 'There's a PDF that likely contains relevant figures, like a 12.2% increase in tasks completed, 25.1% faster performance, and over 40% improved quality. I can cite search14 even if it's not fully opened yet.' Breaks the capability-vs-surface confound: capability is not the explanation.", "answer_text": "The BCG field experiment found that **consultants using GPT-4 performed substantially better on tasks within GPT-4’s capability frontier**: they completed **12.2% more tasks**, finished them **25.1% faster**, and produced work assessed as **over 40% higher quality** than consultants without AI access. The original study was **“Navigating the Jagged Technological Frontier”**, conducted with **758 BCG consultants**.\n\nImportant caveat: for a task **outside** GPT-4’s effective capability frontier, BCG reported that consultants using GPT-4 performed **23% worse** than those who did not use it.", "citations": ["https://www.hbs.edu/ris/Publication%20Files/24-013_d9b45b68-9e74-42d6-a1c6-c72fb70c7282.pdf", "https://www.bcg.com/publications/2023/how-people-create-and-destroy-value-with-gen-ai"], "discriminator_hits": [{"match": "40%", "diagnoses": "preprint (deleted from the version of record)"}, {"match": "23%", "diagnoses": "sponsor marketing figure, not the paper's"}]} {"question_id": "Q5", "question_class": "control", "prompt": "What did the BCG study find about consultants using GPT-4? Give the figure and cite the original source.", "system": "chat_gpt", "model_reported": "consumer interface (model not disclosed)", "run": 1, "started_utc": "2026-07-27T16:58:00+00:00", "surface": "consumer", "collected_via": "dataforseo/chat_gpt_scraper (US, en, force_web_search)", "search_used": true, "error": null, "note": "CITES THE VERSION OF RECORD (4 of 8 slots) and names both versions correctly. Gets the outside-frontier units right ('19 percentage points') where the published abstract does not. STILL reports 'more than 40% higher in quality' and cites the Organization Science page for a claim that paper does not contain.", "answer_text": "The figure most commonly cited from the original BCG/Harvard study is: Consultants using GPT-4 completed 12.2% more tasks, completed them 25.1% faster, and produced work judged to be more than 40% higher in quality than the control group on consulting tasks that were within GPT-4's capabilities. The study also found an important limitation: on a complex business problem intentionally designed to be outside GPT-4's capability frontier, consultants using GPT-4 were 19 percentage points less likely to reach the correct answer than those without AI, illustrating what the authors call the \"jagged technological frontier\". Original source: Dell'Acqua, F., McFowland III, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. \"Navigating the Jagged Technological Frontier...\" Harvard Business School Working Paper No. 24-013 (2023), later published in Organization Science (2026).", "citations": ["https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838", "https://www.hbs.edu/ris/download.aspx?name=24-013.pdf"], "discriminator_hits": [{"match": "40%", "diagnoses": "preprint (deleted from the version of record)"}]} {"question_id": "Q4", "question_class": "control", "prompt": "What did Brynjolfsson, Li and Raymond find about the effect of AI assistance on customer support agent productivity? Give the figure and cite the original source.", "system": "chat_gpt", "model_reported": "consumer interface (model not disclosed)", "run": 1, "started_utc": "2026-07-27T16:59:00+00:00", "surface": "consumer", "collected_via": "dataforseo/chat_gpt_scraper (US, en, force_web_search)", "search_used": true, "error": null, "note": "The only response in the study to volunteer the version revision unprompted. Close to a pass. A leaked internal citation token ('urlNBER working paperturn0search0') appears in the raw markdown — recorded as a surface artefact.", "answer_text": "In the original working paper, Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond found that providing customer support agents with a generative AI conversational assistant increased productivity by 14% on average, where productivity was measured as issues resolved per hour. They studied the staggered rollout of the AI assistant across 5,179 customer support agents. Average productivity increase: 14%. Novice and lower-skilled agents: 34% improvement. Experienced and highly skilled agents: Minimal impact. Original source: Brynjolfsson, E., Li, D., & Raymond, L. R. (2023). Generative AI at Work. National Bureau of Economic Research, Working Paper No. 31161. Note: A later peer-reviewed version published in The Quarterly Journal of Economics reports a 15% average productivity increase using a slightly revised sample and analysis. The 14% figure is the one reported in the original 2023 working paper that is most commonly cited.", "citations": ["https://www.nber.org/papers/w31161", "https://www.gsb.stanford.edu/faculty-research/publications/generative-ai-work"], "discriminator_hits": [{"match": "14%", "diagnoses": "preprint"}, {"match": "5,179", "diagnoses": "preprint"}, {"match": "34%", "diagnoses": "preprint-only figure"}, {"match": "15%", "diagnoses": "version of record"}]}