Causal claim specification sheet
Specify the comparison object before accepting the uplift number.
Reference & Evidence
Source: Author's own worksheet, grounded in DellaVigna and Linos (2022), Gordon, Zettelmeyer, Bhargava, and Chapsky (2019), and Blake, Nosko, and Tadelis (2015). The examples have different settings, designs, and outcomes.
Each line is a claim from the register this journal publishes against, resolved from the register at build time.
- A The published-sample figure, verbatim: "an 8.7 percentage point take-up effect, which is a 33.4% increase over the average control", where "the average impact of a nudge is very large" DellaVigna & Linos (2022), Econometrica 90(1), 81–116, abstract ·
DELLAVIGNA22-C1 - A The decomposition, verbatim: "selective publication in the Academic Journals sample, exacerbated by low statistical power, explains about 70 percent of the difference in effect sizes between the two samples. Different nudge characteristics account for most of the residual difference." A model result, never a raw measurement DellaVigna & Linos (2022), Econometrica 90(1), 81–116, abstract ·
DELLAVIGNA22-C2 - A The at-scale average, verbatim: "still sizable and highly statistically significant, but smaller at 1.4 percentage points, an 8.0% increase" DellaVigna & Linos (2022), Econometrica 90(1), 81–116, abstract ·
DELLAVIGNA22-C3 - A The census, verbatim: "We assemble a unique data set of 126 RCTs covering 23 million individuals, including all trials run by two of the largest Nudge Units in the United States" DellaVigna & Linos (2022), Econometrica 90(1), 81–116, abstract ·
DELLAVIGNA22-C4 - B The randomized benchmark, with its interval: "this implies the ATT lift was 73%", and "the 95% bootstrapped confidence interval for this lift is [49%, 103%]" Gordon et al. (2019), section 7.1 ·
GORDON19-C6 - B The naive comparison, and the authors' own reading of it: exposed and unexposed conversion of 0.061% against 0.019%, "implying an ATT lift of 316%. This estimate represents the combined lift due to treatment and selection and is more than four times the lift due to treatment of 73%" Gordon et al. (2019), section 7.2 ·
GORDON19-C7 - B Matching on demographics is not enough: "exact matching on age and gender alone performs poorly, yielding a lift of 222%" Gordon et al. (2019), section 7.2.1 ·
GORDON19-C8 - B The count, and the qualification the authors attach to it: "the point estimates in 7 of the 14 studies with a checkout-conversion outcome are consistently off by more than a factor of three", while "observational methods do a better job of approximating RCT outcomes for registration and page"-view outcomes Gordon et al. (2019), section 7.3 ·
GORDON19-C9 - B The strongest result is stated as an extreme case, not a rule: "as an extreme case, we show that brand keyword ads have no measurable short-term bene"fits Blake, Nosko & Tadelis (2015), abstract ·
BNT15-C2 - B And the half usually dropped: "for non-brand keywords, we find that new and infrequent users are positively influenced by ads" Blake, Nosko & Tadelis (2015), abstract ·
BNT15-C4 - B The experimental non-brand ROI, and the contrast that makes it worth quoting: OLS "which result in a ROI of over 4,100% without time and geographic controls", while "We then used our experimental methods to control for endogeneity and found a ROI of" -63%, "rejecting the hypothesis that the channel yields any short-run positive returns" Blake, Nosko & Tadelis (2015), section 3.1 ·
BNT15-C5 - B An uplift claim is incomplete until treatment, counterfactual, estimand, interference boundary, outcome window, and measurement object are specified Author framework grounded in DellaVigna and Linos, Gordon et al., and Blake et al. ·
B11-C1 - B A large exposed-versus-unexposed difference can be a selection or measurement contrast rather than a treatment effect Author framework grounded in the observational-versus-RCT comparisons above ·
B11-C2
Grades: A, verified against the printed page of the primary source · B, primary source, text layer only · C, authoritative secondary · D, reported.
Related exhibits
-
What checking did to published numbers, setting by setting
From the essay Evidence over anecdote: what a number has to survive.
-
Estimated treatment effect (ATT) on checkout conversions across model specifications
From the essay The incrementality illusion
-
Non-brand search effectiveness by consumer cohort
From the essay The incrementality illusion