← Every exhibit

Table Table 5 From the research bench

6. Executive Diagnostic Framework and Experimentation Audit Checklist

Audit DimensionCore Diagnostic Evaluation QuestionMaturity Scoring Criteria (1 to 5)Critical Red Flag Warning
1. Hypothesis Pre-RegistrationAre hypotheses, primary metrics, and target sample sizes formally documented prior to test launch?1: No written plans.
5: Comprehensive pre-registration template enforced in Jira.
Teams changing the primary evaluation metric after reviewing preliminary results.
2. Prospective Power SizingIs every experiment sized prospectively to achieve at least 80% statistical power for a realistic MDE?1: Guessed sample sizes.
5: Automated sample size calculations based on baseline variance.
Running tests on low-traffic pages that require two years to reach statistical power.
3. Fixed-Horizon GovernanceAre tests executed for their full pre-determined sample size and complete weekly business cycles?1: Continuous daily peeking.
5: Strict fixed-horizon rules or mathematically certified sequential testing.
Stopping tests early the first morning a dashboard turns green.
4. Automated SRM AuditingDoes the platform run automated chi-square goodness-of-fit tests to detect Sample Ratio Mismatches?1: No SRM checks.
5: Automated daily SRM alerts that lock down reporting on failure.
Evaluating conversion rates on tests with severe variant count imbalances ($p < 10^{-3}$).
5. Guardrail Metric ProtectionAre non-negotiable system and business guardrails (latency, errors, refunds) continuously monitored?1: Only conversion tracked.
5: Comprehensive telemetry dashboards with automated rollback triggers.
Shipping a conversion winner that increased server response times by 300 ms.
6. Telemetry & InstrumentationAre conversion and event tracking pixels verified through automated end-to-end integration tests?1: Manual unverified tags.
5: Automated synthetic testing verifying tracking firing across variants.
Twyman's Law violations: celebrating massive lifts caused by double-firing pixels.
7. Variance Reduction ControlsDoes the platform deploy variance-reduction techniques (such as CUPED) on high-variance metrics?1: Raw noisy metrics.
5: Automated CUPED covariate adjustment on all continuous metrics.
Inability to measure revenue metrics due to overwhelming sample size requirements.
8. Multiple Testing CorrectionAre family-wise error rates or FDR corrections applied when evaluating multiple variants or segments?1: Uncorrected p-hacking.
5: Automated Benjamini-Hochberg FDR adjustments built into reporting.
Cherry-picking obscure post-hoc demographic slices that showed significance by chance.
9. Long-Term Holdout AuditingDoes the organization maintain long-term holdout groups to verify that experimental lifts persist over time?1: Zero holdout tracking.
5: Permanent 1% to 5% holdout cohorts measuring 90-day persistence.
Short-term experimental lifts completely evaporating after 60 days due to novelty decay.
10. Win-Rate Reality CalibrationDoes executive leadership recognize that only one-third of well-formed ideas succeed in practice?1: 90%+ claim win-rates.
5: Rigorous acceptance of negative results as valuable capital protection.
Teams claiming 80%+ experiment win-rates, indicating trivial testing or rigged metrics.

Swipe or scroll horizontally if the table is wider than your screen.

Cite Embed

Reference & Evidence

Source: Table from this essay. Sources and interpretation are given in the article.