Quantile Capital: The Test Says Ship. Should You?
Quantile Capital ran an experiment on its asset management product in UK.
The new checkout flow was tested against control with 20,595 users per arm over 9 days. Control converted at 8%; the variant converted 13% relatively higher. The team reports the result as significant at p < 0.05.
Two details sit further down the document. The team tested 2 variants against the same control, and checked results daily, calling the test when it crossed the threshold. A guardrail metric — refund rate — rose by 2.2 percentage points, which was reported as "not significant".
The PM wants to ship on Monday.
design
- run days
- 9
- stopping rule
- checked daily, stopped when p < 0.05
- users per arm
- 20595
- variants against one control
- 2
results
- absolute lift pts
- 1.04
- relative lift pct
- 13
- control conversion pct
- 8
guardrails
- reported as
- not significant
- refund rate change pts
- 2.2
Advise the PM. Your answer should provide:
- Analysis — whether this result supports the conclusion, given the sample and how the test was run.
- Risks — the specific ways this readout could be wrong.
- Recommendation — ship, iterate, or re-run, with what you would require.
State any assumptions you make.
80 points, 60% to pass.
- recommendation15
- market analysis15
- risk assessment25
- financial analysis25
Reveal suggested structure
Check power, then the stopping rule, then multiple comparisons, then guardrails. A p-value from a peeked test is not a p-value.