learn › Business Stats & Decisions

A/B Tests

lesson · about 3 minutes

Compare two versions fairly, read the result honestly, and ask if it's worth it.

Pen and paper is fine · no calculator needed why?

Opens at level 41. You're level 1. You can read and practice here now.

the idea

An A/B test shows two versions to comparable groups and compares their rates. Say the difference two ways. Percentage points subtract the rates: 4% to 5% is 1 percentage point. Relative lift divides that by the first rate: 1 ÷ 4 = 25%.

A difference in a test can be luck. A 95% interval shows the range of effects the data allow. If it includes zero, "no effect" hasn't been ruled out. And "significant" means the result would be unlikely with no real effect, not that the effect is big or worth its cost.

techniques

Percentage points, then lift

Comparing two conversion rates.

  1. Turn each count into a rate: conversions ÷ recipients.
  2. Percentage points: B's rate minus A's rate.
  3. Relative lift: that difference divided by A's rate.
worked example

Example: A gets 30 purchases from 500 recipients; B gets 45 from 500. What is B's relative lift over A?

  1. A = 30 ÷ 500 = 6%, B = 45 ÷ 500 = 9%.
  2. Difference: 3 percentage points.
  3. Lift: 3 ÷ 6 = 0.5 = 50%.

Answer: 50%

Find both ends

Reading an estimate with a 95% interval.

  1. Work out the low and high ends of the interval.
  2. Both ends above zero: the data point to a real lift. Then check it is worth the cost.
  3. Zero inside: unclear. It proves neither a win nor no effect.
worked example

Example: An estimated lift is 3 percentage points, give or take 4. What is the low end of the interval, in percentage points? (Use a minus sign if negative.)

  1. 3 − 4 = −1, 3 + 4 = 7, so zero is inside.

Answer: −1

Significant, then worth it

Deciding whether a significant result is worth keeping.

  1. Significant means unlikely if there were no real effect.
  2. Net result: what it adds minus what it costs to run.
  3. A real effect can still lose money.
worked example

Example: A significant change adds $400 a month of contribution (revenue minus variable costs) and costs $550 a month to run. What is its net monthly result? (Use a minus sign if negative.)

  1. $400 − $550 = −$150 a month.
  2. Likely real, but it loses money.

Answer: −$150

watch out for

practice

Sign in to try one

Difference in points

worked example

Variant A converts 48 of 200 recipients; B converts 350 of 1,000. What is the difference in percentage points (B minus A)? (Use a minus sign if negative.)

Answer: 11 points

  1. A = 48 ÷ 200 = 24%; B = 350 ÷ 1,000 = 35%.
  2. The difference is 11 percentage points. On its own this doesn't prove the difference is real.

Relative lift

worked example

A gets 4 purchases from 100 recipients; B gets 5 from 100. What is B's relative lift over A?

Answer: 25%

  1. Rates: A = 4%, B = 5%.
  2. (B − A) ÷ A = 1 ÷ 4 = 25%.
  3. That is a 25% relative lift; in absolute terms the rates differ by 1 percentage points.

Read the interval

worked example

An estimated conversion lift is 1.5 percentage points, with a 95% interval from −2.5 to 5.5 points. What can you conclude?

  1. B is proven better
  2. The effect is proven to be zero
  3. There's a 95% chance the lift is exactly 1.5 points
  4. It includes zero, so it doesn't rule out no effect

Answer: It includes zero, so it doesn't rule out no effect

  1. The interval runs from −2.5 to 5.5, so zero is inside it: "no effect" hasn't been ruled out, but it hasn't been proven either.
  2. An interval shows the range of effects the data are consistent with.

Significant, but worth it?

worked example

A statistically significant change adds $3,270 of expected contribution a month and costs $3,920 a month to run. What is its net monthly result? (Use a minus sign if negative.)

Answer: -$650.00

  1. $3,270 − $3,920 = −$650.
  2. "Significant" means the result would be unlikely if there were no real effect, not that the effect is big or worth paying for.

Design a fair test

worked example

You want to compare two pricing pages. Which plan is fair?

  1. Run A this week and B next week
  2. Randomly split comparable site visitors between A and B at the same time
  3. Give A to longtime customers and B to new prospects
  4. Show B only to people who responded last time

Answer: Randomly split comparable site visitors between A and B at the same time

  1. Random assignment at the same time makes the groups alike, so the only systematic difference is the thing you changed.
  2. Decide the sample size and the metric before you look.

Sign in to start