All IssuesResourcesGlossaryAdvertiseWork With Us
Search Performance Marketing
← Marketing Glossary
Definition

Statistical Significance

Also known as: Significance, Confidence Level

A measure of confidence that a difference between two variants in a test is real rather than random chance. Expressed through a confidence level, commonly ninety-five percent, it tells you whether an A/B test result is trustworthy enough to act on before you declare a winner.

Key Takeaways

  • Statistical significance measures confidence that a difference between test variants is real rather than random chance.
  • It is expressed through a confidence level, commonly ninety-five percent, for A/B test results.
  • Reaching significance means a result is unlikely to be a fluke, not that the lift is large.
  • Define sample size and test duration before launch to avoid stopping on noise.
  • Peeking and stopping early when a result looks significant inflates false positives.

How It Works

Statistical significance answers a simple question: is the difference between two variants likely real, or could it be random noise? A test compares a control against a variant and calculates how probable the observed gap would be if there were truly no difference. A ninety-five percent confidence level means there is roughly a five percent chance the result is a fluke.

Significance is not the same as importance. A change can be significant yet tiny, so read it alongside your Conversion Rate lift and whether that lift moves a metric that matters, such as your North Star Metric. A statistically clean win on a trivial number may not be worth shipping.

Significance also is not proof of lasting causal impact. To confirm a change actually drives value at scale, teams pair it with Incrementality analysis. Set sample size and duration in advance, then let the test run to completion instead of stopping the moment it looks like a winner.

Why It Matters

Without significance, you risk shipping changes based on noise, chasing lifts that evaporate at scale. Waiting for significance protects revenue by ensuring test wins are repeatable rather than lucky short-term fluctuations.

Example

An ecommerce store tests a new product-page layout against the current one. It calculates a required sample size in advance and commits to a two-week run. Midway the variant looks like a strong winner, but the team resists stopping early. By the end, results reach ninety-five percent confidence and the lift holds, so they ship the new layout knowing the win is unlikely to be luck.

Common Mistake

Peeking at results and stopping the test the moment it looks significant. Early peeking inflates false positives, so define sample size and duration before launch and let the test run to completion.

Frequently Asked Questions

What does ninety-five percent confidence mean?

It means that if there were truly no difference between variants, you would see a result this extreme only about five percent of the time. It is a threshold for trusting a result, not a guarantee the effect is real.

Why is peeking at test results a problem?

Checking repeatedly and stopping the moment a test looks significant inflates false positives, because random fluctuations occasionally cross the threshold. Set sample size and duration before launch and let the test run to completion.

Does significance mean the change is worth shipping?

Not by itself. A result can be significant but tiny. Weigh the size of the lift, whether it moves a metric that matters, and, for high-stakes decisions, confirm real impact with incrementality testing.

How much traffic do I need to reach significance?

It depends on your baseline conversion rate and the size of the effect you want to detect. Smaller expected lifts need larger samples. Use a sample-size calculator before launching to set a realistic duration.