A B Testing Cheat Sheet for Data Analysts: Hypotheses, Metrics, Statistical Significance and Common Mistakes 

Oct 3, 2026 | Cheat Sheet

An A/B testing question can appear simple: compare two versions and recommend a winner. In practice, you must define success, protect measurement, choose an appropriate method, and explain whether the difference matters. 

This cheat sheet provides a reusable framework for interviews and analytical work. 

An A/B test is a controlled experiment that randomly assigns eligible experimental units to two groups: 

  • Control, or A: the existing experience. 
  • Treatment, or B: the proposed change. 

The experimental unit might be a user, account, session or device. Random assignment aims to make the groups comparable before treatment. If implementation and measurement are valid, differences in outcomes can be attributed to the treatment within the limits of statistical uncertainty. 

Stage Analyst's main question 
Business problem What decision should this experiment support? 
Hypothesis What effect are we testing? 
Metrics How will success and harm be measured? 
Design Who is eligible, and how will assignment work? 
Analysis How large and uncertain is the observed effect? 
Decision Is the evidence strong and valuable enough to act on? 

“Will the new checkout work better?” is not testable until better is connected to a measurable outcome. 

Suppose an ecommerce team wants to simplify checkout. A clearer question is: Does the redesigned checkout change the purchase conversion rate among eligible visitors? 

The hypotheses are: 

  • Null hypothesis, H₀: the treatment and control conversion rates are equal. 
  • Alternative hypothesis, H₁: the conversion rates are different. 

A two-sided alternative tests either direction. A one-sided alternative tests only an increase or decrease and requires advance justification. 

Write the hypothesis before inspecting results. Changing it afterwards turns a planned test into exploratory analysis. 

Good experiments use a metric hierarchy rather than treating every number as equally important. 

Metric type Purpose Checkout example 
Primary metric Determines whether the experiment achieved its main objective Purchase conversion rate 
Secondary metric Explains user behaviour or supports interpretation Checkout completion time 
Guardrail metric Detects an unacceptable negative consequence Refund rate or average order value 

The primary metric should connect the change to the decision. Document its numerator, denominator, inclusion rules and attribution window. “Conversion rate” is incomplete if one analyst uses visitors while another uses sessions. Secondary metrics explain behaviour, while guardrails expose wider harm. 

Decide the following before assigning the first user: 

  • Eligibility: Who can enter the experiment? 
  • Randomisation unit: Is assignment by user, account or session? 
  • Allocation: What traffic split will be used? 
  • Baseline: What is the current primary-metric level? 
  • Minimum detectable effect: What is the smallest worthwhile effect? 
  • Significance level: What false-positive risk is acceptable? 
  • Power: How likely should the test be to detect the target effect? 
  • Sample size: How many units are required? 
  • Duration: Will the test cover relevant weekly or operational cycles? 
  • Stopping rule: When will data collection and analysis end? 

Sample size is not chosen from a universal rule. Detecting a small change, noisy metric or rare event generally requires more observations. 

Before analysing outcomes, validate exposure, tracking and assignment. A sample ratio mismatch occurs when observed allocation differs unexpectedly from the planned split. Microsoft Research identifies it as a symptom of possible experiment or data-quality problems, so investigate its cause before trusting results. 

P-value 

A p-value measures how incompatible the observed data are with the null hypothesis under the selected model. It is not the probability that the null hypothesis is true, nor does it measure the effect's importance. 

At a preselected significance level of 0.05, a p-value below 0.05 is commonly called statistically significant. This is a decision rule, not proof of value or error-free experimentation. The American Statistical Association advises against conclusions based only on crossing a p-value threshold. 

Confidence interval 

A confidence interval reports a range of effect values compatible with the data and method. It helps answer: 

  • Could the true effect be negative? 
  • Is a commercially meaningful improvement plausible? 
  • Is the estimate too uncertain to support a decision? 

Statistical versus practical significance 

Statistical significance concerns evidence against the null hypothesis. Practical significance asks whether the effect matters after considering revenue, user experience, cost and risk. A large experiment can identify a tiny difference with little operational value, so report both the effect and its uncertainty. 

Type I and Type II errors 

Error Meaning Experiment consequence 
Type I Concluding there is an effect when none exists Shipping an ineffective change 
Type II Failing to detect an effect that exists Rejecting a useful change 

Power is 1 - β, where β is the probability of a Type II error for a specified effect. For a broader refresher, see PangaeaX's statistics cheat sheet for data analyst interviews. 

Consider a hypothetical checkout experiment: 

Variant Visitors Purchases Conversion rate 
Control 5,000 400 8.0% 
Treatment 5,000 460 9.2% 

The observed treatment effect is: 

  • Absolute uplift: 9.2% − 8.0% = 1.2 percentage points 
  • Relative uplift: 1.2% ÷ 8.0% = 15% 

For these independent binary outcomes, a two-proportion test is a common approach. The approximate two-sided p-value is 0.032, and the approximate 95% confidence interval for the absolute difference is 0.1 to 2.3 percentage points. 

At a preselected 5% significance level, the result is statistically significant because the p-value is below 0.05 and the interval excludes zero. But the recommendation is not automatically “launch treatment.” First confirm that the planned sample and stopping rule were followed, allocation and tracking were valid, and no guardrail deteriorated. Then assess whether the likely improvement justifies implementation. 

Common mistake Better analytical approach 
Stopping as soon as p < 0.05 Set the sample size, duration and stopping rule in advance 
Choosing the winning metric after seeing results Predefine one primary metric and label additional analysis as exploratory 
Ignoring sample ratio mismatch Investigate assignment, eligibility, exposure and logging 
Reporting only relative uplift Report baseline, absolute change, relative change and confidence interval 
Treating non-significance as proof of no effect Examine the interval to see which positive and negative effects remain plausible 
Checking many metrics and segments without adjustment Limit confirmatory tests and account for multiple comparisons when required 
Focusing only on the primary metric Review guardrails and operational consequences before rollout 
Assuming randomisation fixes broken data Validate experiment assignment, instrumentation and missingness 

Use a decision-focused sequence: 

  1. Clarify the business decision and eligible population. 
  1. Define the null and alternative hypotheses. 
  1. Select the primary metric and relevant guardrails. 
  1. Explain the randomisation unit, allocation and sample-size inputs. 
  1. Validate assignment, exposure and tracking before testing outcomes. 
  1. Report absolute and relative effects, a confidence interval and a p-value. 
  1. Recommend an action based on evidence, practical value and limitations. 

A strong answer might sound like this: 

I would use purchase conversion as the primary metric and monitor order value and refunds as guardrails. I would set the randomisation unit, sample size and stopping rule before launch. After validating allocation and tracking, I would report the absolute effect, relative uplift, confidence interval and p-value, then recommend rollout only if the result is reliable, commercially meaningful and safe on guardrails. 

Analysts also need to show how they reason through realistic business problems. CompeteX offers data and scenario-based challenges that assess analytical thinking, interpretation and decision-making. 

Within the PangaeaX ecosystem, AuthenX supports skill authentication through portfolio screening and AI-led interviews, while ConnectX connects data professionals through a focused community. Together, they support progression from learning to application and credible demonstration. 

A trustworthy A/B test begins before data arrives. Define the hypothesis and metrics, plan the sample and stopping rule, then validate the experiment. 

Statistical significance is only one part of the decision. A capable analyst explains the effect, uncertainty, practical value and validity risks. Use this framework in your next interview answer, then practise applying it through realistic scenarios on CompeteX. 

Stay Updated with PangaeaX

Subscribe to our newsletter for the latest insights, updates, and
opportunities in data science.