Customer Churn Analysis Practice Pack: Dataset, 10 Business Questions and Evaluation Rubric

Oct 8, 2026 | Practice Set

 Customer churn is rarely presented as a clean question with a ready-made formula. A business leader is more likely to say, “Customers are leaving. Find out why and tell us what to do next.” The analyst must define churn, validate the data, identify patterns, estimate the commercial impact and turn the findings into decisions. 

This customer churn analysis practice pack recreates that situation. It gives you a synthetic subscription dataset, 10 business questions and a 100-point evaluation rubric. Complete it in SQL, Python, Excel, Power BI, Tableau or a combination of tools. The goal is to demonstrate accurate analysis, sound business reasoning and clear communication. 

Practice resource: Download the customer churn analysis dataset 

You are a data analyst at a fictional subscription company. Management has noticed uneven customer retention but does not know whether the problem is connected to pricing, product engagement, support experience, failed payments or particular customer segments. 

The leadership team wants your analysis to answer four broad questions: 

  • How large is the churn problem? 
  • Which customers or segments are most affected? 
  • Which behaviours are associated with churn? 
  • What should the company prioritise to improve retention? 

Your submission should be useful to product, customer-success and commercial teams. Avoid stopping at observations such as “Premium customers have a higher churn rate.” Explain the segment size, business relevance and what should be tested next. 

The downloadable CSV contains 3,000 fictional customer records and 20 fields. No real customer or personal information is included. The data represents a customer-level snapshot dated 31 August 2026 and contains a mixture of active and churned accounts. 

The dataset was generated specifically for this exercise with fixed rules and a repeatable random seed. Related patterns were built across engagement, payment, support, tenure and churn fields, but there is no single “correct” explanation. Minor missingness, inconsistent categories and duplicates were inserted intentionally. Because the data is synthetic, your findings demonstrate analytical technique rather than real customer behaviour. 

Field group Included variables What you can examine 
Customer profile customer_id, region, acquisition_channel Segment size and acquisition quality 
Subscription signup_date, snapshot_date, plan_type, monthly_fee, discount_pct, auto_renew, tenure_months Pricing, renewal and tenure patterns 
Engagement logins_30d, last_login_days_ago, features_used Recent activity and product adoption 
Service support_tickets_90d, avg_resolution_hours, csat_score Support demand and customer experience 
Payment and value payment_failures_90d, lifetime_revenue Payment friction and commercial exposure 
Outcome churned, churn_date Churn status and timing 

Before analysing the file, define churn and the population included in the denominator. Treat churn_date carefully. It records an outcome after churn occurred, so using it in a pre-churn prediction model would create target leakage. 

Work through the questions in order. They move from data reliability and descriptive analysis to segmentation, commercial impact and action. 

1. Is the dataset reliable enough to analyse? 

Check customer-level uniqueness, missing values, date logic, category consistency and numerical ranges. Decide how duplicates and missing service metrics should be treated. Document every material cleaning rule instead of silently deleting records. 

2. What is the overall customer churn rate? 

State the calculation before reporting the result. Identify the eligible population, count churned customers and explain the time context. If you calculate more than one version, such as customer churn and revenue churn, label each metric clearly. 

3. Which customer segments have the highest churn? 

Compare churn across plan type, region, acquisition channel and auto-renew status. Show both the rate and the number of customers in each segment. A high percentage based on a very small group should not automatically become the leading priority. 

4. How does tenure relate to churn? 

Create meaningful tenure bands, such as 0–3, 4–6, 7–12, 13–24 and 25-plus months. Determine whether churn is concentrated early in the relationship or among longer-standing customers. Explain how the pattern could influence onboarding or lifecycle communication. 

5. Are less-engaged customers more likely to churn? 

Examine recent logins, days since the last login and the number of features used. Look for thresholds or combinations that distinguish low-engagement customers. Do not assume that reduced engagement causes churn. It may be a warning signal, a consequence of dissatisfaction or both. 

6. Is customer-support experience associated with churn? 

Compare support-ticket volume, average resolution time and satisfaction scores for retained and churned customers. Separate customers with no support contact from customers whose satisfaction score is missing. Those situations have different meanings and should not be grouped automatically. 

7. Do payment problems, fees or discounts reveal retention risks? 

Analyse payment failures, monthly fees, discount levels and renewal settings. Determine whether payment friction appears more important than price alone. Check whether discounting is associated with stronger retention or simply concentrated among already vulnerable segments. 

8. Which combinations identify the most vulnerable groups? 

Move beyond single-variable comparisons. For example, compare low-engagement customers with and without auto-renew, or customers with payment failures across tenure groups. Build two or three understandable risk segments that a business team could actually use. 

9. How much customer value is connected to churn? 

Estimate the monthly fees and lifetime revenue represented by churned customers. Identify whether the largest churn segment is also the most commercially important. State the limitations of the estimate, particularly if future revenue, margins and reactivation are not modelled. 

10. Which three retention actions should management prioritise? 

Turn the evidence into a ranked action plan. For every recommendation, specify the target segment, supporting finding, proposed intervention and success metric. Where causation is uncertain, recommend a controlled test instead of presenting the action as guaranteed to work. 

A complete project should contain: 

  • Reproducible cleaning steps or a clearly documented cleaned dataset 
  • SQL queries, a Python notebook or an analysis workbook 
  • Five to seven visuals, each connected to a business question 
  • A one-page executive summary with the most important findings 
  • Three prioritised and measurable retention recommendations 
  • Assumptions, limitations and suggested next analyses 

Keep the presentation selective. Every chart should clarify a comparison, trend or decision. Ten questions do not require ten charts. 

Use this rubric to score your work or review another analyst’s submission. 

Evaluation area What good work demonstrates Points 
Data validation and cleaning Identifies duplicates, missingness and inconsistencies, with justified treatment 15 
Metric accuracy Defines churn, uses correct denominators and labels calculations clearly 15 
Diagnostic analysis Uses meaningful segments, comparisons and multi-factor analysis 20 
Business interpretation Explains why findings matter without confusing association with causation 20 
Visual communication Uses readable, decision-focused visuals with clear labels 10 
Recommendations Prioritises specific actions supported by findings and measurable outcomes 15 
Reproducibility Documents tools, assumptions and steps well enough to repeat the work 5 
Total  100 

Score interpretation: 

  • 90–100: Portfolio-ready. Accurate, commercially relevant and easy to follow. 
  • 75–89: Strong submission with a few gaps in depth, communication or prioritisation. 
  • 60–74: Technically acceptable, but the business interpretation needs development. 
  • Below 60: Major issues in data treatment, calculation accuracy or decision support. 

Several mistakes can make a polished churn dashboard unreliable: 

  • Calculating churn without defining the population or period 
  • Reporting rates without showing customer counts 
  • Treating correlation as proof that a factor caused churn 
  • Filling every missing value with zero, even when zero has a separate meaning 
  • Using outcome information, such as churn_date,  in a predictive feature set 
  • Recommending broad discounts without estimating cost or identifying a target segment 
  • Presenting many charts but no prioritised conclusion 

Your analysis becomes more credible when you make uncertainty visible. If the dataset cannot prove why customers left, say so and recommend the next data source or experiment needed. 

Do not upload only a notebook or dashboard screenshot. Present the problem, approach, key findings, supported decisions and limitations. Recruiters need to see how you think, not only which tools you used. 

After completing the pack, explore practical data analytics, SQL, Python and business-intelligence challenges on CompeteX. Professionals who want to demonstrate a broader skill profile can explore AuthenX, while ConnectX provides a community for discussing data, analytics and AI. 

The wider PangaeaX ecosystem helps data professionals practise, demonstrate and develop their capabilities. Start with the churn dataset, complete the analysis without looking for a ready-made answer and use the rubric to identify exactly what to improve next. 

Stay Updated with PangaeaX

Subscribe to our newsletter for the latest insights, updates, and
opportunities in data science.