Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation. Use when feeding test results, checking statistical significance, calculating sample sizes, analyzing experiment outcomes, or generating next test ideas based on results. Platform: Google and Meta.
git clone https://github.com/irinabuht12-oss/marketing-skills.git--- name: ab-test-analyzer description: Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation. Use when feeding test results, checking statistical significance, calculating sample sizes, analyzing experiment outcomes, or generating next test ideas based on results. Platform: Google and Meta. metadata: platform: Google and Meta --- # A/B Test Analyzer Evaluate A/B test results with statistical rigor and generate actionable next steps. ## Process 1. **Collect test data** - Variations, sample sizes, conversions, time period 2. **Check validity** - Runtime, sample size, peeking issues 3. **Calculate significance** - Z-score, p-value, confidence interval 4. **Segment analysis** - Device, source, new vs returning 5. **Interpret results** - Statistical vs practical significance 6. **Generate hypotheses** - "Why it worked" and next test ideas ## Sample Size Calculator ``` n = 2 × (Zα/2 + Zβ)² × p(1-p) / δ² Where: - n = sample size per variation - Zα/2 = 1.96 (95% confidence) or 2.58 (99%) - Zβ = 0.84 (80% power) or 1.28 (90%) - p = baseline conversion rate - δ = minimum detectable effect (absolute) ``` **Quick Reference (95% confidence, 80% power)**: | Baseline CR | 10% Relative MDE | 20% Relative MDE | |-------------|------------------|------------------| | 2% | 78,000/var | 19,500/var | | 5% | 30,000/var | 7,500/var | | 10% | 14,300/var | 3,600/var | ## Significance Calculation ``` Z-score = (pB - pA) / √(SE²_A + SE²_B) SE = √(p × (1-p) / n) If |Z| > 1.96 → Significant at 95% If |Z| > 2.58 → Significant at 99% ``` ## Validity Checklist - [ ] Ran for minimum 14 days (day-of-week effects) - [ ] Met required sample size before concluding - [ ] No "peeking" at results during test (inflates false positives to 40%+) - [ ] Traffic split was random and consistent - [ ] No external factors (holidays, outages, campaigns) ## Bayesian vs Frequentist | Use Case | Approach | |----------|----------| | Large samples (>10K), yes/no decision | Frequentist | | Continuous monitoring, early stopping | Bayesian | | Regulatory/legal requirements | Frequentist | | Want probability statements | Bayesian | ## Output Format ``` ## A/B Test Analysis ### Test Summary - **Test Name**: [Name] - **Hypothesis**: [What you expected] - **Runtime**: [X days] - **Sample Size**: Control [n] / Variant [n] ### Results | Metric | Control | Variant | Lift | Significant? | |--------|---------|---------|------|--------------| | Conversion Rate | X% | Y% | +Z% | ✅/❌ | ### Statistical Details - Z-score: [X] - P-value: [X] - Confidence Interval: [X% to Y%] - Power achieved: [X]% ### Validity Assessment - ✅/❌ Minimum runtime met - ✅/❌ Sample size requirement met - ✅/❌ No external confounds identified ### Segment Breakdown | Segment | Control CR | Variant CR | Lift | Notes | |---------|-----------|------------|------|-------| | Mobile | | | | | | Desktop | | | | | ### Interpretation [Statistical significance vs practical significance discussion] ### "Why It Worked" Hypotheses 1. [Hypothesis about mechanism] 2. [Alternative explanation] ### Next Test Ideas 1. [Follow-up test based on learnings] 2. [Adjacent hypothesis to explore] ### Confidence Level: [HIGH/MEDIUM/LOW] ``` ## Example **Input**: "Tested new headline. Control: 10,000 visitors, 250 conversions. Variant: 10,000 visitors, 300 conversions. Ran for 7 days." **Output**: ``` ## A/B Test Analysis ### Test Summary - **Hypothesis**: New headline increases conversions - **Runtime**: 7 days ⚠️ - **Sample Size**: Control 10,000 / Variant 10,000 ### Results | Metric | Control | Variant | Lift | Significant? | |--------|---------|---------|------|--------------| | Conversion Rate | 2.5% | 3.0% | +20% | ✅ Yes (95%) | ### Statistical Details - Z-score: 2.28 - P-value: 0.023 - Confidence Interval: +2.3% to +37.7% - Power achieved: 62% ⚠️ ### Validity Assessment - ❌ Minimum runtime NOT met (7 days < 14 days recommended) - ⚠️ Sample size marginal for 20% MDE - ❓ Cannot assess external confounds without more context ### Interpretation Result is **statistically significant** but validity concerns exist: 1. 7-day runtime may miss day-of-week patterns 2. Wide confidence interval (+2% to +38%) indicates uncertainty 3. Recommend extending test 7 more days to confirm ### "Why It Worked" Hypotheses 1. New headline more clearly communicates value proposition 2. Specificity/numbers in headline increased credibility ### Next Test Ideas 1. Test headline variations that emphasize the winning element 2. Apply same messaging pattern to subheadline ### Confidence Level: MEDIUM Statistical significance achieved, but short runtime reduces confidence. ``` ## Guidelines - Never declare a winner without checking validity - Distinguish statistical significance from practical significance - If test ran <7 days, always recommend extending - If sample size insufficient, calculate required runtime to reach it - Ask for segment data if not provided - results often differ by device/source ## Data access (Ryze MCP) This skill works best with live account data. Connect the free Ryze MCP once and Claude reads your Google Ads, Meta Ads, GA4 and Search Console directly: - claude.ai / Claude Desktop: Settings → Connectors → Add custom connector → `https://connector.get-ryze.ai/mcp` - Claude Code: `claude mcp add ryze --transport http https://connector.get-ryze.ai/mcp` - Cursor: Settings → MCP → add the same URL Setup guide: https://www.get-ryze.ai/how-to-connect-claude-to-google-meta-ads-mcp
[{"step":1,"action":"Gather your A/B test data from Google Ads or Meta Ads. Export the raw performance data including impressions, clicks, conversions, and cost metrics for both variations. Ensure you have the test duration and total visitors.","tip":"Use the platform's native export tools (e.g., Meta Ads Manager's 'Export' button or Google Ads' 'Reports' tab) to get clean CSV files. Include segment dimensions like device type, age, or geography if you plan to analyze them."},{"step":2,"action":"Identify your primary metric and secondary metrics. The primary metric should align with your campaign goal (e.g., conversion rate for a sales campaign, CTR for a brand awareness campaign). Secondary metrics help validate the results (e.g., cost per conversion, ROAS).","tip":"If your primary metric is conversion rate, ensure you're tracking conversions correctly in both variations. For e-commerce, this might mean setting up purchase events in Meta Pixel or Google Analytics 4."},{"step":3,"action":"Input the data into the A/B test analyzer tool. Use the prompt template to structure your request, filling in the placeholders with your specific campaign details, metrics, and segment criteria.","tip":"For platforms like Google Sheets or Excel, use the '=TTEST' function to manually calculate p-values if you prefer a quick check before diving deeper. Tools like Optimizely or VWO can automate this process for web experiments."},{"step":4,"action":"Review the statistical significance results, segment breakdowns, and recommendations. Pay attention to the p-value and effect size to determine if the results are reliable and actionable. Use the segment breakdown to identify high-performing or underperforming groups.","tip":"If the p-value is close to 0.05 (e.g., 0.048), consider running the test longer to confirm the results. Segment data can reveal insights that aren't visible in aggregate metrics, such as device-specific performance differences."},{"step":5,"action":"Implement the recommendations and plan your next test. Scale the winning variation if the results are statistically significant and the lift is meaningful. Use the suggested hypotheses to design your next experiment, ensuring you account for the sample size requirements to achieve reliable results.","tip":"Document your test results and hypotheses in a shared doc (e.g., Notion or Google Docs) to build a knowledge base for your team. This helps avoid repeating tests and accelerates optimization over time."}]
No install command available. Check the GitHub repository for manual installation instructions.
git clone https://github.com/irinabuht12-oss/marketing-skills/tree/main/skills/ab-test-analyzerCopy the install command above and run it in your terminal.
Launch Claude Code, Cursor, or your preferred AI coding agent.
Use the prompt template or examples below to test the skill.
Adapt the skill to your specific use case and workflow.
Analyze the A/B test results for [PLATFORM: Google Ads or Meta Ads] campaign [CAMPAIGN_NAME]. The test compared [VARIATION_A_DESCRIPTION] vs [VARIATION_B_DESCRIPTION] over [TEST_DURATION] days with [TOTAL_VISITORS] visitors. Calculate statistical significance for [PRIMARY_METRIC: e.g., CTR, conversion rate, ROAS], provide a segment breakdown by [SEGMENT_CRITERIA: e.g., device type, audience segment, geography], and recommend whether to scale the winning variation or iterate further. Include sample size requirements for future tests to achieve 95% confidence with 80% power. Suggest 2-3 follow-up hypotheses based on the results.
For a Meta Ads campaign testing a new ad creative (Variation A: video ad with testimonials vs Variation B: static carousel ad) over 14 days with 12,500 visitors, the A/B test results show: **Statistical Significance:** - Primary metric: Conversion rate (purchase) - Variation A: 3.2% (400 conversions) - Variation B: 2.8% (350 conversions) - p-value: 0.042 (statistically significant at 95% confidence) - Effect size: 14.3% lift in conversions for Variation A **Segment Breakdown:** - Mobile users: Variation A (+18% lift), Variation B (+11% lift) - Desktop users: Variation A (+8% lift), Variation B (+6% lift) - Lookalike audiences: Variation A (+22% lift), Variation B (+15% lift) - Top-performing segment: 18-24 age group (Variation A: +25% lift) **Recommendation:** Scale Variation A (video testimonial ad) as the winning variation, but pause the test immediately to avoid potential long-term harm from the losing variation. The lift is statistically significant but modest, suggesting further optimization potential. **Sample Size Requirements:** To detect a 10% lift in conversion rate with 95% confidence and 80% power for future tests: - Minimum sample size: 15,000 visitors per variation - Recommended duration: 21 days (assuming 700 visitors/day) **Follow-up Hypotheses:** 1. Test a hybrid creative combining the testimonial video with carousel elements to appeal to both mobile and desktop users. 2. Run a segment-specific test targeting the 18-24 age group with a UGC-style video ad to maximize lift in this high-performing segment. 3. Test a variation with a stronger CTA overlay on the video ad to see if it further improves conversion rates without increasing cost per acquisition.
skills-collection
Take a free 3-minute scan and get personalized AI skill recommendations.
Take free scan