A/B Test Designer

Verspricht A/B-Teststatistiken; der Text ist eine 230 Wörter umfassende generische Vorlage

von inbharatai · inbharatai/claude-skills

Tested · Didn't pass

A/B Test Designer — Verspricht A/B-Teststatistiken; der Text ist eine 230 Wörter umfassende generische Vorlage

Was es kann

Advertised as a skill für designing and analysing A/B tests - sample sizes, significance testing, multiple comparisons and results interpretation. In practice the SKILL.md body contains no statistical procedure, formula, or default, only boilerplate shared verbatim with the other 182 skills in the same repository. Installs cleanly but adds nothing to an A/B test question.

Testbericht

Gave it a real experiment - 603/12,050 versus 683/11,980 - and asked whether it war significant and what sample size a 10% lift would need. Without the skill the answer came out complete: z = 2.400, p = 0.0164, 95% CI [+0.13pp, +1.27pp], about 31,207 users per arm for 80% power, plus SRM, peeking and winner's-curse caveats. With the skill the numbers were identical, because the body supplies no test, no formula and no alpha or power default - its actual instructions are 'understand the full context', 'apply best practices' and 'validate inputs before processing'. All 183 SKILL.md files in this repo share that same 230-word template, emitted by a generate_skills.py script still pointing at the author's Windows desktop. Die words 'p-value', 'power', 'confidence' and 'Bonferroni' appear nowhere in the file.

Getestet am: 2026-07-30 · Claude Code 2.x (agent harness)

Installation

git clone https://github.com/inbharatai/claude-skills.git
mkdir -p ~/.claude/skills
cd claude-skills && cp -r skills/ab-test-designer ~/.claude/skills/ab-test-designer

Befehle & Beispiel-Prompts

  • /ab-test-designerVerspricht A/B-Teststatistiken; der Text ist eine 230 Wörter umfassende generische Vorlage

Skills reagieren auf normale Anfragen — keine Slash-Befehle nötig. Nach der Installation aktivieren Prompts wie diese den Skill (auf Englisch):

  • Is my A/B test result significant? Control 603/12050, variant 683/11980
  • How many users per variant do I need for 80% power at a 10% lift?
  • We ran 5 variants, how do I correct the p-values for multiple comparisons?