From evals
Analyzes A/B test results with statistical significance, practical significance, and segmentation analysis. Produces ship/no-ship recommendations with statistical justification.
How this skill is triggered — by the user, by Claude, or both
Slash command
/evals:eval-analyzeThis skill is limited to the following tools:
The summary Claude sees in its skill listing — used to decide when to auto-load this skill
You are Eval — Experiment Design Engineer on the Data Science Team.
You are Eval — Experiment Design Engineer on the Data Science Team.
Ask the user for any missing context needed to produce a useful output. If the request is clear, skip questions and proceed.
Gather experiment results (control/treatment metrics, sample sizes), primary metric, and any planned segments.
Output an analysis report: test statistic, p-value, confidence interval, practical significance assessment, segment analysis, and ship/no-ship recommendation.
Output a brief summary:
3plugins reuse this skill
First indexed Jul 25, 2026
npx claudepluginhub tonone-ai/tonone --plugin evalsGuides collaborative design exploration before implementation: explores context, asks clarifying questions, proposes approaches, and writes a design doc for user approval.
Creates structured, bite-sized implementation plans from specs or requirements before writing code. Useful for breaking down multi-step tasks into testable steps with file structure and task boundaries.
Resolves in-progress git merge or rebase conflicts by analyzing history, understanding intent, and preserving both changes where possible. Runs automated checks after resolution.