Slash Command

/eval

Manages eval-driven development: define feature eval criteria, check pass/fail status, generate reports, list all evals, and clean logs.

testing

code-quality

From muggle-ai-teams

Install

Run in your terminal

npx claudepluginhub multiplex-ai/muggle-ai-teams --plugin muggle-ai-teams

Command Content

Other plugins with /eval

/eval

Delegates to eval-harness skill for evaluation lifecycle: define, check, report, list, and cleanup. Legacy shim; prefer skill directly.

everything-claude-code

142.9k

/eval

Evaluates a plugin or skill directory for quality via static analysis and optional LLM judging, producing overall score with badge, dimension breakdowns, anti-patterns, and recommendations.

plugin-eval

32.9k

/eval

Manages eval-driven development: define feature eval criteria, check pass/fail status, generate reports, list all evals, and clean logs.

claude-forge

605

/eval

Manages eval-driven development workflow: define feature eval specs, check pass/fail status, generate reports, list definitions.

everything-claude-code

/eval

Evaluates implemented code for quality (SOLID/DRY), architecture, test coverage, performance, security; suggests iterative improvements via Code Reviewer Agent.

codingbuddy

/eval

debug/

Debugs and evaluates performance issues in the codebase with agent-orchestrated diagnostics, fixes, and metrics like QIS, TES, and Performance Index. Supports <target> and flags like --dry-run, --report-only.

autonomous-agent

Stats

Stars1

Forks0

Last CommitMar 20, 2026

Actions

View Source View Plugin View on GitHub View README

Eval Command

Manage eval-driven development workflow.

Usage

/eval [define|check|report|list] [feature-name]

Define Evals

/eval define feature-name

Create a new eval definition:

Create .claude/evals/feature-name.md with template:

## EVAL: feature-name
Created: $(date)

### Capability Evals
- [ ] [Description of capability 1]
- [ ] [Description of capability 2]

### Regression Evals
- [ ] [Existing behavior 1 still works]
- [ ] [Existing behavior 2 still works]

### Success Criteria
- pass@3 > 90% for capability evals
- pass^3 = 100% for regression evals

Prompt user to fill in specific criteria

Check Evals

/eval check feature-name

Run evals for a feature:

Read eval definition from .claude/evals/feature-name.md
For each capability eval:
- Attempt to verify criterion
- Record PASS/FAIL
- Log attempt in .claude/evals/feature-name.log
For each regression eval:
- Run relevant tests
- Compare against baseline
- Record PASS/FAIL
Report current status:

EVAL CHECK: feature-name
========================
Capability: X/Y passing
Regression: X/Y passing
Status: IN PROGRESS / READY

Report Evals

/eval report feature-name

Generate comprehensive eval report:

EVAL REPORT: feature-name
=========================
Generated: $(date)

CAPABILITY EVALS
----------------
[eval-1]: PASS (pass@1)
[eval-2]: PASS (pass@2) - required retry
[eval-3]: FAIL - see notes

REGRESSION EVALS
----------------
[test-1]: PASS
[test-2]: PASS
[test-3]: PASS

METRICS
-------
Capability pass@1: 67%
Capability pass@3: 100%
Regression pass^3: 100%

NOTES
-----
[Any issues, edge cases, or observations]

RECOMMENDATION
--------------
[SHIP / NEEDS WORK / BLOCKED]

List Evals

/eval list

Show all eval definitions:

EVAL DEFINITIONS
================
feature-auth      [3/5 passing] IN PROGRESS
feature-search    [5/5 passing] READY
feature-export    [0/4 passing] NOT STARTED

Arguments

$ARGUMENTS:

define <name> - Create new eval definition
check <name> - Run and check evals
report <name> - Generate full report
list - Show all evals
clean - Remove old eval logs (keeps last 10 runs)