From rank-eval
Builds ranking evaluation workflows including NDCG/MRR metrics, human relevance labeling, and offline evaluation harnesses. Useful for search relevance and recommender system testing.
How this skill is triggered — by the user, by Claude, or both
Slash command
/rank-eval:rank-evalThe summary Claude sees in its skill listing — used to decide when to auto-load this skill
Build ranking evaluation — NDCG/MRR measurement, human relevance labeling, offline eval harness.
Build ranking evaluation — NDCG/MRR measurement, human relevance labeling, offline eval harness.
Rank — AI Ranking Engineer
Follow the output format defined in docs/output-kit.md.
2plugins reuse this skill
First indexed Jul 25, 2026
npx claudepluginhub tonone-ai/tonone --plugin rank-evalBuilds ranking evaluation workflows including NDCG/MRR metrics, human relevance labeling, and offline evaluation harnesses. Useful for search relevance and recommender system testing.
Audits ranking quality by analyzing metric trends, failure modes, dataset coverage, and reranker performance.
LLM-powered multi-attribute reranking of candidate sets from SQL or lists via pairwise comparisons on clarity, technical depth, insight. Supports custom prompts, model tiers, TopK.