Plugin

mlflow

Name: mlflow
Author: mlflow

Instrument Python and TypeScript code in LLM apps and agents with MLflow tracing for observability, analyze traces and multi-turn sessions to debug issues, evaluate outputs using datasets and judges to optimize accuracy and reduce costs, query aggregated metrics, and iterate improvements.

npx claudepluginhub mlflow/skills

Component Overview

Skills

Component Details

Skills (9)

mlflow-agent

/mlflow-agent

Master dispatcher for all MLflow workflows. Use this skill when the user wants to do anything with MLflow — tracing, evaluating, debugging, or improving an agent. Routes to the right MLflow sub-skill automatically. Triggers on: "use mlflow", "help with mlflow", "mlflow agent", "add mlflow to my project", "trace my agent", "evaluate my agent", or any MLflow task without a specific skill in mind.

agent-evaluation

/agent-evaluation

Use this when you need to EVALUATE OR IMPROVE or OPTIMIZE an existing LLM agent's output quality - including improving tool selection accuracy, answer quality, reducing costs, or fixing issues where the agent gives wrong/incomplete responses. Evaluates agents systematically using MLflow evaluation with datasets, scorers, and tracing. IMPORTANT - Always also load the instrumenting-with-mlflow-tracing skill before starting any work. Covers end-to-end evaluation workflow or individual components (tracing setup, dataset creation, scorer definition, evaluation execution).

instrumenting-with-mlflow-tracing

/instrumenting-with-mlflow-tracing

Instruments Python and TypeScript code with MLflow Tracing for observability. Must be loaded when setting up tracing as part of any workflow including agent evaluation. Triggers on adding tracing, instrumenting agents/LLM apps, getting started with MLflow tracing, tracing specific frameworks (LangGraph, LangChain, OpenAI, DSPy, CrewAI, AutoGen), or when another skill references tracing setup. Examples - "How do I add tracing?", "Instrument my agent", "Trace my LangChain app", "Set up tracing for evaluation"

analyzing-mlflow-trace

/analyze-mlflow-trace

Analyzes a single MLflow trace to answer a user query about it. Use when the user provides a trace ID and asks to debug, investigate, find issues, root-cause errors, understand behavior, or analyze quality. Triggers on "analyze this trace", "what went wrong with this trace", "debug trace", "investigate trace", "why did this trace fail", "root cause this trace".

analyzing-mlflow-session

/analyze-mlflow-chat-session

Analyzes an MLflow session — a sequence of traces from a multi-turn chat conversation or interaction. Use when the user asks to debug a chat conversation, review session or chat history, find where a multi-turn chat went wrong, or analyze patterns across turns. Triggers on "analyze this session", "what happened in this conversation", "debug session", "review chat history", "where did this chat go wrong", "session traces", "analyze chat", "debug this chat".

retrieving-mlflow-traces

/retrieving-mlflow-traces

Retrieves MLflow traces using CLI or Python API. Use when the user asks to get a trace by ID, find traces, filter traces by status/tags/metadata/execution time, query traces, or debug failed traces. Triggers on "get trace", "search traces", "find failed traces", "filter traces by", "traces slower than", "query MLflow traces".

querying-mlflow-metrics

/querying-mlflow-metrics

Fetches aggregated trace metrics (token usage, latency, trace counts, quality evaluations) from MLflow tracking servers. Triggers on requests to show metrics, analyze token usage, view LLM costs, check usage trends, or query trace statistics.

mlflow-onboarding

/mlflow-onboarding

Onboards users to MLflow by determining their use case (GenAI agents/apps or traditional ML/deep learning) and guiding them through relevant quickstart tutorials and initial integration. If an experiment ID is available, it should be supplied as input to help determine the use case. Use when the user asks to get started with MLflow, set up tracking, add observability, or integrate MLflow into their project. Triggers on "get started with MLflow", "set up MLflow", "onboard to MLflow", "add MLflow to my project", "how do I use MLflow".

searching-mlflow-docs

/searching-mlflow-docs

Searches and retrieves MLflow documentation from the official docs site. Use when the user asks about MLflow features, APIs, integrations (LangGraph, LangChain, OpenAI, etc.), tracing, tracking, or requests to look up MLflow documentation. Triggers on "how do I use MLflow with X", "find MLflow docs for Y", "MLflow API for Z".

README

MLflow Skills

MLflow Skills for Coding Agents

Turn your favorite coding agent into an LLMOps expert with MLflow skills.

Build, debug, and evaluate GenAI applications with confidence. These skills give your AI coding assistant deep knowledge of MLflow's tracing, evaluation, and observability features.

Works with any coding agent that support Skills, including Claude Code, Cursor, Codex CLI, Gemini CLI, and OpenCode.

Why MLflow Skills?

Building production-ready AI agents is hard. You need observability to understand what your agent is doing, evaluation to measure quality, and debugging tools when things go wrong. MLflow provides SDKs and best practices for all of these operations, and with skills we bring them directly into the environment where LLM agent development happens. Now you can go to your favorite coding agent and just ask:

"Add tracing to my LangChain app" → Instruments your code automatically
"Why did this trace fail?" → Analyzes spans, finds root causes, suggests fixes
"Evaluate my agent's accuracy" → Sets up datasets, scorers, and runs evaluation
"Improve my agent and verify your work" → Gives the coding agent reproducible and verifiable mechanism with MLflow eval to hill climb on quality
"Show me token usage trends" → Queries metrics and analyze trends

Available Skills

Observability & Debugging

Skill	Description
instrumenting-with-mlflow-tracing	Instruments Python and TypeScript code with MLflow Tracing. Supports OpenAI, Anthropic, LangChain, LangGraph, LiteLLM, and more.
analyze-mlflow-trace	Debugs issues by examining spans, assessments, and correlating with your codebase.
analyze-mlflow-chat-session	Debugs multi-turn chat conversations by reconstructing session history and finding where things went wrong.
retrieving-mlflow-traces	Powerful trace search and filtering by status, session, user, time range, or custom metadata.

Evaluation & Metrics

Skill	Description
agent-evaluation	End-to-end agent evaluation workflow — dataset creation, scorer selection, evaluation execution, and results analysis.
querying-mlflow-metrics	Fetches aggregated metrics (token usage, latency, error rates) with time-series analysis and dimensional breakdowns.

Helping New Users

Skill	Description
mlflow-onboarding	Guides new users through MLflow setup based on their use case (GenAI apps vs traditional ML).
searching-mlflow-docs	Searches official MLflow documentation efficiently using the llms.txt index.

Installation

Using `skills` installer

npx skills add mlflow/skills

Direct Installation from Source

git clone https://github.com/mlflow/mlflow-skills.git
cp -r mlflow-skills/* ~/.claude/skills/

Change the ~/.claude/skills/ directory to the appropriate location for your coding agent, e.g., ~/.codex/skills/ for Codex.

Project-Level Installation

Add skills to your project for team sharing:

cd your-project
git clone https://github.com/mlflow/mlflow-skills.git .skills/mlflow
# Or as a submodule:
git submodule add https://github.com/mlflow/mlflow-skills.git .skills/mlflow

Auto-Suggestion Hook (Optional)

The hooks/ directory contains a UserPromptSubmit hook that automatically detects MLflow-related patterns in your prompts and surfaces the right skill before the agent responds — no need to remember which skill does what.

Install the Hook

Step 1: Copy the hook somewhere permanent:

cp hooks/mlflow-suggest-hook.py ~/.claude/hooks/mlflow-suggest-hook.py

Step 2: Add it to your Claude Code settings (~/.claude/settings.json):

{
  "hooks": {
    "UserPromptSubmit": [
      {
        "type": "command",
        "command": "python3 ~/.claude/hooks/mlflow-suggest-hook.py"
      }
    ]
  }
}

Step 3: Start a new session. When you ask something like "Add tracing to my app", you'll see:

💡 Use the `instrumenting-with-mlflow-tracing` skill to add MLflow tracing.

See hooks/README.md for the full keyword-to-skill mapping and troubleshooting.

Quick Examples

Instrument Your App with Tracing

> Add MLflow tracing to my OpenAI app

The coding agent will:
1. Detect your LLM framework
2. Add the right autolog call
3. Configure experiment tracking
4. Verify traces are being captured

Debug a Failed Trace

> Analyze trace tr-abc123 — why did it return the wrong answer?

View full README on GitHub

Similar Plugins

team-skills-platform

167.4k

1.5K

Team-oriented workflow plugin with role agents, 27 specialist agents, ECC-inspired commands, layered rules, and hooks skeleton.

Stats

Version1.0.0

Stars20

Forks8

MaintenanceExcellent

AddedMar 31, 2026

Actions

View on GitHub View README Plugin Marketplace JSON

MLflow Skills for Coding Agents

Turn your favorite coding agent into an LLMOps expert with MLflow skills.

Build, debug, and evaluate GenAI applications with confidence. These skills give your AI coding assistant deep knowledge of MLflow's tracing, evaluation, and observability features.

Works with any coding agent that support Skills, including Claude Code, Cursor, Codex CLI, Gemini CLI, and OpenCode.

Why MLflow Skills?

"Add tracing to my LangChain app" → Instruments your code automatically
"Why did this trace fail?" → Analyzes spans, finds root causes, suggests fixes
"Evaluate my agent's accuracy" → Sets up datasets, scorers, and runs evaluation
"Improve my agent and verify your work" → Gives the coding agent reproducible and verifiable mechanism with MLflow eval to hill climb on quality
"Show me token usage trends" → Queries metrics and analyze trends

Available Skills

Observability & Debugging

Skill	Description
instrumenting-with-mlflow-tracing	Instruments Python and TypeScript code with MLflow Tracing. Supports OpenAI, Anthropic, LangChain, LangGraph, LiteLLM, and more.
analyze-mlflow-trace	Debugs issues by examining spans, assessments, and correlating with your codebase.
analyze-mlflow-chat-session	Debugs multi-turn chat conversations by reconstructing session history and finding where things went wrong.
retrieving-mlflow-traces	Powerful trace search and filtering by status, session, user, time range, or custom metadata.

Evaluation & Metrics

Skill	Description
agent-evaluation	End-to-end agent evaluation workflow — dataset creation, scorer selection, evaluation execution, and results analysis.
querying-mlflow-metrics	Fetches aggregated metrics (token usage, latency, error rates) with time-series analysis and dimensional breakdowns.

Helping New Users

Skill	Description
mlflow-onboarding	Guides new users through MLflow setup based on their use case (GenAI apps vs traditional ML).
searching-mlflow-docs	Searches official MLflow documentation efficiently using the llms.txt index.

Installation

Using `skills` installer

npx skills add mlflow/skills

Direct Installation from Source

git clone https://github.com/mlflow/mlflow-skills.git
cp -r mlflow-skills/* ~/.claude/skills/

Change the ~/.claude/skills/ directory to the appropriate location for your coding agent, e.g., ~/.codex/skills/ for Codex.

Project-Level Installation

Add skills to your project for team sharing:

cd your-project
git clone https://github.com/mlflow/mlflow-skills.git .skills/mlflow
# Or as a submodule:
git submodule add https://github.com/mlflow/mlflow-skills.git .skills/mlflow

Auto-Suggestion Hook (Optional)

Install the Hook

Step 1: Copy the hook somewhere permanent:

cp hooks/mlflow-suggest-hook.py ~/.claude/hooks/mlflow-suggest-hook.py

Step 2: Add it to your Claude Code settings (~/.claude/settings.json):

{
  "hooks": {
    "UserPromptSubmit": [
      {
        "type": "command",
        "command": "python3 ~/.claude/hooks/mlflow-suggest-hook.py"
      }
    ]
  }
}

Step 3: Start a new session. When you ask something like "Add tracing to my app", you'll see:

💡 Use the `instrumenting-with-mlflow-tracing` skill to add MLflow tracing.

See hooks/README.md for the full keyword-to-skill mapping and troubleshooting.

Quick Examples

Instrument Your App with Tracing

> Add MLflow tracing to my OpenAI app

The coding agent will:
1. Detect your LLM framework
2. Add the right autolog call
3. Configure experiment tracking
4. Verify traces are being captured

Debug a Failed Trace

> Analyze trace tr-abc123 — why did it return the wrong answer?

mlflow

Component Overview

Component Details

Skills (9)

README

MLflow Skills for Coding Agents

Why MLflow Skills?

Available Skills

Observability & Debugging

Evaluation & Metrics

Helping New Users

Installation

Using skills installer

Direct Installation from Source

Project-Level Installation

Auto-Suggestion Hook (Optional)

Install the Hook

Quick Examples

Instrument Your App with Tracing

Debug a Failed Trace

Similar Plugins

team-skills-platform

mlflow

Component Overview

Component Details

Skills (9)

README

MLflow Skills for Coding Agents

Why MLflow Skills?

Available Skills

Observability & Debugging

Evaluation & Metrics

Helping New Users

Installation

Using skills installer

Direct Installation from Source

Project-Level Installation

Auto-Suggestion Hook (Optional)

Install the Hook

Quick Examples

Instrument Your App with Tracing

Debug a Failed Trace

Similar Plugins

team-skills-platform

claude-buddy

episodic-memory

Using `skills` installer

Using `skills` installer