Plugins listed here are tagged for this technology stack and auto-indexed from public GitHub repositories.
Plugins listed here are tagged for this technology stack and auto-indexed from public GitHub repositories.
Claude Code plugins tagged for Prometheus development. Browse commands, agents, skills, and more.
Profile, optimize, and monitor application performance across frontend and backend systems using OpenTelemetry, Prometheus, Grafana, and Datadog. Includes React/Next.js optimization, caching, load testing, and observability setup with SLI/SLO management.
Implement production observability with Prometheus metrics, Grafana RED/USE dashboards, distributed tracing via Jaeger/Tempo, and SLO/error-budget management. Includes agents for performance profiling, database tuning, and network optimization.
Automate database migrations with zero-downtime SQL scripts, observability via CDC pipelines and Grafana, and expert DBA agents for multi-cloud operations, performance tuning, and reliability engineering across PostgreSQL, MySQL, MongoDB, and more.
Set up, monitor, and optimize production systems with distributed tracing, SLI/SLO management, incident response, and performance profiling using Prometheus, Grafana, OpenTelemetry, and other observability tools.
Orchestrate production incident response with structured runbooks, multi-agent triage and debugging, blameless postmortems, and on-call shift handoffs. Automate detection, investigation, mitigation, and learning from incidents using SRE practices and observability tools.
Debug distributed microservices by tracing requests across services, analyzing logs, and correlating errors to identify root causes. Set up distributed tracing, diagnose production incidents, and optimize performance using Kubernetes, Docker, and observability tools.
Implement production-grade observability and monitoring with distributed tracing, SLI/SLO management, incident response, performance optimization, and postmortem writing.
Expert-level observability, incident response, and SRE capabilities for diagnosing production issues, implementing monitoring and tracing with Prometheus, Grafana, Datadog, and OpenTelemetry, conducting post-incident reviews, and managing SLIs/SLOs.
Accelerate cloud infrastructure management and DevOps operations with skills for AWS serverless, Kubernetes, Docker, Terraform, CI/CD with GitHub Actions, production deployment strategies, incident response, and observability using Prometheus, Grafana, and Datadog.
Orchestrate multi-agent teams for complex projects: decompose tasks, match agent capabilities, coordinate shared state, manage error handling, parallel execution, load balancing, and business process workflows with saga patterns and monitoring.
Configure Prometheus for comprehensive metric collection, monitoring, storage, alerting, and service discovery across infrastructure and applications, including scrapers, recording rules, and alerting rules setup
Configure Prometheus-based Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alert policies for SRE practices and reliability monitoring.
Idiomatic Go patterns and production tooling for the full development lifecycle: coding conventions, concurrency, performance optimization, testing, debugging, CI/CD, database access, GraphQL, gRPC, monitoring, and security auditing.
Create and manage production Grafana dashboards for visualizing system and application metrics using RED/USE methods and PromQL queries, with SLO tracking and infrastructure monitoring.
Set up Qdrant monitoring and observability with Prometheus, Grafana, and health checks. Debug production issues like optimizer stuck, memory growth, and slow requests.
Manage and optimize ClickHouse databases across the full lifecycle — schema design, data ingestion, query optimization, production deployment, monitoring, security, and cost management. Includes error diagnostics, CI setup, RBAC, migrations, and streaming ingestion.
Develop, deploy, and maintain Intercom integrations with skills for API development, CI/CD, monitoring, security, compliance, and migration from other platforms.
Manage and optimize GPU workloads on CoreWeave's Kubernetes cloud: deploy inference services, run distributed training, monitor GPU health, control costs, and secure multi-team access.
Collect infrastructure performance metrics across compute, storage, network, containers, load balancers, and databases using Prometheus, Datadog, or CloudWatch, and produce Grafana dashboard configurations and alert rules for monitoring and troubleshooting.
Integrate Mistral AI into your development and production workflows with skills for API setup, chat completions, embeddings, RAG, security best practices, CI/CD, deployment, observability, cost optimization, and migration from other LLM providers.
Detect, analyze, and prevent database deadlocks across PostgreSQL, MySQL, and MongoDB with lock queries, log parsing, code tracing, and monitoring scripts.
Deploy production-ready monitoring stacks (Prometheus, Grafana, Datadog) on Kubernetes or Docker with exporters, alerting rules, and pre-built dashboards, while also generating configuration code.
Build and operate Exa neural search integrations end-to-end: SDK setup, search execution, RAG pipelines, caching, rate limiting, CI/CD, monitoring, cost optimization, and production deployment.
Generate intelligent alerting rules for Prometheus, Grafana, Datadog, and PagerDuty with configurable thresholds, routing, escalation policies, and runbook generation for production systems.
Build, deploy, and maintain Apollo.io sales intelligence integrations with full lifecycle support: authentication, search/enrichment, email sequences, webhooks, rate limiting, caching, monitoring, CI/CD, and compliance.
Integrate and manage Cohere API v2 across the full development lifecycle—from SDK setup and local mocking to production deployment, monitoring, and incident response. Includes RAG pipelines, tool-using agents, cost optimization, security compliance, rate limiting, and migration from OpenAI/Anthropic or Cohere v1.
Centralize performance metrics from apps, systems, databases, caches, and queues into Prometheus, StatsD, or CloudWatch with unified naming, and generate instrumentation code, dashboards, and alert definitions.
Track and analyze response times across API endpoints, database queries, and service calls with percentile reporting and SLO compliance monitoring. Output Prometheus/Grafana dashboards and optimization strategies.
Manage Langfuse LLM observability across the full lifecycle: installation, configuration, tracing, evaluation, cost monitoring, deployment, scaling, and incident response. Supports OpenAI, LangChain, and other LLM frameworks with CI/CD integration and production-grade patterns.
Integrate, automate, and maintain Navan travel and expense management platform—from OAuth setup and API calls to data extraction, deployment, monitoring, and incident response.
Apply 28 prioritized best practice rules for ClickHouse schema design, query optimization, and data ingestion with supporting skills for deploying to ClickHouse Cloud, setting up local development environments, integrating Node.js and Python clients, and troubleshooting performance issues.
Set up comprehensive monitoring and observability for your applications, including APM instrumentation, custom metrics, alerting, centralized logging, distributed tracing, infrastructure monitoring, and dashboards using tools like New Relic, Datadog, Prometheus, Grafana, Elasticsearch, and AWS.
Manage Qdrant vector search deployments end-to-end: scaling, performance tuning, search relevance optimization, multitenancy, zero-downtime upgrades, embedding model migration, monitoring with Prometheus/Grafana, and Edge deployment for local vector search
Debug application issues, manage dashboards, SLOs, alerts, and synthetic monitoring checks in Grafana Cloud using the gcx CLI, with support for GitOps workflows, Go code generation, and observability-driven root cause analysis.
Load-test websites and APIs with k6, instrument apps with OpenTelemetry and eBPF, ship telemetry to Grafana Cloud, and troubleshoot performance, cardinality, and cost issues across Prometheus, Loki, Tempo, and Pyroscope.
Write, review, and debug Go services with patterns for API design, testing, concurrency, observability, and project structure. Covers gRPC, REST, CLI tools, dependency injection, database access, CI/CD, security audits, and performance optimization.
Delegate observability infrastructure: set up OpenTelemetry distributed tracing, Prometheus/Grafana monitoring dashboards, Datadog alerting, structured logging pipelines, and SLO incident runbooks
Audit observability posture across your services: scan for RED metrics, SLOs, alerts, runbooks, tracing, and structured logging. Get a coverage matrix and identify critical gaps to ensure monitoring is sufficient before launch.
Inventory observability tools across services, map metrics, tracing, logging, and alerting coverage, and highlight monitoring blind spots
Generates SLO-based alert rules with burn-rate thresholds and paired runbooks for Prometheus, Grafana, Datadog, and CloudWatch, outputting actual configs when asked to set up alerts, create runbooks, or define SLOs.
Diagnose production incidents end-to-end: automatically detect the environment, gather symptoms, read logs, check metrics, trace requests to find root cause, and propose a fix with rollback.
Treat vigil as a dedicated SRE engineer that instruments services with OpenTelemetry, configures SLO-based alerting and runbooks for Prometheus, Grafana, and Datadog, audits observability posture across your stack, and diagnoses production incidents by combining logs, metrics, and traces.
Orchestrates multi-step development workflows by decomposing tasks, assigning specialized agents (code review, QA, DevOps, documentation, architecture, dependency management), and executing them in parallel waves with plan-mode approval and post-implementation cleanup.
Manage site reliability with incident response workflows, Prometheus-based monitoring and alerting, and SRE templates for SLOs, error budgets, and distributed system resilience patterns.
Designs and audits service mesh deployments on Kubernetes (Istio/Linkerd) — mTLS policy, traffic management (canary, circuit breaker, retry), and observability integration with Prometheus/Grafana. Delegate mesh infrastructure decisions to an agent.
Review and scaffold Go code with expertise in web architecture, data persistence, concurrency, BubbleTea TUI, Wish SSH, and Prometheus instrumentation — enforcing idiomatic patterns, security, testing, and middleware practices.
Manage Cloud SQL for PostgreSQL databases end-to-end: provision instances, explore schemas, execute SQL, monitor performance with PromQL, audit health, manage backups and upgrades, optimize vector search, and monitor replication.
Design service mesh observability by generating Prometheus/Grafana dashboards and configuring Jaeger/Tempo for distributed tracing and golden signal metrics.
Define SLI/SLO/SLA targets, error budgets, and incident severity levels for your services. Delegate reliability planning and post-mortem analysis to an SRE expert that enforces monitoring best practices with Prometheus.
Set up enterprise-grade monitoring, observability, and alerting for B2B applications, including APM, distributed tracing, logging, metrics, and SLA compliance with Prometheus, Grafana, Datadog, and cloud-native infrastructure.
Manage the full lifecycle of AlloyDB for PostgreSQL on GCP: provision clusters and instances, create and secure users, explore schemas and run queries, monitor health and replication, and troubleshoot performance via Cloud Monitoring metrics.
Scan projects for vulnerabilities, audit dependencies, validate IaC (Docker, K8s, Terraform), optimize Kubernetes clusters, and deploy services to cloud (GCP/Azure) using the Syncable CLI.
Deploy specialized AI agents for system architecture, database schema design, performance engineering, OpenTelemetry-based monitoring, and systematic web research within your SDLC workflow.
Set up production monitoring and observability across Datadog, CloudWatch, Prometheus, and Grafana — configure dashboards, alerts, SLOs, and distributed tracing, then automate incident response with SRE best practices and compliance audits for SOC2, HIPAA, and GDPR.
Generates SQL queries, transforms data with jq or pandas, builds ETL pipelines (batch and streaming), and performs time series forecasting or anomaly detection from natural language requirements.
Analyze and optimize VictoriaMetrics by identifying high-cardinality metrics, unused time series, and slow query traces, then generating actionable relabeling rules and aggregation configs to reduce storage costs and improve performance.
Query the full VictoriaMetrics observability stack: run PromQL/MetricsQL queries against metrics, search logs with LogsQL, discover distributed traces via Jaeger API, and manage AlertManager alerts and silences.
Automate infrastructure, CI/CD, and deployment workflows across Kubernetes, Docker, and cloud platforms with staged rollouts, monitoring, and cost optimization
Automate the operations phase of the AI Development Lifecycle (AIDLC) on AWS with self-improving loops: autonomous canary deployments using Kubernetes/Argo, continuous agent evaluation for quality and cost, incident response from CloudWatch alarms, and budget governance—all with human approval at checkpoints.
Investigate traces, logs, and metrics across an OpenSearch and Prometheus observability stack using PPL and PromQL queries to correlate data, diagnose errors, monitor RED metrics, define SLOs, and troubleshoot GenAI agent performance.
Identify cost drivers in Grafana Cloud by analyzing Prometheus metric DPM rates with per-series and per-label breakdowns, using gcx for automatic stack discovery and environment setup.
Provision, manage, and monitor Cloud SQL for SQL Server instances on GCP, including backups, cloning, schema exploration, SQL query execution, and performance monitoring with PromQL metrics.
Monitors database transactions in real time, detecting long-running queries, lock contention, and rollback anomalies with diagnostic queries and alert configuration.
Instrument API endpoints, database queries, and external services to track response times, identify bottlenecks, monitor SLOs, and generate optimization recommendations with Prometheus/Grafana dashboards.
Deploy production-ready monitoring stacks (Prometheus, Grafana, Datadog) with best-practice configurations and CI/CD integration, plus generate secure DevOps setup code tailored to your infrastructure.
Detect, analyze, and prevent database deadlocks with automated monitoring, alerting, and resolution strategies including lock contention analysis, monitoring scripts, and dashboards.
Centralize performance metrics from applications, systems, databases, caches, and queues into a unified monitoring view with automated dashboards, alerts, and instrumentation code.
Collect infrastructure performance metrics across compute, storage, network, containers, load balancers, and databases, then configure agents and create dashboards with alert rules for health monitoring and capacity tracking.
Generate alerting rules for latency, error rate, throughput, resource usage, availability, and SLO violations with defined thresholds, routing, escalation policies, and runbooks for Prometheus and Datadog.
Monitors PostgreSQL and MySQL database health with real-time metrics, predictive alerts, and automated remediation to detect performance degradation, replication lag, and resource exhaustion before production impact.
Orchestrates full-stack feature development from architecture through deployment, with automated CI/CD pipelines, performance tuning, security auditing, and AI-powered test generation.
Provides expertise for building IoT systems: device lifecycle management, MQTT messaging, edge computing, digital twins, time-series databases, security for constrained devices, and cloud platform integration (AWS, Azure, GCP).
Run comprehensive Kubernetes cluster health diagnostics and operator-specific checks (ArgoCD, Cert-Manager, Crossplane, Prometheus) with scored reports and kubectl-assisted debugging
Analyze Kubernetes cluster resource efficiency across nodes, workloads, Karpenter, and costs with Prometheus integration. Produces reports with utilization stats, issues, and actionable recommendations, supporting deep analysis and comparisons.
Optimize CI/CD pipelines and infrastructure with automated analysis of pipeline configurations, parallel job optimization, Docker best practices, Kubernetes deployment strategies, and full-stack observability setup including Prometheus, Grafana, and Loki.
Orchestrates multi-agent development workflows across architecture, code generation, testing, CI/CD, security, and documentation using a constructor-pattern skill system with parallel agents and sleep-sync session persistence.
Deploy and manage OpenTelemetry Collector pipelines shipping to Coralogix, instrument applications with OTel SDKs, write and debug OTTL transformations, and resolve telemetry semantic issues across Kubernetes and cloud environments.
Query Prometheus and Loki billing metrics via Grafana API for active series, ingestion rates, storage usage, and observability costs. Scaffold and develop Grafana v12.x plugins (panels, data sources, apps, backends) with automated Docker dev environments and full lifecycle support.
Accelerate enterprise full-stack development with automated CI/CD, multi-cloud infrastructure as code, code quality enforcement, API and test generation, documentation from OpenAPI specs and Mermaid diagrams, and performance monitoring using Prometheus and Grafana.
Manage the full SDLC for Go API services: scaffold projects, generate code from PostgreSQL schemas, run Playwright browser tests, review GitLab MRs with multi-persona analysis, monitor production via Sentry/Prometheus/Grafana/Loki/Kibana, investigate incidents, and resolve YouTrack tasks with automated commits.
Respond to production incidents with structured runbooks, automated triage, and multi-agent debugging. Orchestrate incident response playbooks, perform deep root cause analysis via code tracing and git bisect, write blameless postmortems, and generate test suites for verified fixes.
Modern Python development with async patterns, FastAPI/Django scaffolding, production best practices, and tooling for testing, packaging, observability, and code quality
Run multi-layered code reviews combining static analysis tools with AI reasoning, and delegate performance engineering and test automation to specialized agents for load testing, observability, caching, and self-healing test generation.
Troubleshoot distributed systems by configuring debugging environments, diagnosing production incidents, and correlating errors across logs and codebases.
Generate and orchestrate CI/CD pipelines across GitHub Actions, GitLab CI, and Azure DevOps with multi-stage workflows, approval gates, security scans, and Kubernetes deployments. Includes secrets management, infrastructure-as-code automation, and production incident response.
Profile, optimize, and monitor application performance end-to-end, from React/Next.js frontend tuning to backend observability with Prometheus, Grafana, and Datadog. Includes load testing, distributed tracing, caching, and SLI/SLO management.
Set up production-grade observability with Prometheus metrics, Grafana dashboards, distributed tracing via Jaeger/Tempo, and SLO-based alerting for microservices and Kubernetes environments.
Implement SRE practices for production systems: set up monitoring alerts with Prometheus, define SLOs and error budgets, and follow incident response workflows with templates for severity levels, triage, and communication.
Set up and manage observability infrastructure with specialized agents for logging, monitoring, and distributed tracing, covering metrics, dashboards, alerts, and OpenTelemetry instrumentation.