Plugins listed here are tagged for this technology stack and auto-indexed from public GitHub repositories.
Plugins listed here are tagged for this technology stack and auto-indexed from public GitHub repositories.
Claude Code plugins tagged for OpenTelemetry development. Browse commands, agents, skills, and more.
Profile, optimize, and monitor application performance across frontend and backend systems using OpenTelemetry, Prometheus, Grafana, and Datadog. Includes React/Next.js optimization, caching, load testing, and observability setup with SLI/SLO management.
Implement production observability with Prometheus metrics, Grafana RED/USE dashboards, distributed tracing via Jaeger/Tempo, and SLO/error-budget management. Includes agents for performance profiling, database tuning, and network optimization.
Debug production errors by correlating stack traces, logs, and distributed tracing data across systems, performing root cause analysis, and implementing fixes with automated observability steps. Also sets up structured logging, alerts, and error tracking integrations for real-time monitoring. Includes agents that isolate test failures and search logs for anomalies.
Automate performance analysis, test coverage review, and AI-powered code quality assessment across PRs and codebases using multi-agent reviews, static analysis tools, and observability agents for scaling, caching, and load testing.
Set up, monitor, and optimize production systems with distributed tracing, SLI/SLO management, incident response, and performance profiling using Prometheus, Grafana, OpenTelemetry, and other observability tools.
Deploy full-stack applications to Azure Container Apps with Bicep infrastructure, develop Azure Functions with production patterns, integrate Azure OpenAI and AI Search, instrument Node.js with OpenTelemetry monitoring, and authenticate using Azure Identity SDK.
Guides AI-assisted debugging from error triage through root cause analysis, collecting observability data and generating ranked hypotheses. Automates environment setup, optimizes workflows, and improves tooling and documentation for faster development.
Orchestrate production incident response with structured runbooks, multi-agent triage and debugging, blameless postmortems, and on-call shift handoffs. Automate detection, investigation, mitigation, and learning from incidents using SRE practices and observability tools.
Debug distributed microservices by tracing requests across services, analyzing logs, and correlating errors to identify root causes. Set up distributed tracing, diagnose production incidents, and optimize performance using Kubernetes, Docker, and observability tools.
Expert-level observability, incident response, and SRE capabilities for diagnosing production issues, implementing monitoring and tracing with Prometheus, Grafana, Datadog, and OpenTelemetry, conducting post-incident reviews, and managing SLIs/SLOs.
Implement production-grade observability and monitoring with distributed tracing, SLI/SLO management, incident response, performance optimization, and postmortem writing.
Deploy and manage Azure cloud applications with skills for Azure Container Apps deployment via Azure Developer CLI, Azure OpenAI integration for .NET, Azure Functions development patterns, Azure Identity authentication, Azure Monitor telemetry for Node.js, and Azure AI Search vector/hybrid search capabilities.
Add DeepEval evaluation loops, tracing, and dataset generation to AI applications, enabling span-by-span observability in Confident AI and iterative improvement on failures.
Idiomatic Go patterns and production tooling for the full development lifecycle: coding conventions, concurrency, performance optimization, testing, debugging, CI/CD, database access, GraphQL, gRPC, monitoring, and security auditing.
Build Python applications on Azure using SDK best practices across AI, storage, identity, messaging, monitoring, and management — including agents, content safety, translation, vision, speech, Cosmos DB, Event Hubs, Service Bus, Key Vault, and OpenTelemetry instrumentation.
Automates distributed tracing setup for microservices using OpenTelemetry, enabling end-to-end request visibility with context propagation, sampling, and deployment instructions for Jaeger, Zipkin, or Datadog APM.
Integrate and manage Cohere API v2 across the full development lifecycle—from SDK setup and local mocking to production deployment, monitoring, and incident response. Includes RAG pipelines, tool-using agents, cost optimization, security compliance, rate limiting, and migration from OpenAI/Anthropic or Cohere v1.
Build, deploy, and maintain Apollo.io sales intelligence integrations with full lifecycle support: authentication, search/enrichment, email sequences, webhooks, rate limiting, caching, monitoring, CI/CD, and compliance.
Provides Azure SDK patterns and best practices for Java developers to build applications across AI, communication, storage, identity, monitoring, and management services.
Debug production issues by instrumenting your project with OpenTelemetry to report to a Traceway instance, then use the Traceway CLI to query exceptions, logs, and metrics, diagnose performance bottlenecks, and trace root causes.
Build full-stack Azure applications from TypeScript/Node.js: AI services (content safety, document intelligence, translation, voice, search), storage (blobs, files, queues, Cosmos DB), messaging (Event Hubs, Service Bus, Web PubSub), identity, monitoring, and key management. Includes React UI components and Zustand state management patterns.
Manage the full Elastic stack lifecycle from Claude Code: provision and configure Elastic Cloud projects, query and analyze data with ES|QL, manage security (RBAC, audit logs, authentication), create Kibana dashboards and alerts, instrument applications with OpenTelemetry, triage security alerts, and investigate Kubernetes and LLM performance.
Debug application issues, manage dashboards, SLOs, alerts, and synthetic monitoring checks in Grafana Cloud using the gcx CLI, with support for GitOps workflows, Go code generation, and observability-driven root cause analysis.
Run a complete software development lifecycle with 18 specialized agents covering backend (Go/TypeScript), frontend (React/Next.js), DevOps, QA, and security — including gated development cycles, TDD, code review, accessibility/E2E/performance testing, dependency auditing, and infrastructure-as-code validation.
Manage LLM evaluation datasets on LangSmith, build and run LLM-as-judge evaluation pipelines, and enable distributed tracing for LLM applications using LangChain auto-tracing or OpenTelemetry.
Configure and deploy OpenTelemetry Collector pipelines, instrument applications with traces/metrics/logs, write OTTL transformations, and validate semantic conventions for observability.
Configure, instrument, and troubleshoot OpenTelemetry observability across multiple languages and the Collector pipeline, including SDK setup, YAML configuration, OTTL debugging, version lookups, and synthetic telemetry generation.
Provides a comprehensive .NET development workflow covering architecture design, code generation, testing, security review, performance analysis, deployment, and documentation for modern C#, ASP.NET, Blazor, MAUI, and cloud-native applications.
Write, review, and debug Go services with patterns for API design, testing, concurrency, observability, and project structure. Covers gRPC, REST, CLI tools, dependency injection, database access, CI/CD, security audits, and performance optimization.
Query, index, and manage Elasticsearch clusters and Kibana dashboards via curl REST API, including aggregations, ES|QL, ILM, OpenTelemetry patterns, and cluster health troubleshooting.
Instrument LLM applications with Arize AX observability: auto-instrument traces, manage datasets and experiments, run LLM-as-judge evaluators, optimize prompts from production data, and audit for regulatory compliance.
Delegate observability infrastructure: set up OpenTelemetry distributed tracing, Prometheus/Grafana monitoring dashboards, Datadog alerting, structured logging pipelines, and SLO incident runbooks
Audit observability posture across your services: scan for RED metrics, SLOs, alerts, runbooks, tracing, and structured logging. Get a coverage matrix and identify critical gaps to ensure monitoring is sufficient before launch.
Inventory observability tools across services, map metrics, tracing, logging, and alerting coverage, and highlight monitoring blind spots
Treat vigil as a dedicated SRE engineer that instruments services with OpenTelemetry, configures SLO-based alerting and runbooks for Prometheus, Grafana, and Datadog, audits observability posture across your stack, and diagnoses production incidents by combining logs, metrics, and traces.
Instrument any service with OpenTelemetry — add RED metrics, structured logging, distributed tracing, and health checks by generating code and configuration for Node.js, Python, and Go stacks.
Add Logfire observability to Python, JavaScript/TypeScript, and Rust applications with auto-instrumentation for FastAPI, httpx, asyncpg, SQLAlchemy, and more. Query traces, logs, and metrics via SQL, debug production issues, and start local dev sessions with temporary credentials.
Set up local-first LLM agent observability with GreptimeDB, including telemetry configuration for Claude Code via OTLP exporter. Query token usage, costs, traces, events, errors, and model comparisons through SQL.
Manage site reliability with incident response workflows, Prometheus-based monitoring and alerting, and SRE templates for SLOs, error budgets, and distributed system resilience patterns.
Design, instrument, and configure OpenTelemetry-based observability for distributed systems, with guidance on sampling, cardinality management, security, and OTTL transformations.
Autonomously design, configure, deploy, and troubleshoot production-grade AI agents on AWS using Bedrock AgentCore, Strands Agents SDK, with Terraform-first IaC and CloudWatch/OpenTelemetry observability.
Add full observability, evaluation, and safety to LLM/agent apps: trace all calls with OpenTelemetry, run automated evals (faithfulness, consistency, RAG quality), detect hallucinations, block prompt injection and PII leakage, optimize costs, and prevent regressions with production trace replay.
Provides a set of skills for code review, documentation, project design, and automation. Includes conventions for Go and Python, SQLite database guidance, OpenTelemetry instrumentation, and AI model fine-tuning with Unsloth.
Analyze AI agent execution traces to detect failures like grounding, goal drift, tool failures, and guardrail violations, then apply context remediation to fix system prompts and tool descriptions.
Guides Domain-Driven Design with Spring Boot 4, including bounded contexts, aggregates, REST APIs, Spring Data, Modulith modules, security, testing, observability, and upgrade verification from 3.x to 4.x.
Scaffold and wire the full Arbiter .NET ecosystem: mediator, CQRS with EF Core or MongoDB, REST endpoints, Blazor dispatcher, Azure Service Bus messaging, OpenTelemetry monitoring, and email/SMS communication — all with pre-built generic commands, queries, pipeline behaviors, and source-generated mappers.
Manage TrueFoundry AI Gateway for unified LLM access: configure model routing, guardrails, MCP servers, prompts, and observability; migrate codebases to use the gateway; verify credentials and diagnose issues.
Run battle-tested TDD workflows, systematic debugging, code review, and parallel task execution using reusable skill packs, agents, and commands.
Investigate traces, logs, and metrics across an OpenSearch and Prometheus observability stack using PPL and PromQL queries to correlate data, diagnose errors, monitor RED metrics, define SLOs, and troubleshoot GenAI agent performance.
Enables autonomous AI-driven development workflows within Claude Code, enforcing best practices across code quality, security, testing, documentation, and git operations. Coordinates multi-instance agents for issue resolution, code review, and epic management with crash-safe durability.
Run a multi-dimensional quality audit on your codebase: detect dead code, supply-chain risks, failure-mode gaps, telemetry issues, coupling problems, performance bottlenecks, and schema/API drift. Includes auto-cadence hooks and a measurement dashboard.
Run nine quality-canary audits covering code health, supply chain, resilience, telemetry, testability, performance, and contract drift. Each skill scans for specific issues (dead code, N+1 queries, missing observability, etc.) and reports findings with optional fixes. Includes a dashboard and self-update.
Query, instrument, and debug production systems using Honeycomb observability — from OpenTelemetry setup and migration to SLO management and root-cause analysis, including AI/LLM instrumentation.
Automates distributed tracing setup for microservices using OpenTelemetry, configuring context propagation, span creation, and trace collection with support for Jaeger, Zipkin, or Datadog backends.
Orchestrates full-stack feature development from architecture through deployment, with automated CI/CD pipelines, performance tuning, security auditing, and AI-powered test generation.
Provides a complete playbook of opinionated agent skills for enterprise software development, covering architecture decomposition, git-based risk analysis, observability with OpenTelemetry, security guardrails, contract-first API design, testing strategies, and code quality enforcement.
Orchestrates multi-agent development workflows across architecture, code generation, testing, CI/CD, security, and documentation using a constructor-pattern skill system with parallel agents and sleep-sync session persistence.
Deploy and manage OpenTelemetry Collector pipelines shipping to Coralogix, instrument applications with OTel SDKs, write and debug OTTL transformations, and resolve telemetry semantic issues across Kubernetes and cloud environments.
Administer Oracle Cloud Infrastructure end-to-end: manage IAM, networking, compute, OKE, databases, storage, security, cost, observability, Terraform, and disaster recovery with safety preflight checks, redaction, and risk-specific approvals.
Standardize OpenTelemetry instrumentation using semantic conventions for span naming, attributes, and status, and validate tracing with in-memory trace testing.
Configure OpenTelemetry environment variables and resource attributes for Claude Code sessions, enabling telemetry export to an OTel collector with optional disable support
Develop Effect-TS applications with skills covering core types, error handling, concurrency, streams, schema validation, dependency injection, testing, and observability. Includes commands for compliance checking and agents for code review and migration from imperative patterns.
Generate and orchestrate CI/CD pipelines across GitHub Actions, GitLab CI, and Azure DevOps with multi-stage workflows, approval gates, security scans, and Kubernetes deployments. Includes secrets management, infrastructure-as-code automation, and production incident response.
Analyze production errors from stack traces, logs, and distributed traces to perform root cause analysis, generate ranked hypotheses, and suggest fixes with observability tooling
Troubleshoot distributed systems by configuring debugging environments, diagnosing production incidents, and correlating errors across logs and codebases.
Respond to production incidents with structured runbooks, automated triage, and multi-agent debugging. Orchestrate incident response playbooks, perform deep root cause analysis via code tracing and git bisect, write blameless postmortems, and generate test suites for verified fixes.
Set up production-grade observability with Prometheus metrics, Grafana dashboards, distributed tracing via Jaeger/Tempo, and SLO-based alerting for microservices and Kubernetes environments.
Run multi-layered code reviews combining static analysis tools with AI reasoning, and delegate performance engineering and test automation to specialized agents for load testing, observability, caching, and self-healing test generation.
Profile, optimize, and monitor application performance end-to-end, from React/Next.js frontend tuning to backend observability with Prometheus, Grafana, and Datadog. Includes load testing, distributed tracing, caching, and SLI/SLO management.
Guide AI-assisted debugging from error triage through root cause analysis, leveraging observability data to generate ranked hypotheses. Automate environment setup and optimize workflows for faster, more enjoyable development.
Set up and manage observability infrastructure with specialized agents for logging, monitoring, and distributed tracing, covering metrics, dashboards, alerts, and OpenTelemetry instrumentation.
Implement SRE practices for production systems: set up monitoring alerts with Prometheus, define SLOs and error budgets, and follow incident response workflows with templates for severity levels, triage, and communication.