From o11y-dev-opentelemetry-skill
Guides OpenTelemetry collector configuration, pipeline design, production instrumentation, sampling, cardinality management, security, and OTTL transformations.
How this skill is triggered — by the user, by Claude, or both
Slash command
/o11y-dev-opentelemetry-skill:opentelemetry-skillThe summary Claude sees in its skill listing — used to decide when to auto-load this skill
Use these defaults:
Use these defaults:
Stability over Features: Check otelcol-contrib stability (Alpha/Beta/Stable) and warn on non-stable components in production.
Convention over Configuration: Prefer OpenTelemetry Semantic Conventions over custom attribute names.
Protocol Unification: Default to OTLP gRPC (4317); use OTLP HTTP (4318) when gRPC is blocked by the agent, proxy, browser, or backend.
Deterministic Routing Keys: Use stable routing keys for load-balancing exporters (traceID for tail sampling, tenant_id or cluster for tenant/shard routing). Normalize non-string attributes first.
Safety First: Prefer collector stability (memory limiters, persistent queues, backpressure) over completeness. Dropping data is better than crashing the collector.
Cardinality Awareness: High-cardinality attributes (>100 unique values) must not be metric dimensions; use traces or logs instead.
Security by Default: Redact PII, enable TLS for cross-network communication, and authenticate all collector endpoints.
Cross-Field Consistency: Treat collector reviews as systems reviews. Check processor order, memory limits, replica strategy, routing, metric temporality, queue storage, PDB/HPA settings, and OTTL types together before calling a config safe.
Before generating config/code, confirm these. If unknown, ask first:
file_storage + persistent queues. → collector.mdWhen user requests match these patterns, include these points explicitly:
memory_limiter; keep it first in each pipeline processors list; explain OOM-prevention rationale.user_id: refuse; explain time-series explosion risk; suggest traces and bounded metric dimensions.loadbalancing with routing_key: traceID, Headless Service (clusterIP: None), error+10% policies, Beta stability caution.CLAUDE_CODE_ENABLE_TELEMETRY=1, OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE=cumulative, ~/.claude/settings.json persistence; OTEL_LOG_USER_PROMPTS/OTEL_LOG_TOOL_DETAILS default to false — warn against enabling in shared/production environments without PII controls; avoid session.id as a metric dimension.gen_ai.* traces, name execute-tool spans with the actual tool name (for example bash or search_code) and preserve gen_ai.tool.name; do not use a generic execute_tool span name.When the user provides an existing collector config, Helm values file, or Kubernetes manifest, audit it for internal contradictions before proposing edits.
Check these interactions together:
limit_mib must leave headroom for the Go runtime and buffers.tail_sampling, spanmetrics, and servicegraph need sticky routing above one replica.hostPort fits DaemonSet/node-local patterns, not scaled gateway Deployments.replicaCount, HPA minReplicas, PodDisruptionBudget, and rolling updates together.deltatocumulative / cumulativetodelta need source/backend/restart checks.file_storage needs local locking-safe storage; RWX, EFS, and NFS are unsafe defaults.Load detailed reference documentation only when the user's request matches a trigger. This keeps context lean.
| Trigger keywords | Load | Key topics |
|---|---|---|
| Kubernetes, Helm, values.yaml, audit, review, DaemonSet, Sidecar, Gateway, Scaling, Load Balancing | architecture.md | DaemonSet vs Gateway vs Sidecar, Target Allocator, HPA, rollout consistency |
| Pipeline, Receiver, Processor, Exporter, Queue, Batch, Memory, Extensions, existing config | collector.md | Processor ordering, memory_limiter, file_storage, config audit heuristics, temporality/state audits, stability levels |
| SDK, Instrumentation, Spans, Attributes, Semantic Conventions, Cardinality | instrumentation.md | Auto vs manual, SemConv, cardinality Rule of 100 |
| Sampling, Cost, Volume, Head Sampling, Tail Sampling, Probabilistic | sampling.md | Head/tail sampling, sticky sessions, sampling math |
| Security, PII, GDPR, Redaction, TLS, Authentication, Credentials | security.md | PII redaction, mTLS, RBAC, extension exposure risks |
| Monitor the collector, Health, Alerts, Self-monitoring, Collector metrics | monitoring.md | otelcol_* metrics, dashboards, alert rules |
| Lambda, Azure Functions, GCP Functions, Serverless, FaaS, Mobile, Browser | platforms.md | FaaS patterns, Lambda extension layer, client-side apps |
| OTTL, Transform, Transformation, Modify, Filter attributes, Parse, Extract | ottl.md | OTTL syntax, context types, built-in functions, error handling |
| Connector, spanmetrics, servicegraph, routing connector, failover connector | connectors.md | R.E.D. metrics, service graph, routing, failover, stickiness |
| Claude Code, Codex, Gemini CLI, Copilot, AI agent, coding agent, MCP | ai-agents.md | Agent OTel support matrix, unified collector config, GenAI SemConv |
| validate, dry-run, startup error, pipeline error, dropped data, queue full, recovery | validation.md | Config validation commands, live checks, symptom→cause→fix recovery guidance |
| playbook, production playbook, blog, 2025 blog, 2026 blog, real world | playbooks.md | Production patterns from opentelemetry.io blogs |
| anti-pattern, common mistake, what to avoid, pitfall | anti-patterns.md | Full annotated anti-pattern catalogue: pipeline, metrics, Kubernetes, AI agents, OTTL |
Use this copy-paste-ready baseline unless the user wants something different:
extensions:
health_check:
endpoint: "0.0.0.0:13133"
file_storage/queue:
directory: /var/lib/otelcol/queue
timeout: 10s
compaction:
on_start: true
on_rebound: false
receivers:
otlp:
protocols:
grpc:
endpoint: "0.0.0.0:4317"
http:
endpoint: "0.0.0.0:4318"
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
batch:
timeout: 10s
send_batch_size: 1024
exporters:
otlp:
endpoint: "your-backend:4317"
sending_queue:
enabled: true
storage: file_storage/queue
num_consumers: 4
queue_size: 1024
retry_on_failure:
enabled: true
initial_interval: 1s
max_interval: 30s
max_elapsed_time: 300s
# otlphttp: # HTTP exporter — use when backend requires HTTP
# endpoint: "https://your-backend:4318"
# sending_queue: { enabled: true, storage: file_storage/queue }
service:
extensions: [health_check, file_storage/queue]
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlp]
metrics:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlp]
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlp]
Key defaults:
memory_limiter must be first in every processor chain.batch reduces exporter network calls.file_storage preserves queues across restarts only on the same host/volume. In Kubernetes, back /var/lib/otelcol/queue with a ReadWriteOnce block-backed PVC, not RWX/network storage.health_check binds to localhost (not 0.0.0.0) in shared networks.Always include at least one validation checkpoint: otelcol validate --config <path>.
For container dry-run commands, live pipeline checks, and symptom→cause→fix recovery, load validation.md.
❌ Placing memory_limiter anywhere except first in the processor chain
❌ Using high-cardinality attributes (user_id, trace_id) as metric dimensions
❌ Exposing pprof (1777), zpages (55679) on 0.0.0.0 in production
❌ Using tail_sampling without sticky session load balancing (loadbalancing exporter)
❌ Omitting batch processor (causes excessive network calls)
❌ Calling a config "fine" because it parses, without checking memory limits, sticky routing, exporter durability, and rollout settings together
See anti-patterns.md for the full annotated catalog.
SKILL.md focused on routing logic, guardrails, and production defaults rather than inline release tracking.npx claudepluginhub joshuarweaver/cascade-code-devops-misc-2 --plugin o11y-dev-opentelemetry-skillProvides deep operational guidance for OpenTelemetry SDK and Collector tuning, including propagator selection, sampling design, semantic-convention migration, and cardinality control.
Configures and deploys the OpenTelemetry Collector: receivers, exporters, processors, pipelines, sampling, RED metrics, and custom distributions. Use for agent/gateway patterns, Kubernetes (operator, Helm chart, raw manifests), and Dash0 forwarding.
Reviews OpenTelemetry Collector configurations managed by the Operator for pipeline correctness, deployment mode, memory safety, sampling integrity, and auto-instrumentation coverage.