From msp-ops-kit
Manages MSP proactive operations: patching, backup monitoring, alert triage, on-call rotation, change management, and vendor renewals. Triggered by maintenance window, backup, or change requests.
How this skill is triggered — by the user, by Claude, or both
Slash command
/msp-ops-kit:msp-maintenanceThe summary Claude sees in its skill listing — used to decide when to auto-load this skill
> **Defaults you must review.** The specific numbers in this skill are shipped example defaults
Defaults you must review. The specific numbers in this skill are shipped example defaults from a working MSP. Review and replace them with your own before anything goes client-facing.
This skill is the source of truth for the recurring work {{COMPANY_NAME}} does when nothing is broken: patching, backup verification, monitoring, on-call coverage, and controlled change. This is the substance behind the managed fee. The reactive desk (msp-helpdesk) is what clients see; this cycle is why they see it rarely.
One liability fact shapes the whole skill, inherited from msp-legal: the MSA's ransomware cost allocation carries a negligence carveout. {{COMPANY_NAME}}'s protection in a bad scenario is a dated log showing the routine ran: patches applied, backups verified, alerts handled. Every section below ends in a record because the record is part of the service.
Workstations: automated weekly patch cycle through the RMM. Reboots enforced monthly at minimum; a machine that has dodged its reboot past the deadline gets scheduled with the user, not skipped. Third-party application patching rides the same weekly cycle where the RMM covers the app.
Servers: monthly maintenance window, the third Thursday, 8:00 pm to midnight, {{TIMEZONE}}. Sequence per server: snapshot or verified backup point first, then patches, then reboot, then service verification against a per-server checklist (services up, backups scheduled, key app responds). Clients with affected services get the planned-maintenance notice (msp-client-comms template 1) at least 3 business days ahead.
Out-of-band: critical or actively-exploited vulnerabilities do not wait for the window. Test, then deploy as soon as reasonable, any day. If the fix is disruptive, use the emergency maintenance variant of the notice: one honest sentence on why it cannot wait.
Exceptions: a client or app that cannot tolerate a patch gets a documented exception with a review date, not a quiet skip. A client who refuses patching against recommendation is a Risk Acceptance Waiver conversation (msp-legal).
Record: patch status per endpoint lives in the RMM; the monthly window gets a closing note (what was patched, what was deferred and why). This feeds the QBR scorecard's Updates row.
Record: the restore log (date, client, scope, result, tech) is the single most important document in this skill. In a ransomware dispute it is the difference between a defensible posture and the negligence carveout.
The rule that keeps a small team sane: an alert is either actionable or it is tuned out of existence. An alert nobody would act on is noise, and noise trains the desk to ignore the one that matters.
For risky changes (firewall and network changes, DNS, server configuration, tenant-wide settings, anything that can take a client offline):
A change that skipped these steps and went fine is still a process failure; the one that goes wrong without them is an outage plus a liability problem.
The rule from the documentation-system school, adopted: update the doc in the same session as the change. Passwords rotated, configs changed, hardware swapped, vendor contacts updated: the client's documentation reflects it before the ticket or change closes. Stale documentation surfaces at the worst moments (an incident, an offboarding handoff) and both of those moments are already covered by promises this suite makes (msp-helpdesk, msp-offboarding).
The values below shipped as example defaults from a working MSP, and the items after them are still genuinely open questions upstream. Settle all of it for your own shop before this skill goes client-facing:
npx claudepluginhub rtfm-it-services-llc/msp-claude-skillsDefines MSP helpdesk operations: ticket priority, response targets, escalation, after-hours, and PSA workflow. Use when triaging tickets, drafting SLAs, or training new hires.
Automates PagerDuty incident management, alert inspection, and on-call operations via Composio's PagerDuty toolkit through Rube MCP.
Correlates data across PSA, RMM, documentation, and config monitoring tools during incident investigation. Produces a unified incident summary from vendor-agnostic workflows.