By datathings
Run high-throughput LLM inference using vLLM — batch generation, chat completions, structured outputs, LoRA adapters, multimodal inputs, and OpenAI-compatible serving.
Own this plugin?
Verify ownership to unlock analytics, metadata editing, and a verified badge. GitHub access is read-only (username + org membership).
Sign in to claimOwn this plugin?
Verify ownership to unlock analytics, metadata editing, and a verified badge. GitHub access is read-only (username + org membership).
Sign in to claimBased on adoption, maintenance, documentation, and repository signals. Not a security audit or endorsement.
npx claudepluginhub datathings/marketplace --plugin vllmNVIDIA CUDA C/C++ skill - Runtime API, cuBLAS, cuFFT, cuSPARSE, cuRAND, cuSolver, Thrust, and Cooperative Groups for GPU-accelerated computing
AMD ROCm 7.2.4 GPU computing stack: HIP kernel development, rocBLAS/rocFFT/rocRAND/rocSOLVER compute libraries, profiling, and CUDA-to-HIP porting
OpenCL SDK (Khronos Group) v2026.05.29 (OpenCL 3.1) C/C++ skill — cross-platform GPU/CPU parallel computing with ~60 C API functions, C++ wrapper, and SDK utilities
Comprehensive reference for GreyCat C API and GCL Standard Library. Covers native function implementation, tensor operations, scheduling, I/O, statistics, and all std modules.
Run and manage local LLMs via Ollama's REST API, with support for text generation, chat, embeddings, and custom model creation.
Run and manage local LLMs via Ollama's REST API, with support for text generation, chat, embeddings, and custom model creation.
Deploy and benchmark vLLM with Claude Code
Agent Skills for Together AI platform — inference, training, embeddings, audio, video, images, function calling, and infrastructure
Inference-time scaling for LLMs — generate multiple candidates and select the best using voting, scoring, or search
Agent-ready playbooks for LLM serving benchmarks, capacity planning, torch-profiler triage, pipeline analysis, compute simulation, SGLang/vLLM SOTA Humanize loops, human code review, production incident triage, and model PR-history dossiers.
Ultra-compressed communication mode. Cuts 65% of output tokens (measured) while keeping full technical accuracy by speaking like a caveman.