By JiusiServe
Reproduce Long Video Sparse Attention (LVSA) paper headline results including SotA grid comparisons and latency scaling using bundled benchmarks, VQeval, and VBench-Long scoring, with figure regeneration.
Based on adoption, maintenance, documentation, and repository signals. Not a security audit or endorsement.
npx claudepluginhub jiusiserve/longvideosparseattention --plugin lvsa-reproduce-paperConfigure and run the LVSA vllm-omni plugin: env vars, geometry overrides, hooks for Wan vs HunyuanVideo, Cosmos 3.0 backend seam-swap (LVSA_COSMOS3_BACKEND — engages sparse single-GPU and under Ulysses SP), multi-GPU (tensor-parallel and Ulysses engage LVSA; Ring SP falls back to dense).
Diagnose LVSA failure modes: silent dense fallback, OOM at long sequences, missing output mp4 in Docker, quality regression at training reference.
Install LVSA and generate your first long video. Picks SDPA vs FlashInfer, sets the right LVSA_REFERENCE_LATENT_FRAMES per model (Wan/HunyuanVideo/Cosmos 3.0/CogVideoX), verifies the sparse path engaged.
Add LVSA support for a new video diffusion model by implementing the ModelAdapter ABC and wiring it into examples/.
Tune LVSA's sparsity_scale, window_size, n_first_frames, and reference_latent_frames for a target model and quality/speed trade-off.
Install LVSA and generate your first long video. Picks SDPA vs FlashInfer, sets the right LVSA_REFERENCE_LATENT_FRAMES per model (Wan/HunyuanVideo/Cosmos 3.0/CogVideoX), verifies the sparse path engaged.
Skills for finding, comparing, running, and prompting AI models on Replicate
vLLM v0.19.0 skill: offline batch inference, OpenAI-compatible server, LoRA adapters, multimodal inputs, embeddings, classification, structured outputs, and tool calling.
Agent Skills for Together AI platform — inference, training, embeddings, audio, video, images, function calling, and infrastructure
Agent-ready playbooks for LLM serving benchmarks, capacity planning, torch-profiler triage, pipeline analysis, compute simulation, SGLang/vLLM SOTA Humanize loops, human code review, production incident triage, and model PR-history dossiers.
Video generation at scale. Generate videos, images, and audio with Runway's API — batch ad campaigns, product videos, multishot stories, and creative iteration. Supports seedance2, gen4.5, veo3, Nano, Banana Pro, and more.