If you used WarpFusion back in the days and have missed on the VibeWarp -
github.com/Sxela/VibeWarp
WarpFusion, unpacked from the notebook
VibeWarp is WarpFusion v0.37 consolidated into an installable Python package with a local web UI — no Colab, no runtime git clone, every dependency vendored in or pinned.
It feeds each frame of a video through Stable Diffusion and other image edit models, warping the previous render forward along optical flow so every frame builds on the last instead of being redrawn from scratch. The result is painterly and alive — it drifts, dissolves and reconstitutes.
That texture is the point, not an artefact of it. A video model gives you a clean, coherent
clip; this gives you something that looks made by hand.
This is a cumulative overview of everything in v0.7.x. If you're arriving now, this is the
whole surface.
The render engine
Stable Diffusion 1.5 and SDXL, with the full WarpFusion loop: optical-flow warping, consistency masks, reconstructed-noise mode, and per-frame scheduling of essentially every knob.
- ControlNet — the full SD1.5 and SDXL stacks, multi-net with per-net weights and scheduling. Annotators are built in: depth (Midas / Zoe / Depth-Anything), softedge, scribble, canny, MLSD, normalbae, segmentation, lineart and lineart-anime, shuffle, tile, ip2p, temporalnet, inpaint, and OpenPose / DWPose with per-part body, hand and face detection.
- AnimateDiff — SD1.5 v1/v2/v3 plus SDXL / HotshotXL motion modules, sliding context windows, seam handling and stylized reinjection. The motion module is fully vendored.
- LORA — A1111-format <lora:name:weight> parsing in prompts, plus per-frame weight scheduling.
- IP-Adapter — a canonical model catalog driving a dedicated editor, with the right image encoder (ViT-H or bigG) picked from the adapter itself. Multiple adapters at once, per-frame reference sources, and the full ComfyUI layer-weight preset set including style transfer, composition and the precise variants.
- Gradient guidance — pixel-space (LPIPS + MSE) and latent-space guidance steering sampling toward a chosen temporal target, RMS-clamped, with the target optionally noised to the current sigma.
- Consistency masks — per-component weights, blur, dilate, softening, all schedulable.
- Prompts — per-frame keyframes, multi-prompt blending with weights ( a:0.7 | b:0.3 ), and {caption} interpolation from BLIP frame captioning.
- Content-aware scheduling — scene-change detection (RMSE / LPIPS) driving steps, CFG, style strength and flow blend.
- Plus FreeU, softcap, background masking, tiled VAE, colormatch, deflicker, RealESRGAN upscaling, audio preservation, and fixed_code / reconstructed / pingpong noise modes.
Speed. A multiscale sampler runs early steps at reduced resolution and later steps at full. Optional U-Net caching (DeepCache, First Block Cache) reuses work across denoiser evaluations, and compile_unet torch.compiles the U-Net, ControlNets and VAE paths. A tiled sampler handles resolutions that won't fit otherwise.
Image-edit models
FLUX.2 Klein Edit (4B / 9B)
HiDream-O1 Edit
Qwen Image Edit 2511 (+ GGUF)
Mage-Flow Edit (4B, + turbo)
Alongside the diffusion pipeline, VibeWarp drives four instruction-edit models through an external ComfyUI server over HTTP — nothing is vendored, and live progress is relayed back into the UI.
All four share an ordered Reference Images editor — the first image is always sent, later slots take raw or temporal frames or an uploaded style reference via drag-and-drop.
Temporal contact sheets (experimental) tile several frames into one image and edit them together, so the model sees a frame's neighbours instead of each frame in isolation, then split the result back into frames. Adaptive or fixed layouts, a configurable gutter, and an optional per-sheet instruction. Works across all four edit families, snapping to each model's preferred resolutions.
The interface
A local web UI (Svelte + FastAPI) at http://localhost:7860, plus a full CLI (
github.com/Sxela/VibeWarp/bl…). The form is generated from the config itself, so it never drifts from what the engine accepts.
Render. Settings grouped by how you actually use them — what you set once per project, what you change every run, and what you touch once or never. The input video is probed on the spot, so the resolution and frame count shown are what the render will genuinely produce, not an estimate. Queued rendering, live progress and logs, and full config import/export. The ControlNet panel scans your model directory and tells you which checkpoints are actually present before you start.
History & Comparison. Every past run, inspectable frame by frame and layer by layer: the init frame, warped init, processed consistency mask, each ControlNet's source and detected map, the diffusion input, and the output — side by side on a frame slider, so you can see exactly what a ControlNet was fed and what it produced.
• Loop playback — press play on the frame stepper to judge motion without assembling a video first, including part-way through a render.
• Resume — cancelling never discards rendered frames. Continue a run from its first missing frame with the previous stylized frame restored, or assemble a partial video at any time and play it in place.
• Compare — Ctrl/Cmd-click any two runs to diff every saved setting, including settings one run predates entirely.
• Load settings — pull any past run's config back into the form to iterate on it. Supports vanilla WarpFusion settings.
• Errors and settings search - you now see which input field failed validation and can just search for a setting you need.
The run gallery stays usable with hundreds of runs: thumbnails are cached per run rather than served as full renders, and a run still rendering updates its frame count live.
Hope you like it!
For more updates follow me on patreon -
patreon.com/c/sxela