A hybrid-attention text-to-video model that generates 14.4-second clips in 11.23 seconds on 8× B200 GPUs — faster than the video plays.
VDN-H3 adds a linear attention branch alongside the existing softmax attention, delivering massive speedups without sacrificing quality.
Frame-wise linear attention that scales O(n) with sequence length. Handles the bulk of the computation, making generation dramatically faster than dense softmax attention.
Preserves the original MiniMax H3 backbone's visual quality and temporal consistency. Ensures the output retains the rich expressiveness of the full model.
Two small LoRA adapters merge into the backbone during inference. The original MiniMax H3 weights remain untouched — zero compromise on the base model.
Weights, optimized inference stack with FP8 and kernel compilation, and training code — all released. No black boxes.
What makes VDN-H3 stand out from other video generation models.
14.4s video in 11.23s on 8× B200 — faster than the video itself plays. The model outruns real-time playback.
Linear attention for speed + softmax attention for quality. Best of both worlds in a single architecture.
Separate adapters merge at inference time. Backbone stays untouched — easy to update, swap, or revert.
FP8 inference kernels compiled via Triton. Nearly 3× faster than the dense baseline on a single GPU.
Near-linear scaling across 8 GPUs. From 90.5s to 18.3s (H200) with minimal overhead.
Weights, training code, inference stack — all public. Built on Diffusers, FlashAttention, and Triton.
Steady-state denoising speed on the 768p, 14.4-second video generation workload. Excludes model loading, warm-up, VAE decoding, and MP4 encoding.
| GPU | Configuration | GPUs | s/NFE | 50 NFE | 8 NFE (DMD) |
|---|---|---|---|---|---|
| H200 | dense MiniMax-H3 | 1 | 32.7 | 27.3 min | 4.4 min |
| H200 | VDN-H3 FP8 | 1 | 11.2 | 9.4 min | 90.5 s 3× |
| H200 | VDN-H3 FP8 Distributed | 8 | 2.29 | 1.9 min | 18.3 s 9× |
| B200 | dense MiniMax-H3 (cuDNN) | 1 | 16.74 | 13.95 min | 2.23 min |
| B200 | VDN-H3 FP8 | 1 | 6.41 | 5.3 min | 51 s 2.6× |
| B200 | VDN-H3 FP8 Distributed | 8 | 1.40 | 1.2 min | 11.23 s 12× |
★ Highlighted row = headline configuration: 8-step generation faster than real-time playback.
Generate your own videos with VDN-H3. The model runs on a single GPU (FP8) or scales to 8 GPUs for maximum speed.
Rewrite your prompt using H3-Context-IR or the official prompt-writing skills before encoding with the VLM. This can greatly improve generated video quality.