⚡ Video generation faster than real-time

VDN-MiniMax-H3

A hybrid-attention text-to-video model that generates 14.4-second clips in 11.23 seconds on 8× B200 GPUs — faster than the video plays.

11.23s
8-step generation (8×B200)
14.4s
Video length
82GB
Model weights
8
Denoising steps (DMD)

Hybrid Architecture

VDN-H3 adds a linear attention branch alongside the existing softmax attention, delivering massive speedups without sacrificing quality.

⚡

Linear Attention Branch

Frame-wise linear attention that scales O(n) with sequence length. Handles the bulk of the computation, making generation dramatically faster than dense softmax attention.

🎯

Softmax Attention Branch

Preserves the original MiniMax H3 backbone's visual quality and temporal consistency. Ensures the output retains the rich expressiveness of the full model.

🔌

Plug-and-Play LoRA

Two small LoRA adapters merge into the backbone during inference. The original MiniMax H3 weights remain untouched — zero compromise on the base model.

📦

Fully Open-Source

Weights, optimized inference stack with FP8 and kernel compilation, and training code — all released. No black boxes.

Text Prompt Qwen3-VL-32B Prompt Encoder Linear Attn O(n) · fast path Softmax Attn O(n²) · quality LoRA Adapters Hybrid Fusion 🎬 Video 768p · 14.4s VAE → MP4

Key Features

What makes VDN-H3 stand out from other video generation models.

⏱️

Real-Time Generation

14.4s video in 11.23s on 8× B200 — faster than the video itself plays. The model outruns real-time playback.

🪶

Hybrid Attention

Linear attention for speed + softmax attention for quality. Best of both worlds in a single architecture.

🧩

LoRA Merging

Separate adapters merge at inference time. Backbone stays untouched — easy to update, swap, or revert.

🔬

FP8 Optimization

FP8 inference kernels compiled via Triton. Nearly 3× faster than the dense baseline on a single GPU.

📐

Distributed Scaling

Near-linear scaling across 8 GPUs. From 90.5s to 18.3s (H200) with minimal overhead.

🌐

Fully Open

Weights, training code, inference stack — all public. Built on Diffusers, FlashAttention, and Triton.

Benchmarks

Steady-state denoising speed on the 768p, 14.4-second video generation workload. Excludes model loading, warm-up, VAE decoding, and MP4 encoding.

GPU Configuration GPUs s/NFE 50 NFE 8 NFE (DMD)
H200 dense MiniMax-H3 1 32.7 27.3 min 4.4 min
H200 VDN-H3 FP8 1 11.2 9.4 min 90.5 s 3×
H200 VDN-H3 FP8 Distributed 8 2.29 1.9 min 18.3 s 9×
B200 dense MiniMax-H3 (cuDNN) 1 16.74 13.95 min 2.23 min
B200 VDN-H3 FP8 1 6.41 5.3 min 51 s 2.6×
B200 VDN-H3 FP8 Distributed 8 1.40 1.2 min 11.23 s 12×

★ Highlighted row = headline configuration: 8-step generation faster than real-time playback.

Run It Yourself

Generate your own videos with VDN-H3. The model runs on a single GPU (FP8) or scales to 8 GPUs for maximum speed.

1. Clone & Set Up

# Install the code repository
git clone https://github.com/OpenVDN/vdn-minimax-h3.git
cd vdn-minimax-h3

# Set up environment (see repo README)
bash setup.sh

# Download model weights (~82 GB)
hf download OpenVDN/vdn-minimax-h3 --local-dir ckpts

2. Render Your First Video

# Run the 8-step (turbo) model on a single GPU
bash scripts/inference/8nfe_tuned_fp8.sh

# Or use your own prompt:
# 1. Encode with Qwen3-VL-32B VLM
python src/inference/encode_prompt.py \
  --prompt "A serene mountain lake at sunrise" \
  --out prompts/mine.pt

# 2. Run inference
python src/inference/infer.py \
  --config configs/inference/8nfe_tuned_fp8.yaml \
  checkpoint=ckpts/stage-dmd-step-250 \
  render.prompt_file=prompts/mine.pt \
  render.out=results/mine.mp4

3. For Best Quality

Rewrite your prompt using H3-Context-IR or the official prompt-writing skills before encoding with the VLM. This can greatly improve generated video quality.

Get the Code Read the Blog →

📖 Citation

@misc{xi2026videodeltanet,
  title = {VideoDeltaNet on MiniMax H3},
  author = {Haocheng Xi and Yiming Xie and Hexu Zhao and Yiwen Zhang and
          Michael Liu and Thomas Creavin and Kurt Keutzer and
          Xiuyu Li and Zhaoyang Lv and Chenfeng Xu and Haiwen Feng},
  year = {2026},
  url = {https://openvdn.github.io/}
}