OMNIVBENCH / BENCHMARK

OmniVBench

Evaluating how models preserve, disentangle, and compose reference factors according to the instruction.

01

Comprehensive R2V Task Coverage

813 evaluation cases span 7 task families and 18 fine-grained tasks, covering heterogeneous reference types and multi-reference compositions.

02

Factor-Grounded Evaluation

12,172 case-specific checklist items assess factor preservation, disentanglement, and target binding. Manually verified checklists are shared across models.

03

Fine-Grained Capability Analysis

Evaluation of 11 open- and closed-source models reveals task-specific strengths and separates reference fidelity from instruction realization.

OmniVBench taxonomy: seven task families, 18 sub-tasks and 813 evaluation cases
TASK TAXONOMY

Single and Multi-Reference Control

Tasks distinguish what to preserve, what to change, and how to combine references.

Single-reference

Extract the designated content, motion, style, structure, or narrative from an image or video reference.

Multi-reference

Combine multiple entities or complementary factors while binding each reference to its intended target.

Explore model outputs

Task coverage compared with existing benchmarks

BenchmarkContent Ref.Motion Ref.Style Ref.Structure Ref.Narrative Ref.Multiple Refs.
ObjectCharacterSceneActionCamera MotionStyleGreyboxLine ArtRough StoryboardMulti-Panel StoryboardStoryPreceding-ShotMulti-ContentCross-Aspect
OpenS2V-Eval
VACE-Bench
UniVBench
IntelligentVBench
FashionVideoBench
OmniVBench (Ours)

Model comparison across task families

No model leads across all seven task families. Seedance 2.5 achieves the highest overall score, closely followed by MiniMax H3.

Models are ordered by Overall score. Each task score averages Reference Fidelity (RF), Instruction Realization (IR), and Video Quality (VQ). Orange marks the highest score in each column.

ModelGroupContentMotionStyleStructureNarrativeMulti-contentCross-aspectOverall
Seedance 2.5Closed78.8866.1067.1373.3573.9777.7771.5372.68
MiniMax H3Open78.8665.3266.2472.9473.9077.2272.3672.41
Seedance 2.0Closed79.0063.0563.7368.3673.7676.6371.8370.91
Happy Horse 1.0Closed75.9068.5865.6269.4672.3474.6269.7170.89
Gemini OmniClosed75.0262.2569.2470.1375.1173.3670.1870.76
Kling 3.0 OmniClosed75.9964.2055.9071.1268.5975.9868.3368.59
Vidu-Q2-ProClosed71.8747.6559.8751.8150.6872.0762.2159.45
BerniniOpen69.2755.9250.6269.4442.0260.3251.0356.95
LoomVideoOpen61.2452.0942.0362.5346.9849.0244.8651.25
UniVideoOpen66.0144.5247.4048.8436.9758.0046.0949.69
OmniWeavingOpen62.2350.5841.4063.1938.0347.4939.2548.88

Reference inputs and generated outputs

View full size ↗
Representative OmniVBench examples pairing reference inputs with generated outputs across content, motion, style, structure, narrative, multi-content, and cross-aspect tasks

Explore benchmark cases

Two selected cases per sub-task, with model outputs compared using the same references and instruction.

References
Instruction

Model outputs