FilmBench Reveals Cinematic Gaps in Video Generators
A director-designed taxonomy and evaluation agent track human model rankings at Spearman ρ=0.95–0.96 while exposing multi-shot and dynamic-aesthetics failures.
Underlying Paper
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
Video-generation benchmarks commonly ask whether a clip is plausible, text-aligned, and temporally smooth. FilmBench argues that those tests miss the criteria used to make and assess filmed scenes: shot design, staging, performance, editing continuity, and the visual choices that make a sequence read as cinema rather than a competent isolated clip. The paper therefore evaluates text-to-video (T2V) and reference-to-video (R2V) systems against prompts derived from real film material rather than generic web prompts or synthetic templates.
Core Contribution
FilmBench is a benchmark built around professional Cinematic Language. Its 1,169 prompts are reverse-engineered from selected clips spanning 20 genres, with directors choosing the source material and prompts following real shot lists. The central design choice is consequential: 1,056 prompts contain multiple shots. That makes the benchmark test whether a model can maintain cinematic intent across a scene, not merely render a short attractive moment.
The evaluation taxonomy has three L1 axes, 12 L2 components, and 35 sub-metrics for T2V, plus three R2V-specific sub-metrics. This is a more granular target than a single visual-quality score. Figure 1 lays out that hierarchy and its grounding in film-school practice. The benchmark also introduces FilmOps, an open-source core suite of Cinematic Language operators, alongside an expert-grade automatic evaluator intended to score the taxonomy at scale.
Technical Approach
The paper’s assessment pipeline combines a film-derived prompt set with fine-grained automated scoring and human validation. For T2V, models must realize the scripted scene from text. For R2V, they must preserve and extend a reference scene’s cinematic attributes while generating the requested output. The distinction matters because a reference can constrain visual identity and composition, but it does not remove the need to handle shot transitions, motion, or narrative continuity.
Rather than treating all errors as equivalent, FilmBench scores dimensions separately. This lets the authors locate a model’s weakness within a scene. In the R2V example reverse-engineered from La La Land, Figure 3 compares Seedance 2.0 with Vidu Q2 Pro over the benchmark’s per-dimension measures; the reported aggregate scores are 94.97 and 73.61, respectively. The figure is useful because it turns a 21.36-point aggregate difference into an account of which cinematic dimensions separate the systems.
Results and Analysis
The authors benchmark nine T2V systems and seven R2V systems. Their automatic evaluator reproduces the human ranking at the model level with Spearman for T2V and for R2V. That is strong evidence that the evaluator can order the tested models similarly to the expert human assessments, although it is evidence about ranking across this fixed model set rather than proof that every individual clip score is reliable.
The benchmark’s substantive finding is less flattering for current generators: scores are materially lower under film-oriented evaluation than under prior web-style benchmarks. The reported gaps concentrate in dynamic aesthetics and become more pronounced when a task moves from one shot to several. This is a useful stress test because the degradation is not merely a failure to create a visually acceptable frame; it concerns the continuity and shot-level control needed for a coherent sequence.
The multi-shot T2V example in Figure 4 makes the model spread concrete. On a science-fiction mech-battle scene, Seedance 2.0 scores 86.11 while Grok Imagine Video scores 55.97. The 30.14-point difference suggests that high-level prompt compliance alone is not enough to identify systems that can preserve cinematic structure through a multi-shot sequence. At the same time, a benchmark score should not be read as a universal measure of video quality: it is deliberately optimized for the paper’s film-language taxonomy and curated scenario distribution.
Limits of the Evidence
FilmBench is a demanding and well-specified evaluation, but its coverage is bounded by 20 curated genres, selected award-winning-film references, and the 16 evaluated model configurations. The high rank correlations validate the proposed evaluator at model level, not necessarily its calibration for rare visual styles, long-form productions, interactive workflows, or new model families. The paper also emphasizes gaps in dynamic aesthetics and multi-shot generation, but the reported evidence establishes comparative benchmark performance rather than a causal explanation for why particular architectures fail.
Evidence Box
strongKey Claims
- •Film-oriented criteria expose weaknesses missed by generic video benchmarks
- •FilmOps-based automatic scoring follows expert human model rankings
- •Multi-shot generation creates a larger performance gap than single-shot generation
Key Results
- •1,169 prompts across 20 genres, including 1,056 multi-shot prompts
- •Model-level Spearman ρ=0.95 for 9 T2V systems against human rankings
- •Model-level Spearman ρ=0.96 for 7 R2V systems against human rankings
- •R2V example: Seedance 2.0 scores 94.97 versus Vidu Q2 Pro at 73.61
Limitations & Caveats
- •Coverage limited to clips curated from award-winning films in 20 genres
- •Human-alignment evidence is reported at model-ranking level rather than per-clip calibration
- •Evaluation covers 9 T2V and 7 R2V model configurations
- •Benchmark scores do not establish causal reasons for dynamic-aesthetics or multi-shot failures