Verifiable Rewards Improve Native Visual Reasoning Training
A 300-task procedural suite replaces preference-only judging with task-specific scorers, raising matched reinforcement-learning performance from 0.509 to 0.548.
Sep 4, 20264 min2608.26105
Multimedia systems and multimodal content processing.
A 300-task procedural suite replaces preference-only judging with task-specific scorers, raising matched reinforcement-learning performance from 0.509 to 0.548.
VBVR pairs 200 curated reasoning tasks with rule-based, human-aligned evaluation to study whether video models generalize across spatiotemporal reasoning problems.