Danyang Zhang

2 articles on SOTA Papers

Million-Clip Dataset Tests Video Reasoning Beyond Visual Quality

VBVR pairs 200 curated reasoning tasks with rule-based, human-aligned evaluation to study whether video models generalize across spatiotemporal reasoning problems.

Aug 30, 20264 min2602.20159

Qwen-CUA Reaches 86.2 on Verified Computer Use

A 397B-A17B mixture-of-experts agent learns screenshot-only keyboard and mouse control from verifiable interactive trajectories, improving OSWorld-Verified performance to 86.2.

Aug 8, 20264 min2608.02352