Verifiable Rewards Improve Native Visual Reasoning Training
A 300-task procedural suite replaces preference-only judging with task-specific scorers, raising matched reinforcement-learning performance from 0.509 to 0.548.
Image processing, computer vision, pattern recognition, and scene understanding.
A 300-task procedural suite replaces preference-only judging with task-specific scorers, raising matched reinforcement-learning performance from 0.509 to 0.548.
VBVR pairs 200 curated reasoning tasks with rule-based, human-aligned evaluation to study whether video models generalize across spatiotemporal reasoning problems.
An ontology, evaluation scheme, and 600,000-skill repository let agents retrieve and compose prior procedures, reporting 40% higher rewards with 30% fewer steps.
A director-designed taxonomy and evaluation agent track human model rankings at Spearman ρ=0.95–0.96 while exposing multi-shot and dynamic-aesthetics failures.
Boogu-Image-0.1 combines a stronger multimodal encoder, agentic prompt rewriting, and curated data to train competitive generation and editing models for about $400K.
Smartphone capture and a modular processing toolchain turn 2,000 hours of human manipulation video into structured supervision for embodied models.
Progressive causal distillation and a co-designed streaming stack produce 720P rollouts at up to 16 FPS on an RTX 5090.
Next-State-Prediction trains a shared latent space from video, events, and VQA data, improving balanced downstream readouts with a frozen backbone.
ABot-3DWorld 0 maps text, images, multi-view photos, and video into a shared Spatial Generative Primitive before generating panoramic exploration paths and reconstructing 3D Gaussian Splatting worlds.
Explicit reasoning and pixel-goal anchors decouple cognition from control, raising POI arrival to 77.3% with a reported 35.0% gain.