Verifiable Rewards Improve Native Visual Reasoning Training
A 300-task procedural suite replaces preference-only judging with task-specific scorers, raising matched reinforcement-learning performance from 0.509 to 0.548.
Covers all areas of AI except Vision, Robotics, Machine Learning, Multiagent Systems, and NLP.
A 300-task procedural suite replaces preference-only judging with task-specific scorers, raising matched reinforcement-learning performance from 0.509 to 0.548.
VBVR pairs 200 curated reasoning tasks with rule-based, human-aligned evaluation to study whether video models generalize across spatiotemporal reasoning problems.
A shared feed-forward memory and iterative attention reasoners retain comparable downstream scores while reducing continual-pretrained 35B-model inference time by nearly 4×.
An ontology, evaluation scheme, and 600,000-skill repository let agents retrieve and compose prior procedures, reporting 40% higher rewards with 30% fewer steps.
Kimi K2.5 couples joint text-vision training with dynamically scheduled parallel subagents, improving agentic search scores while reducing time to target quality.
An 8B vision-language model predicts image-space waypoints from a single RGB stream, reaching 77.4% unseen-environment success while cutting supervised training tokens 22×.
A director-designed taxonomy and evaluation agent track human model rankings at Spearman ρ=0.95–0.96 while exposing multi-shot and dynamic-aesthetics failures.
ROCS separates request-side and candidate-side computation, reaching up to 3× higher retrieval QPS and 50% higher ranking QPS in production.
A sparse 229.9B-parameter MoE activates 9.8B parameters per token and pairs agent-native data with RL for coding, cowork, and reasoning tasks.
Boogu-Image-0.1 combines a stronger multimodal encoder, agentic prompt rewriting, and curated data to train competitive generation and editing models for about $400K.