Computer Science

Research across all areas of computer science, from theoretical foundations to applied systems.

Sort:

Verifiable Rewards Improve Native Visual Reasoning Training

A 300-task procedural suite replaces preference-only judging with task-specific scorers, raising matched reinforcement-learning performance from 0.509 to 0.548.

Sep 4, 20264 min2608.26105

Million-Clip Dataset Tests Video Reasoning Beyond Visual Quality

VBVR pairs 200 curated reasoning tasks with rule-based, human-aligned evaluation to study whether video models generalize across spatiotemporal reasoning problems.

Aug 30, 20264 min2602.20159

Mobius Separates Knowledge From Reasoning for Faster Inference

A shared feed-forward memory and iterative attention reasoners retain comparable downstream scores while reducing continual-pretrained 35B-model inference time by nearly 4×.

Aug 25, 20264 min2608.14290

SkillNet Cuts Agent Steps Through Reusable Skills

An ontology, evaluation scheme, and 600,000-skill repository let agents retrieve and compose prior procedures, reporting 40% higher rewards with 30% fewer steps.

Aug 23, 20264 min2603.04448

XPolicyLab Cuts Robot Policy Integration From NM to N Plus M

A shared policy contract and isolated client/server runtime let 42 robot policies connect to simulations and real-robot evaluation without pairwise adapters.

Aug 18, 20264 min2608.09892

DREAM Improves Taobao Feed Metrics Without Replacing Rankers

An intent engine and agentic meta-control layer steer existing recommendation stages, lifting GMV 1.31% when extended through fine ranking.

Aug 15, 20264 min2608.09408

Agent Swarm Cuts Complex Task Latency Up To 4.5×

Kimi K2.5 couples joint text-vision training with dynamically scheduled parallel subagents, improving agentic search scores while reducing time to target quality.

Aug 13, 20264 min2602.02276

Kimi K3 Brings Open Weights Closer to Frontier Agents

A 2.8T-parameter MoE combines Delta Attention, 16-of-896 expert routing, and agentic reinforcement learning to approach leading proprietary systems at lower task cost.

Aug 10, 20264 min2607.24653

Qwen-CUA Reaches 86.2 on Verified Computer Use

A 397B-A17B mixture-of-experts agent learns screenshot-only keyboard and mouse control from verifiable interactive trajectories, improving OSWorld-Verified performance to 86.2.

Aug 8, 20264 min2608.02352

Monocular Navigation Surpasses Depth-Based Systems on R2R-CE

An 8B vision-language model predicts image-space waypoints from a single RGB stream, reaching 77.4% unseen-environment success while cutting supervised training tokens 22×.

Aug 5, 20264 min2607.20785