Artificial Intelligence

cs.AI

Covers all areas of AI except Vision, Robotics, Machine Learning, Multiagent Systems, and NLP.

Sort:

Mobius Separates Knowledge From Reasoning for Faster Inference

A shared feed-forward memory and iterative attention reasoners retain comparable downstream scores while reducing continual-pretrained 35B-model inference time by nearly 4×.

Aug 25, 20264 min2608.14290

SkillNet Cuts Agent Steps Through Reusable Skills

An ontology, evaluation scheme, and 600,000-skill repository let agents retrieve and compose prior procedures, reporting 40% higher rewards with 30% fewer steps.

Aug 23, 20264 min2603.04448

Agent Swarm Cuts Complex Task Latency Up To 4.5×

Kimi K2.5 couples joint text-vision training with dynamically scheduled parallel subagents, improving agentic search scores while reducing time to target quality.

Aug 13, 20264 min2602.02276

Monocular Navigation Surpasses Depth-Based Systems on R2R-CE

An 8B vision-language model predicts image-space waypoints from a single RGB stream, reaching 77.4% unseen-environment success while cutting supervised training tokens 22×.

Aug 5, 20264 min2607.20785

FilmBench Reveals Cinematic Gaps in Video Generators

A director-designed taxonomy and evaluation agent track human model rankings at Spearman ρ=0.95–0.96 while exposing multi-shot and dynamic-aesthetics failures.

Aug 3, 20264 min2607.24241

Request-Level Sharing Cuts Recommendation Inference Cost

ROCS separates request-side and candidate-side computation, reaching up to 3× higher retrieval QPS and 50% higher ranking QPS in production.

Aug 1, 20265 min2607.27744

Mini Activations Narrow Frontier Model Gaps

A sparse 229.9B-parameter MoE activates 9.8B parameters per token and pairs agent-native data with RL for coding, cowork, and reasoning tasks.

Jul 31, 20265 min2605.26494

Hybrid Mamba MoE Trades Dense Scale for Throughput

Soofi S activates 3B of 30B parameters per token and matches larger open models while reaching 4.8k TPS/GPU at 40K context.

Jul 29, 20265 min2607.09424

Open Image Model Narrows Gap Under Tight Compute

Boogu-Image-0.1 combines a stronger multimodal encoder, agentic prompt rewriting, and curated data to train competitive generation and editing models for about $400K.

Jul 29, 20265 min2607.13125 Code available

Long-Horizon Workflows Expose Agent Reliability Gaps

OSWorld 2.0 packages 108 multi-application workflows into executable tests, where the best tested agent completes only 20.6%.

Jul 28, 20265 min2606.29537