Machine Learning

cs.LG

All aspects of machine learning research including supervised, unsupervised, and reinforcement learning.

Sort:

SkillNet Cuts Agent Steps Through Reusable Skills

An ontology, evaluation scheme, and 600,000-skill repository let agents retrieve and compose prior procedures, reporting 40% higher rewards with 30% fewer steps.

Aug 23, 20264 min2603.04448

Agent Swarm Cuts Complex Task Latency Up To 4.5×

Kimi K2.5 couples joint text-vision training with dynamically scheduled parallel subagents, improving agentic search scores while reducing time to target quality.

Aug 13, 20264 min2602.02276

Kimi K3 Brings Open Weights Closer to Frontier Agents

A 2.8T-parameter MoE combines Delta Attention, 16-of-896 expert routing, and agentic reinforcement learning to approach leading proprietary systems at lower task cost.

Aug 10, 20264 min2607.24653

Qwen-CUA Reaches 86.2 on Verified Computer Use

A 397B-A17B mixture-of-experts agent learns screenshot-only keyboard and mouse control from verifiable interactive trajectories, improving OSWorld-Verified performance to 86.2.

Aug 8, 20264 min2608.02352

Request-Level Sharing Cuts Recommendation Inference Cost

ROCS separates request-side and candidate-side computation, reaching up to 3× higher retrieval QPS and 50% higher ranking QPS in production.

Aug 1, 20265 min2607.27744

Mini Activations Narrow Frontier Model Gaps

A sparse 229.9B-parameter MoE activates 9.8B parameters per token and pairs agent-native data with RL for coding, cowork, and reasoning tasks.

Jul 31, 20265 min2605.26494

Hybrid Mamba MoE Trades Dense Scale for Throughput

Soofi S activates 3B of 30B parameters per token and matches larger open models while reaching 4.8k TPS/GPU at 40K context.

Jul 29, 20265 min2607.09424

ABot-World-0 Runs Interactive Worlds on One Desktop GPU

Progressive causal distillation and a co-designed streaming stack produce 720P rollouts at up to 16 FPS on an RTX 5090.

Jul 28, 20265 min2607.19191

Six-Stage Filtering Exposes Weak Drug Generators

HEDGEHOG runs generated molecules through medicinal-chemistry, synthesis, docking, and 3D pose checks, leaving only 0.65% of 230,000 candidates.

Jul 28, 20265 min2607.13155 Code available

SlimPer Scales Recommendation Histories Without Quadratic Attention

A compact knowledge-base backbone queries full user histories at each layer, improving offline ranking metrics while cutting per-example FLOPs by 8×–25×.

Jul 28, 20265 min2607.12281