Computer Vision

cs.CV

Image processing, computer vision, pattern recognition, and scene understanding.

Sort:

SkillNet Cuts Agent Steps Through Reusable Skills

An ontology, evaluation scheme, and 600,000-skill repository let agents retrieve and compose prior procedures, reporting 40% higher rewards with 30% fewer steps.

Aug 23, 20264 min2603.04448

FilmBench Reveals Cinematic Gaps in Video Generators

A director-designed taxonomy and evaluation agent track human model rankings at Spearman ρ=0.95–0.96 while exposing multi-shot and dynamic-aesthetics failures.

Aug 3, 20264 min2607.24241

Open Egocentric Data Lowers Barriers for Robot Learning

Smartphone capture and a modular processing toolchain turn 2,000 hours of human manipulation video into structured supervision for embodied models.

Jul 29, 20265 min2607.14183

Open Image Model Narrows Gap Under Tight Compute

Boogu-Image-0.1 combines a stronger multimodal encoder, agentic prompt rewriting, and curated data to train competitive generation and editing models for about $400K.

Jul 29, 20265 min2607.13125 Code available

Xiaomi Scales Robot Policies With Real Trajectories

A two-stage VLA recipe combines 100K hours of UMI trajectories with cross-embodiment post-training, reaching 57.4% on RoboCasa365.

Jul 28, 20265 min2607.15330

ABot-World-0 Runs Interactive Worlds on One Desktop GPU

Progressive causal distillation and a co-designed streaming stack produce 720P rollouts at up to 16 FPS on an RTX 5090.

Jul 28, 20265 min2607.19191

Orca Unifies World Modeling Across Text, Vision, and Action

Next-State-Prediction trains a shared latent space from video, events, and VQA data, improving balanced downstream readouts with a frozen backbone.

Jul 28, 20265 min2606.30534

Unified 3D Pipeline Turns Mixed Inputs Into Explorable Worlds

ABot-3DWorld 0 maps text, images, multi-view photos, and video into a shared Spatial Generative Primitive before generating panoramic exploration paths and reconstructing 3D Gaussian Splatting worlds.

Jul 28, 20264 min2607.11673

Slow-Fast Navigation Model Improves Urban POI Arrival

Explicit reasoning and pixel-goal anchors decouple cognition from control, raising POI arrival to 77.3% with a reported 35.0% gain.

Jul 28, 20264 min2607.10383

OmniDreams Runs Generative AV Simulation in Real Time

A Cosmos-derived causal diffusion model uses simulator state, action cues, and a streaming KV cache to render 704×1280 rollouts at 68–105 FPS.

Jul 28, 20265 min2606.03159 Code available