SkillNet Cuts Agent Steps Through Reusable Skills
An ontology, evaluation scheme, and 600,000-skill repository let agents retrieve and compose prior procedures, reporting 40% higher rewards with 30% fewer steps.
Image processing, computer vision, pattern recognition, and scene understanding.
An ontology, evaluation scheme, and 600,000-skill repository let agents retrieve and compose prior procedures, reporting 40% higher rewards with 30% fewer steps.
A director-designed taxonomy and evaluation agent track human model rankings at Spearman ρ=0.95–0.96 while exposing multi-shot and dynamic-aesthetics failures.
Smartphone capture and a modular processing toolchain turn 2,000 hours of human manipulation video into structured supervision for embodied models.
Boogu-Image-0.1 combines a stronger multimodal encoder, agentic prompt rewriting, and curated data to train competitive generation and editing models for about $400K.
A two-stage VLA recipe combines 100K hours of UMI trajectories with cross-embodiment post-training, reaching 57.4% on RoboCasa365.
Progressive causal distillation and a co-designed streaming stack produce 720P rollouts at up to 16 FPS on an RTX 5090.
Next-State-Prediction trains a shared latent space from video, events, and VQA data, improving balanced downstream readouts with a frozen backbone.
ABot-3DWorld 0 maps text, images, multi-view photos, and video into a shared Spatial Generative Primitive before generating panoramic exploration paths and reconstructing 3D Gaussian Splatting worlds.
Explicit reasoning and pixel-goal anchors decouple cognition from control, raising POI arrival to 77.3% with a reported 35.0% gain.
A Cosmos-derived causal diffusion model uses simulator state, action cues, and a streaming KV cache to render 704×1280 rollouts at 68–105 FPS.