Archive
Papers
55 robotics papers the field upvoted. The homepage shows this week; this page remembers all of them.
- PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
- Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
- Memory as Plans: World-Action Modeling with Memory-Grounded Planning
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
- Show-Harness: Just a VLM Agent Can Play Robots
- ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
- SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem
- Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work
- Exploring Collaboration between a language and a non-language agent
- RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning
- Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
- Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
- LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
- ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
- Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
- Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
- Scaling Automatic Research Agents via World Models
- UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
- Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
- Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
- Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
- DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
- Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
- H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
- ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
- Self-Evolving Embodied Agents via Skill-Harness Evolution
- SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
- RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
- CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
- 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
- GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
- DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
- From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models
- Uncertainty-Aware World Model for Aerial Image-Goal Navigation
- SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
- When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents
- World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
- SkillJack: Persistent Skill Backdoors in Self-Evolving Agents
- UniWorld-Design: From Pixel Generation to Layer-Native Design
- Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
- Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
- Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
- MiniWorld: Democratizing the Training of Video World Models from Scratch
- GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization
- DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents
- WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
- ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
- EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
- HumanCLAW: Can Vision-Language Models Act Through a Body?
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents
- N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation