1. PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models 2026-09-14 ▲ 172 on Hugging Face
  2. Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models 2026-09-11 ▲ 77 on Hugging Face
  3. Memory as Plans: World-Action Modeling with Memory-Grounded Planning 2026-09-10 ▲ 39 on Hugging Face
  4. The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement 2026-09-10 ▲ 88 on Hugging Face
  5. Show-Harness: Just a VLM Agent Can Play Robots 2026-09-09 ▲ 157 on Hugging Face
  6. ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs 2026-09-09 ▲ 51 on Hugging Face
  7. SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem 2026-09-07 ▲ 140 on Hugging Face
  8. Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work 2026-09-04 ▲ 86 on Hugging Face
  9. Exploring Collaboration between a language and a non-language agent 2026-09-02 ▲ 5 on Hugging Face
  10. RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning 2026-09-02 ▲ 14 on Hugging Face
  11. Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching 2026-09-01 ▲ 26 on Hugging Face
  12. Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling 2026-08-31 ▲ 104 on Hugging Face
  13. LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation 2026-08-31 ▲ 27 on Hugging Face
  14. ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training 2026-08-31 ▲ 47 on Hugging Face
  15. Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions 2026-08-31 ▲ 14 on Hugging Face
  16. Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory 2026-08-30 ▲ 17 on Hugging Face
  17. Scaling Automatic Research Agents via World Models 2026-08-29 ▲ 453 on Hugging Face
  18. UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City 2026-08-27 ▲ 109 on Hugging Face
  19. Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models 2026-08-27 ▲ 88 on Hugging Face
  20. Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization 2026-08-26 ▲ 20 on Hugging Face
  21. Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models 2026-08-24 ▲ 29 on Hugging Face
  22. DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation 2026-08-13 ▲ 87 on Hugging Face
  23. Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence 2026-08-13 ▲ 35 on Hugging Face
  24. H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models 2026-08-13 ▲ 11 on Hugging Face
  25. ComBodied Agents: a New Paradigm of Human-Centric Agentic AI 2026-08-11 ▲ 189 on Hugging Face
  26. Self-Evolving Embodied Agents via Skill-Harness Evolution 2026-08-11 ▲ 11 on Hugging Face
  27. SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding 2026-08-10 ▲ 26 on Hugging Face
  28. RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance 2026-08-10 ▲ 13 on Hugging Face
  29. CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems 2026-08-10 ▲ 6 on Hugging Face
  30. 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents 2026-08-09 ▲ 10 on Hugging Face
  31. GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? 2026-08-06 ▲ 43 on Hugging Face
  32. DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation 2026-08-06 ▲ 22 on Hugging Face
  33. From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models 2026-08-06 ▲ 33 on Hugging Face
  34. Uncertainty-Aware World Model for Aerial Image-Goal Navigation 2026-08-06 ▲ 9 on Hugging Face
  35. SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries 2026-08-06 ▲ 74 on Hugging Face
  36. When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents 2026-08-05 ▲ 12 on Hugging Face
  37. World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation 2026-08-05 ▲ 25 on Hugging Face
  38. SkillJack: Persistent Skill Backdoors in Self-Evolving Agents 2026-08-04 ▲ 22 on Hugging Face
  39. UniWorld-Design: From Pixel Generation to Layer-Native Design 2026-08-04 ▲ 21 on Hugging Face
  40. Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs 2026-08-03 ▲ 138 on Hugging Face
  41. Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data 2026-08-03 ▲ 23 on Hugging Face
  42. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills 2026-08-03 ▲ 11 on Hugging Face
  43. MiniWorld: Democratizing the Training of Video World Models from Scratch 2026-08-02 ▲ 18 on Hugging Face
  44. GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization 2026-08-02 ▲ 12 on Hugging Face
  45. DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents 2026-08-01 ▲ 15 on Hugging Face
  46. WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning 2026-07-31 ▲ 25 on Hugging Face
  47. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine 2026-07-30 ▲ 36 on Hugging Face
  48. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them 2026-07-30 ▲ 21 on Hugging Face
  49. ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow 2026-07-30 ▲ 14 on Hugging Face
  50. EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents 2026-07-30 ▲ 12 on Hugging Face
  51. TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM 2026-07-29 ▲ 128 on Hugging Face
  52. HumanCLAW: Can Vision-Language Models Act Through a Body? 2026-07-29 ▲ 71 on Hugging Face
  53. HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone 2026-07-28 ▲ 149 on Hugging Face
  54. CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents 2026-07-28 ▲ 103 on Hugging Face
  55. N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation 2026-07-26 ▲ 35 on Hugging Face