From Pixels to States: Rethinking Interactive World Models as Game Engines Paper • 2607.14076 • Published 5 days ago • 32
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Paper • 2607.11523 • Published 7 days ago • 12
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents Paper • 2607.02255 • Published 18 days ago • 64
YoCausal: How Far is Video Generation from World Model? A Causality Perspective Paper • 2605.30346 • Published May 28 • 55
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency Paper • 2501.04931 • Published Jan 9, 2025
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents Paper • 2508.17198 • Published Aug 24, 2025 • 10
SFHand: A Streaming Framework for Language-guided 3D Hand Forecasting and Embodied Manipulation Paper • 2511.18127 • Published Nov 22, 2025 • 1
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions Paper • 2510.27195 • Published Oct 31, 2025 • 1
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 171
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference Paper • 2603.25730 • Published Mar 26 • 53