OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning Paper • 2610.12458 • Published 3 days ago • 12
Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation Paper • 2610.02368 • Published 10 days ago • 4
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining Paper • 2606.20521 • Published Jun 18 • 15
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Paper • 2605.18287 • Published May 18 • 15
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation Paper • 2606.28128 • Published Jun 26 • 55
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining Paper • 2606.20521 • Published Jun 18 • 15
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Paper • 2605.18287 • Published May 18 • 15
HumanNet: Scaling Human-centric Video Learning to One Million Hours Paper • 2605.06747 • Published May 7 • 55
Enhancing Spatial Understanding in Image Generation via Reward Modeling Paper • 2602.24233 • Published Feb 27 • 60
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head Paper • 2601.07832 • Published Jan 12 • 53