Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains Paper • 2608.09873 • Published Aug 10 • 29
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published Jul 18 • 99
One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding Paper • 2606.30084 • Published Jun 29 • 8
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models Paper • 2605.31597 • Published May 29 • 9
OpenComputer: Verifiable Software Worlds for Computer-Use Agents Paper • 2605.19769 • Published May 19 • 68
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration Paper • 2605.20025 • Published May 19 • 90
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence Paper • 2605.12882 • Published May 13 • 65
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Paper • 2605.06130 • Published May 7 • 74
Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment Paper • 2604.19548 • Published Apr 21 • 11
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation Paper • 2604.08540 • Published Apr 9 • 5
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver Paper • 2604.08377 • Published Apr 9 • 226