UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 1 day ago • 162
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents Paper • 2609.27334 • Published 9 days ago • 51
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 15 days ago • 191
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published Aug 24 • 211
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 29 days ago • 187
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 29 days ago • 104
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published Aug 10 • 796
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published Aug 31 • 97
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 30 days ago • 407
ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29 Text Generation • 9B • Updated about 1 month ago • 44
ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29 Text Generation • 9B • Updated about 1 month ago • 44
UserBench: An Interactive Gym Environment for User-Centric Agents Paper • 2507.22034 • Published Jul 29, 2025 • 31
ipfipfipf/Qwen3.5-4B-sdpo-react-mathcodesearch-sdpo-arm-e Image-Text-to-Text • 4B • Updated Aug 23 • 11
ipfipfipf/Qwen3.5-4B-sdpo-react-mathcodesearch-sdpo-arm-e Image-Text-to-Text • 4B • Updated Aug 23 • 11
ipfipfipf/Qwen3.5-4B-sdpo-react-mathcodesearch-sdpo-arm-a Image-Text-to-Text • 4B • Updated Aug 22 • 10
ipfipfipf/Qwen3.5-4B-sdpo-react-mathcodesearch-sdpo-arm-a Image-Text-to-Text • 4B • Updated Aug 22 • 10
ipfipfipf/Qwen3.5-4B-sdpo-react-mathcodesearch-grpo-arm-e Image-Text-to-Text • 4B • Updated Aug 22 • 18