Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Paper • 2608.12149 • Published 6 days ago • 30
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Paper • 2608.12149 • Published 6 days ago • 30
meta-models/Muse-Glimmer-30B Image-Text-to-Text • 30B • Updated 6 days ago • 334k • • 1.66k
OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Paper • 2605.19660 • Published May 19 • 40
OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Paper • 2605.19660 • Published May 19 • 40
Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models Paper • 2604.01622 • Published Apr 2 • 7
Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models Paper • 2604.01622 • Published Apr 2 • 7
EpochX: Building the Infrastructure for an Emergent Agent Civilization Paper • 2603.27304 • Published Mar 28 • 47
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence Paper • 2511.18538 • Published Nov 23, 2025 • 306
view article Article Announcing NeurIPS 2025 E2LM Competition: Early Training Evaluation of Language Models tiiuae • Jul 4, 2025 • 11