SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 4 days ago • 27
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 17 days ago • 84
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization Paper • 2607.04988 • Published 21 days ago • 27
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents Paper • 2604.26752 • Published Apr 29 • 113
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control Paper • 2601.05138 • Published Jan 8 • 19
NitroGen: An Open Foundation Model for Generalist Gaming Agents Paper • 2601.02427 • Published Jan 4 • 46
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models Paper • 2512.02556 • Published Dec 2, 2025 • 271
Running on CPU Upgrade Featured 3.25k The Smol Training Playbook 📚 3.25k The secrets to building world-class LLMs
VIST3A: Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator Paper • 2510.13454 • Published Oct 15, 2025 • 10
Learning an Image Editing Model without Image Editing Pairs Paper • 2510.14978 • Published Oct 16, 2025 • 9
FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers Text-to-Video • 5B • Updated Nov 25, 2025 • 67.1k • 65