SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 1 day ago • 56
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 1 day ago • 56
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published 15 days ago • 9
LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation Paper • 2608.00267 • Published Jul 31 • 3
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities Paper • 2607.24821 • Published Jul 17 • 18
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published about 1 month ago • 183
Beacon: Knowing When and How to Perform Agentic Visual Reasoning Paper • 2607.28595 • Published Jul 30 • 55
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Paper • 2607.28509 • Published Jul 30 • 30
ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders Paper • 2607.21217 • Published Jul 23 • 8
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Paper • 2607.10350 • Published Jul 15 • 85
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Paper • 2607.10350 • Published Jul 15 • 85
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published Jul 14 • 102
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning Paper • 2607.07708 • Published Jul 8 • 89
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 160
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents Paper • 2606.22883 • Published Jun 22 • 37
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents Paper • 2606.22883 • Published Jun 22 • 37