DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published 20 days ago • 165
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space Paper • 2607.25675 • Published Jul 28 • 69
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 157
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published Jul 21 • 114
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published Jul 20 • 63
Search and Refine During Think: Autonomous Retrieval-Augmented Reasoning of LLMs Paper • 2505.11277 • Published May 16, 2025 • 12
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 120
Motion4Motion: Motion Transfer Across Subjects at Inference Paper • 2607.11644 • Published Jul 13 • 8
RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation Paper • 2607.06558 • Published Jul 7 • 27
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots Paper • 2607.02501 • Published Jul 2 • 55
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Paper • 2606.22873 • Published Jun 22 • 16
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Paper • 2606.24428 • Published Jun 23 • 15
Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems Paper • 2606.18837 • Published Jun 17 • 11
Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts Paper • 2606.05922 • Published Jun 4 • 37
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales Paper • 2606.12736 • Published Jun 10 • 6
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling Paper • 2606.18023 • Published Jun 16 • 161
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Paper • 2606.19195 • Published Jun 17 • 80
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Paper • 2606.09426 • Published Jun 8 • 49