Jeesup/svd-safety-l2_remove40_swapdiscnet_b010_r04 Text Generation • 7B • Updated 13 days ago • 371 • 7
Autobrik/tangram-square-rectangle-house-panda-300ep Franka Panda • Updated 15 days ago • 300 episodes • 83 • 1
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 19 days ago • 174
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution Paper • 2608.25593 • Published Aug 26 • 69
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published Aug 26 • 174
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published Aug 12 • 110
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 178
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Paper • 2607.28802 • Published Jul 30 • 11
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Paper • 2607.21655 • Published Jul 22 • 111
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Paper • 2607.25895 • Published Jul 28 • 95
baohao/SFT_Math_Qwen3-4B-Instruct-2507_SFT_on_Math_Qwen3-4B-Instruct-2507_SFT-RL_trace 4B • Updated Jul 24 • 5 • 1