tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8-GPTQ Image-Text-to-Text • 175B • Updated about 11 hours ago • 683 • 18
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published Aug 17 • 50
TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs Paper • 2607.22639 • Published Jun 22 • 5
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published Jul 19 • 93
PraMem: Practice-derived Experiential Memory for Long-horizon Behavior Prediction Paper • 2607.02881 • Published Jul 3 • 7
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Paper • 2607.04033 • Published Jul 4 • 55