sanjaydoss/Multi-Agent_Reinforcement_Learning_Trading_System_Data Viewer • Updated Sep 3 • 5.28k • 217 • 19
crumb/bloom-560m-RLHF-SD2-prompter-aesthetic Text Generation • 0.6B • Updated Mar 19, 2023 • 330 • 29
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents Paper • 2610.01892 • Published 9 days ago • 27
HuatuoGPT-3: RL-Only Domain Adaptation from Base Models Paper • 2610.05966 • Published 5 days ago • 35
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning Paper • 2610.02824 • Published 8 days ago • 33
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 24 days ago • 30
jaehyeokdoo2/openpi-droid-pnpcarrot-singetask-qflow-offlinerl-criticwarmup2000-alpha100-bs8-test Updated Mar 4 • 5
All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts Paper • 2609.24058 • Published 19 days ago • 55
Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 21 days ago • 18
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents Paper • 2609.23377 • Published 20 days ago • 50