Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning: model and rubric dataset.
AI & ML interests
Machine Learning, Natural Language Processing, Optimization, Multi-Modality, Artificial Intelligence
Recent Activity
View all activity
Papers
Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models
Skip a Layer or Loop It? Learning Program-of-Layers in LLMs
models 9
umd-zhou-lab/AudioRubrics
Audio-Text-to-Text • 12B • Updated • 78 • 3
umd-zhou-lab/controllable-wizardlm-7b
Text Generation • 7B • Updated • 20
umd-zhou-lab/controllable-llama2-7b
Text Generation • 7B • Updated • 10
umd-zhou-lab/recycled-wizardlm-7b-v2.0
Text Generation • Updated • 57 • 2
umd-zhou-lab/recycled-alpaca-7b-v2.0
Text Generation • Updated • 56 • 1
umd-zhou-lab/claude2-alpaca-7B
Text Generation • Updated • 20 • 2
umd-zhou-lab/claude2-alpaca-13B
Text Generation • Updated • 68 • 5
umd-zhou-lab/recycled-wizardlm-7b-v1.0
Text Generation • Updated • 12
umd-zhou-lab/recycled-alpaca-7b-v1.0
Text Generation • Updated • 13
datasets 14
umd-zhou-lab/AVQA-Audio-Rubrics
Viewer • Updated • 40.4k • 89 • 1
umd-zhou-lab/TSRBench
Updated • 323 • 3
umd-zhou-lab/V-REX
Viewer • Updated • 702 • 32
umd-zhou-lab/ChartAlignBench
Viewer • Updated • 22.6k • 107
umd-zhou-lab/ColorBench
Viewer • Updated • 5.89k • 353 • 4
umd-zhou-lab/CoSTAR
Viewer • Updated • 121 • 308 • 4
umd-zhou-lab/sRecycled_Wiz70
Viewer • Updated • 46.6k • 23 • 2
umd-zhou-lab/sRecycled_Alpaca
Viewer • Updated • 37.1k • 16 • 2
umd-zhou-lab/Reflect_Alpaca_All
Viewer • Updated • 208k • 58 • 1
umd-zhou-lab/Reflect_Wiz70_All
Viewer • Updated • 280k • 122 • 1