arxiv:2607.23802
Huazheng Wang
huazhengwang
AI & ML interests
Reinforcement Learning, Information Retrieval, LLM Agent.
Recent Activity
authored a paper about 1 month ago
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement upvoted a paper about 1 month ago
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs upvoted a paper about 1 month ago
Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline ParallelismOrganizations
None yet