arxiv:2508.21104
Penghong Zhao
DDDDrop
AI & ML interests
RL,Multimodal,Machine Learninh
Recent Activity
authored a paper about 1 year ago
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic
Reasoning upvoted a paper about 1 year ago
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic
ReasoningOrganizations
None yet