StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks
Abstract
Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at https://github.com/amazon-science/StructRL.
Community
StructRL is an online RL framework for long-horizon VLA tasks. Instead of rewarding only final task success, it rewards verifiable progress:
- an LLM decomposes each task into subtasks with simulator-checkable completion criteria and a prerequisite structure,
- a completion is rewarded only after its prerequisites are done (structure-aware gating),
- and the reward is scaled by completion pace relative to the SFT demonstrations (dynamic pacing).
On RoboCasa365 and LIBERO-Long with GR00T-N1.5 and π0.5, StructRL outperforms the evaluated online RL baselines, e.g. 49.1% vs. 41.5% (SimpleVLA-RL) on RoboCasa365 and 96.6% vs. 92.4% on LIBERO-Long, both with GR00T-N1.5.
- Code: https://github.com/amazon-science/StructRL
- Project page (videos): https://amazon-science.github.io/StructRL/
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning (2026)
- PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models (2026)
- Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation (2026)
- SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations (2026)
- From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention (2026)
- LexiconVLA: Learning Reusable Atomic Action Codebooks for Unseen Tasks (2026)
- TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.36352 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper