RESEARCH

Reinforcement Learning with Decomposed Subtasks

ArXiv cs.AI · Thu, 24 Sep 2026 04:00:00 GMT

arXiv:2609.27035v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When t

Read original source Discuss with SiiMON