Group Relative Policy Optimization (GRPO) assigns trajectory-level advantages uniformly to all policy tokens, which masks which intermediate decisions actually caused success or failure. ProVer addresses this by targeting potentially pivotal decisions for fine-grained credit assignment. An agentic judge contrasts successful and failed rollouts to propose a candidate segment responsible for divergent outcomes; ProVer then verifies that segment by sampling current-policy continuations started before and after the segment and estimating the segment’s advantage from the difference in terminal success rates. When the estimated advantage is positive, it is added into the GRPO advantages for policy tokens within that segment. Crucially, model judgment is used only to select where to verify, so local credit is grounded in observed outcome differences without exhaustively evaluating every intermediate state.
Across ALFWorld, WebShop, and SearchQA, ProVer produces the strongest average performance at both model scales evaluated, yielding relative improvements over GRPO of 9.91% for Qwen3.5-2B and 7.12% for Qwen3.5-4B. Further analyses show that informed segment selection improves policy training with modest extra generation overhead and remains beneficial even when the judge model is not frontier-scale. The approach demonstrates an effective and efficient way to assign credit to pivotal decisions in agentic reinforcement learning by selectively verifying and rewarding impactful policy segments.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.