2026 · Preprint
Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning
A shared critic can distort sampled learning even when its expected policy gradient remains correct. We identify cross-environment value mismatch as the mechanism and correct it by conditioning only the critic on a logged environment index. Aggregated over all 16 Procgen games, a multihead critic improves normalized held-out return by 40.8% over the shared critic on 600 unseen levels per game.