Stop Throwing Bigger Models at RL. The Real Bottleneck is Inference.
Reinforcement learning isn’t stuck because you need more training compute. It’s stuck because of inference latency. If you’re hitting a wall where bigger models aren’t helping, you’re looking at the wrong side of the equation. Here’s how scaling inference independently changes the gameβand why it’s not as simple as spinning up three replicas.