According to Nick Heiner, Surge’s head of RL environments, the key post-training method for taste is reinforcement learning ...