LLM

Simpler is Better: Finding the Best Reward Function in Long Chain-of-Thought Reinforcement Learning for Small Language Models

Luning Wang*, Zichen Zhang*, Junkuan Liu*.

Reward functions decide whether a reasoning model is pushed toward correct answers, shorter chains, or both. We trained Qwen2.5-3B with GRPO using normal, cosine, and dynamic rewards, asking whether length penalties that help larger models also work for small ones. In our runs, the normal reward steadily improved MATH500 and GSM8K accuracy. Cosine and dynamic rewards shortened the reasoning but made accuracy worse and less stable.
Simpler is Better: Finding the Best Reward Function in Long Chain-of-Thought Reinforcement Learning for Small Language Models