GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

Abstract

This study evaluates Group Relative Policy Optimization (GRPO) with verifiable rewards across multiple base models, training languages, and reasoning-language rewards. Native-language reasoning often performs close to English reasoning, and training in one language can improve performance in others. Outcomes depend on the model and language, with some training settings causing severe regressions in other languages’ out-of-domain capabilities. These findings highlight the need for broad multilingual evaluation.

Publication
Preprint 2026
Konstantin Dobler
Konstantin Dobler
Ph.D. Student in ML & NLP

I’m an ELLIS Ph.D. student at Hasso Plattner Institute.