From Reasoning to Code: GRPO for Underrepresented Languages
Reinforcement learning with interpreter feedback to teach small LLMs to write Prolog
Large Language Models write fluent Python but struggle with underrepresented languages such as Prolog and Lisp, where public training data is scarce. This project introduces a generalizable reinforcement learning approach: small versions of Qwen2.5-Coder are fine-tuned with Group Relative Policy Optimization (GRPO), and the reward comes from actually executing the generated code.
How it works. The model is trained on GSM8K math word problems and must answer each one in a fixed structure:
<reasoning> ...chain-of-thought reasoning... </reasoning>
<code> ...Prolog facts and rules... </code>
<query> ...Prolog query for the answer... </query>
At training time the generated knowledge base and query are run against a live SWI-Prolog interpreter (via pyswip, with a 5-second timeout). A composite reward combines a correctness signal (the numerical answer matches the ground truth) with partial credit for respecting the output format, so the model learns both to reason and to produce valid, executable code.
Setup. Qwen2.5-Coder-Instruct at 0.5B, 1.5B, 3B and 7B parameters, trained with 4-bit quantization and LoRA (rank 32), in zero-, one- and five-shot prompting modes.
Generating accurate and executable code using Large Language Models (LLMs) remains a significant challenge for underrepresented programming languages, such as Prolog and Lisp, due to the scarcity of public training data compared to high-resource languages like Python. This paper introduces a generalizable Reinforcement Learning (RL) approach that combines small-scale versions of the Qwen2.5-Coder model with Group Relative Policy Optimization (GRPO) to enable effective code generation through reasoning. To address the limitations of sparse datasets, we integrate execution-driven feedback directly into the RL loop, utilizing a reward system that exploits both logical correctness and structural formatting. Experimental results on GSM8K dataset demonstrate significant improvements in reasoning quality and code accuracy across underrepresented languages. These findings underscore the potential of our approach to benefit a wide range of programming languages lacking extensive training resources by leveraging symbolic reasoning and interpreter-based feedback.
@article{pennino2026reasoning,title={From Reasoning to Code: {GRPO} Optimization for Underrepresented Languages},author={Pennino, Federico and Raimondi, Bianca and Rondelli, Massimo and Gurioli, Andrea and Gabbrielli, Maurizio},journal={Theory and Practice of Logic Programming},publisher={Cambridge University Press},pages={1--14},year={2026},month=jul,doi={10.1017/S1471068426100489},}