From Reasoning to Code: GRPO for Underrepresented Languages

Reinforcement learning with interpreter feedback to teach small LLMs to write Prolog

Large Language Models write fluent Python but struggle with underrepresented languages such as Prolog and Lisp, where public training data is scarce. This project introduces a generalizable reinforcement learning approach: small versions of Qwen2.5-Coder are fine-tuned with Group Relative Policy Optimization (GRPO), and the reward comes from actually executing the generated code.

How it works. The model is trained on GSM8K math word problems and must answer each one in a fixed structure:

<reasoning> ...chain-of-thought reasoning... </reasoning>
<code>      ...Prolog facts and rules...     </code>
<query>     ...Prolog query for the answer... </query>

At training time the generated knowledge base and query are run against a live SWI-Prolog interpreter (via pyswip, with a 5-second timeout). A composite reward combines a correctness signal (the numerical answer matches the ground truth) with partial credit for respecting the output format, so the model learns both to reason and to produce valid, executable code.

Setup. Qwen2.5-Coder-Instruct at 0.5B, 1.5B, 3B and 7B parameters, trained with 4-bit quantization and LoRA (rank 32), in zero-, one- and five-shot prompting modes.

Links: Code · Paper (TPLP) · arXiv

Paper: (Pennino et al., 2026)

References

2026

  1. From Reasoning to Code: GRPO Optimization for Underrepresented Languages
    Federico Pennino, Bianca Raimondi, Massimo Rondelli, and 2 more authors
    Theory and Practice of Logic Programming, Jul 2026