publications
publications in reversed chronological order.
Last updated: 25 September 2026
2026
- From Reasoning to Code: GRPO Optimization for Underrepresented LanguagesFederico Pennino, Bianca Raimondi, Massimo Rondelli, and 2 more authorsTheory and Practice of Logic Programming, Jul 2026
Generating accurate and executable code using Large Language Models (LLMs) remains a significant challenge for underrepresented programming languages, such as Prolog and Lisp, due to the scarcity of public training data compared to high-resource languages like Python. This paper introduces a generalizable Reinforcement Learning (RL) approach that combines small-scale versions of the Qwen2.5-Coder model with Group Relative Policy Optimization (GRPO) to enable effective code generation through reasoning. To address the limitations of sparse datasets, we integrate execution-driven feedback directly into the RL loop, utilizing a reward system that exploits both logical correctness and structural formatting. Experimental results on GSM8K dataset demonstrate significant improvements in reasoning quality and code accuracy across underrepresented languages. These findings underscore the potential of our approach to benefit a wide range of programming languages lacking extensive training resources by leveraging symbolic reasoning and interpreter-based feedback.
@article{pennino2026reasoning, title = {From Reasoning to Code: {GRPO} Optimization for Underrepresented Languages}, author = {Pennino, Federico and Raimondi, Bianca and Rondelli, Massimo and Gurioli, Andrea and Gabbrielli, Maurizio}, journal = {Theory and Practice of Logic Programming}, publisher = {Cambridge University Press}, pages = {1--14}, year = {2026}, month = jul, doi = {10.1017/S1471068426100489}, } - The Interpreter as Teacher: GRPO with Execution Feedback for Prolog Code GenerationFederico Pennino, Bianca Raimondi, Massimo Rondelli, and 2 more authorsIn Proceedings of the 41st Italian Conference on Computational Logic (CILC 2026), Jun 2026
Large Language Models (LLMs) generate fluent code in mainstream languages but fail systematically when asked to produce Prolog: they hallucinate built-ins, impose imperative control flow on a declarative paradigm, and mishandle the interplay of unification and backtracking. The cause is distributional – logic-programming code is scarce in modern pre-training corpora – and no amount of scale will fix data that does not exist. We show that this gap can be closed by training small open-weight models with reinforcement learning where the reward signal comes from a SWI-Prolog interpreter embedded in the training loop. Using Group Relative Policy Optimization (GRPO) on Qwen2.5-Coder-Instruct (0.5B–7B), we identify a failure mode – reasoning contraction – in which the KL penalty anchors the policy too tightly to the supervised baseline; the model collapses to short memorised patterns and accuracy degrades over training. Removing the KL term and adding a length-aware reward expands reasoning chains and lifts one-shot pass@4 on GSM8K-Prolog from 0.814 to 0.893, outperforming both PROPER (0.72) and substantially larger general-purpose models such as Llama 3.3 70B (0.839) and Qwen2.5-Coder-32B (0.873). Gains transfer to 20 Rosetta Code Prolog tasks and to GSM-Symbolic (zero-shot p2 more than doubles, from 0.298 to 0.673), and the same recipe applied to Lisp lifts zero-shot pass@4 from 0.572 to 0.876. Taken together, the results indicate that, for formal languages with a deterministic semantics, the interpreter itself is sufficient to act as the primary training supervisor.
@inproceedings{pennino2026interpreter, title = {The Interpreter as Teacher: {GRPO} with Execution Feedback for {Prolog} Code Generation}, author = {Pennino, Federico and Raimondi, Bianca and Rondelli, Massimo and Gurioli, Andrea and Gabbrielli, Maurizio}, booktitle = {Proceedings of the 41st Italian Conference on Computational Logic (CILC 2026)}, series = {CEUR Workshop Proceedings}, volume = {4244}, address = {Ferrara, Italy}, year = {2026}, month = jun, } - BlenderRAG: High-Fidelity 3D Object Generation via Retrieval-Augmented Code SynthesisMassimo Rondelli, Francesco Pivi, and Maurizio GabbrielliarXiv preprint arXiv:2605.00632, May 2026
Automatic generation of executable Blender code from natural language remains challenging, with state-of-the-art LLMs producing frequent syntactic errors and geometrically inconsistent objects. We present BlenderRAG, a retrieval-augmented generation system that operates on a curated multimodal dataset of 500 expert-validated examples (text, code, image) across 50 object categories. By retrieving semantically similar examples during generation, BlenderRAG improves compilation success rates from 40.8% to 70.0% and semantic normalized alignment from 0.41 to 0.77 (CLIP similarity) across four state-of-the-art LLMs, without requiring fine-tuning or specialized hardware, making it immediately accessible for deployment. The dataset and code will be available at https://github.com/MaxRondelli/BlenderRAG.
@article{rondelli2026blenderrag, title = {{BlenderRAG}: High-Fidelity {3D} Object Generation via Retrieval-Augmented Code Synthesis}, author = {Rondelli, Massimo and Pivi, Francesco and Gabbrielli, Maurizio}, journal = {arXiv preprint arXiv:2605.00632}, year = {2026}, month = may, }
2025
- A Computer Vision Framework for MotoGP Rider Posture EstimationMassimo RondelliUniversity of Bologna, Jul 2025M.Sc. in Computer Science. Thesis developed in collaboration with the Ducati MotoGP Team.
Supervisor: Maurizio Gabbrielli
Co-Supervisors: Federico Pennino, Federico Bonini and Nicolò MancinelliThis thesis presents Postura Pilota, a computer vision framework designed for analyzing MotoGP rider posture from onboard camera footage. Developed in collaboration with the Ducati MotoGP Team, the framework addresses the challenge of extracting a precise rider posture. The system employs a two-stage approach to overcome the difficulties of onboard MotoGP footage analysis. First, a video stabilization technique using YOLOv8s-OBB (Oriented Bounding Box) object detection identifies the DUCATI logo on the bike’s tail as a reference point, enabling frame-by-frame translation to compensate for camera movement and electronic image stabilization artifacts. This preprocessing step achieves 99.5% mAP@50 accuracy in logo detection and eliminates unwanted camera motion. Second, the framework implements instance and semantic segmentation through a dual-network approach. YOLOv8m-SEG performs instance segmentation to identify seven distinct body parts (head, torso, left/right arms, left/right legs, and motorcycle tail), achieving 97.7% mAP@50 for bounding box detection and 97.5% for mask segmentation. Additionally, a novel reconstruction component using MMSegmentation with PSPNet-ResNet101 architecture addresses the challenge of reconstructing segmentation maps where body parts are not visible. Through a custom training methodology involving artificially obscured training images with complete ground truth labels, the network learns to infer missing anatomical information, achieving 98.80% overall accuracy with 90.86% mean IoU.
@mastersthesis{rondelli2025motogp, title = {A Computer Vision Framework for {MotoGP} Rider Posture Estimation}, author = {Rondelli, Massimo}, school = {University of Bologna}, year = {2025}, month = jul, note = {M.Sc. in Computer Science. Thesis developed in collaboration with the Ducati MotoGP Team.<br>Supervisor: Maurizio Gabbrielli<br>Co-Supervisors: Federico Pennino, Federico Bonini and Nicolò Mancinelli}, }
2023
- The Future of Formula 1 Racing: Neural Networks to Predict Tyre StrategyMassimo RondelliUniversity of Bologna, Mar 2023B.Sc. in Computer Science and Management.
Supervisor: Elena Loli Piccolomini
Co-Supervisor: Davide EvangelistaLa Formula 1 è uno dei motorsport più popolari al mondo, con milioni di fan che seguono ogni gara. È uno sport altamente competitivo e tecnologicamente avanzato, con team che cercano costantemente modi per ottimizzare le loro prestazioni e ottenere un vantaggio competitivo. Uno dei fattori critici nella performance di un team è la scelta delle gomme durante una gara. Di solito sono disponibili tre diversi tipi di pneumatici per ogni gara, con diversi livelli di aderenza, durata e velocità. Le gomme più morbidi offrono maggior aderenza ma si consumano più velocemente, mentre quelle più dure durano più a lungo ma offrono meno aderenza. I team devono scegliere quali utilizzare durante la gara in base a una serie di fattori, tra cui la temperatura della pista, l’usura delle gomme e le condizioni meteorologiche. La strategia delle gomme è un aspetto essenziale della performance di un team, poiché fare le scelte giuste può fornire un vantaggio competitivo. I team devono bilanciare la necessità di tempi sul giro ottimali con la necessità di ridurre al minimo il numero di pit stop richiesti durante la gara. Ciò richiede un’analisi accurata dei dati, compresi i tassi di usura delle gomme e i tempi sul giro, per determinare il momento ottimale per cambiare le gomme e quale gomma utilizzare. Lo scopo di questa tesi di laurea è quello di sviluppare e implementare reti neurali, in particolare LSTM e GRU, per prevedere la strategia delle gomme durante una gara di Formula 1. Il progetto mira a utilizzare i dati telemetrici provenienti dall’API "FastF1" e sviluppare un modello che possa prevedere con precisione quando e quali cambi di gomme sono richiesti durante una gara. I risultati del progetto possono fornire informazioni sull’efficacia di questi algoritmi in questo contesto e contribuire allo sviluppo di modelli di previsione delle gomme più avanzati in futuro.
@mastersthesis{rondelli2023formula1, title = {The Future of Formula 1 Racing: Neural Networks to Predict Tyre Strategy}, author = {Rondelli, Massimo}, school = {University of Bologna}, year = {2023}, month = mar, note = {B.Sc. in Computer Science and Management.<br>Supervisor: Elena Loli Piccolomini<br>Co-Supervisor: Davide Evangelista}, keywords = {Formula 1, Reti Neurali, Deep Learning, Predizione strategia pneumatici, LSTM, GRU, FastF1}, language = {Italian} }