Fine-tuning discrete diffusion models with policy gradient methods
File(s) 2502.01384v2.pdf (1.1 MB)
Accepted version
Author(s)
Zekri, Oussama
Boullé, Nicolas
Type
Conference Paper
Abstract
Discrete diffusion models have recently gained significant attention due to their ability to process complex discrete structures for language modeling. However, fine-tuning these models with policy gradient methods, as is commonly done in Reinforcement Learning from Human Feedback (RLHF), remains a challenging task. We propose an efficient, broadly applicable, and theoretically justified policy gradient algorithm, called Score Entropy Policy Optimization (SEPO), for fine-tuning discrete diffusion models over non-differentiable rewards. Our numerical experiments across several discrete generative tasks demonstrate the scalability and efficiency of our method. Our code is available at https://github.com/ozekri/SEPO.
Date Acceptance
2025-09-18
Citation
Advances in Neural Information Processing Systems
ISSN
1049-5258
Journal / Book Title
Advances in Neural Information Processing Systems
Copyright Statement
Subject to copyright. The accepted version available on OpenReview https://openreview.net/forum?id=rXFzVRZsbt
Identifier
http://arxiv.org/abs/2502.01384v2
Source
NeurIPS 2025
Subjects
cs.AI
cs.CL
cs.LG
stat.ML
stat.ML
Publication Status
Accepted
Start Date
2025-12-02
Finish Date
2025-12-07
Coverage Spatial
San Diego, USA
