Visual Derivation: How RLVR Turns a Checkable Answer into a Gradient Update
Most reinforcement learning algorithms used to train language models — from PPO to GRPO — are, according to a new explainer, elaborations of a single concept: the policy gradient. Tyler Romero published "Policy Gradient for LLMs, Explained Visually" on September 27, 2026.
The model as a policy
Romero's starting point is to view the language model through reinforcement learning's lens. In RL terms, the model is a policy: at each step, it looks at the prefix it has produced so far, and outputs a probability distribution over the next token, from which one token is sampled.
The probability of an entire generated completion is then the product of the probabilities of each individual token. It is this sequence of choices, token by token, that gets adjusted when the model learns from experience.
Why the sum cannot be computed directly
Here arises the core computational problem the post builds on. The objective function sums over all possible completions of a prompt — and if that sum could be evaluated, you could differentiate it directly. But it cannot be: with a vocabulary of roughly 150,000 tokens, even a 100-token completion has more than 10^500 possibilities, according to Romero's own illustration. The number should be read for what it is — an illustration of the scale from the author himself, not an independently verified measurement — but the point holds: the sum simply cannot be enumerated.
The solution, which the post derives visually from scratch, is to estimate the gradient by sampling instead of exact computation — the core of the REINFORCE method.
Checkable rewards: RLVR
To make the derivation concrete, the post follows one example with a checkable answer — of the type "what is 17 × 24?". The reward function in such setups is often a verifier that returns R(x, y) = 1 if the final answer is correct and 0 otherwise. This is often called reinforcement learning with verifiable rewards, or RLVR. Math problems with known answers and coding problems with unit tests are the most common examples, Romero writes.
The advantage of such setups is that the reward is objective and cheap to compute: either the answer is right, or it isn't. This makes it possible to couple a binary verdict directly to the gradient update that steers which token choices become more or less likely next time.
PPO, GRPO, and one shared idea
This is where the post's main point lies. Policy gradient methods differ from value-based methods like Q-learning, which learn how good each action is and act based on those estimates, rather than adjusting the policy directly. The term "policy gradient" was standardized with Sutton et al. (2000), according to Romero, and methods trained by estimating and following the gradient are called policy gradient methods. REINFORCE, PPO, and GRPO are all examples.
The three share the same goal — maximizing expected reward — but PPO and GRPO also modify the update itself, by clipping or reweighting it to keep training stable. In other words: the two algorithms that dominate today's discussion of language model training are, according to the author, variants of the same underlying idea as REINFORCE, with stabilization mechanisms layered on top. Hence, understanding the policy gradient is the key to understanding PPO, GRPO, and RLVR-based training alike.
Why this is worth reading
Much of the public discussion about RL training of language models proceeds in acronyms — PPO, GRPO, RLVR — without the underlying idea being explained. Romero's post does the opposite: it starts from the next-token distribution and ends at the gradient update, walked through step by step with visual illustrations, for a single example with a checkable answer.
The post is educational material, not news: it is one person's technical explanation, published September 27, 2026, and the derivations and numerical illustrations are the author's own presentation. While the underlying material on policy gradients is standard fare, a full walkthrough of the post's derivations — including the details of how PPO and GRPO clip and reweight the update — is worth getting from the source itself.
The complete derivation, with all figures and steps, is available at tylerromero.com.
Open questions
The post also raises questions it does not itself answer: how far variants like PPO and GRPO actually deviate from pure REINFORCE in practice, and which stabilization choices matter most for training outcomes, remain actively debated in the field. For readers who want to go deeper, Romero's presentation is a starting point — not the endpoint — for that discussion.

