Research connects in context learning with policy gradient methods

Paste scored outputs into a prompt and many models climb. That trick powered OPRO-style prompt search and self-rewarding loops, but the theory lagged the demos. A COLM 2026 paper from Masahiro Kaneko and Timothy Baldwin (MBZUAI) argues the mechanism is not magic few-shot mimicry. Score-conditioned in-context learning can act like an implicit policy gradient step.
If that mapping holds, iterative self-improvement without weight updates starts looking like lightweight RL at inference time.
The structural link from score-conditioned ICL to REINFORCE
Classic ICL theory mostly explains supervised $(x, y)$ pairs: transformers as gradient descent, ridge regression, Bayesian inference. Score-conditioned setups are different. The input $x$ is fixed. Context is a bag of self-generated outputs paired with scalar scores $(y, r)$. The model should shift toward high-$r$ behavior, not copy a labeled answer key.
The authors give a constructive proof: under specific weight configurations, linear self-attention aggregates output embeddings weighted by their scores, matching the score-weighted term in REINFORCE. Softmax attention becomes Boltzmann weighting, so high scores dominate exponentially. With mean-centered rewards and a simplified identity output map, the hidden-state update matches REINFORCE exactly; more generally it is a structural analogy, not a claim that every pretrained head implements the textbook matrices.
Scope is careful. This is an existence result plus empirical checks, not a claim that production models secretly run PPO in the residual stream.
Trust-region-style bounds on distribution shift
Policy gradient methods often add KL penalties so one update cannot yank the policy into oblivion. In a simplified output-layer model, the paper derives an exact pathwise upper bound on the KL divergence induced by a bounded attention update. Magnitude scales with learning-rate-like factors, max score, and embedding norms. The analogy is trust-region flavored: a single score-conditioned pass cannot move infinitely far.
That bound covers one step under stated assumptions. It does not promise monotonic reward gains across iterations. It does explain why practitioners see steady early climbs that plateau instead of immediate chaos.
What multi-LLM experiments measured
Open-weight runs cover Llama-3 (8B/70B), Olmo-3 (7B/32B), and Qwen3-4B; black-box checks add GPT-4o, Gemini 2.5, and Claude Sonnet 4. Tasks include prompt optimization on DROP and GSM8K, plus jailbreak rewriting on HH-RLHF and JailbreakBench, with $N=8$ scored samples for up to $T=10$ iterations.
Findings line up with the theory. Example scores positively rank-correlate with log-probability changes under ICL (often 0.5–0.7 Spearman on larger models; shuffle the scores and the correlation dies). Attention weights track scores, especially in late layers. Iterative ICL lifts normalized scores early, then flattens. Given identical samples, one-step REINFORCE and score-conditioned ICL shift $\Delta\log p$ vectors in the same direction (cosine similarity above ~0.45). Ablating score-token representations at the peak layers cuts the effect.
Why this matters for iterative self-improvement
Analysis: if attention can implement reward-weighted aggregation, then OPRO, self-refine, and self-rewarding stacks are not mysterious prompt folklore. They are approximate, KL-bounded policy updates you can run without opening the optimizer.
The stakes cut both ways. The same machinery that steers math prompts can steer jailbreaks when the score is attack success. The authors say as much in the ethics note: understanding the mechanism is how you detect score-chasing contexts. For product teams, numeric score formats work; natural-language scores weaken the signal; no-score controls barely move. Put the numbers where late layers can see them, keep batches modest, and stop expecting free lunch after the early plateau.



