Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

UNREAL uses a language model to select its own evidence

A research framework links corpus retrieval with long context reasoning.

Listen to this article

An October 6 preprint introduces UNREAL, which uses a frozen language model to select evidence and answer questions. Training updates fewer than 500000 added parameters.

UNREAL source diagram showing evidence retrieval and answer generation through the same frozen language model
Source diagram by Edan Kinderman and coauthors from Figure 2 of the UNREAL paper, converted to WebP under Creative Commons Attribution 4.0. The diagram shows retrieval and generation using one model.

What the results measure

The paper reports a best HotpotQA complete evidence recall of 73.2 percent against a 49.1 percent baseline. This measures whether all required passages appear in the top ten results, rather than final answer accuracy.

Tests focus on sparse evidence tasks. Transfer beyond Wikipedia training to other domains and languages remains untested. The index needs extra storage and still uses BM25 for initial context.

Speed tests used random weights and inputs, with some serving overhead omitted.

Featured chart by Edan Kinderman and coauthors from UNREAL Figure 11, under Creative Commons Attribution 4.0. White margins cropped and converted to WebP.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.