EN
Research Preprint

From Likelihood Shifts to Defensible Claims

A formal and practical audit of fixed-target likelihood measures: what they identify, where common normalizations fail, and what evidence is needed for defensible RAG claims.

Barak Or
ArtificialGate Ltd.
2026
RAG Evaluation

Abstract

Likelihood shifts are increasingly used to ask whether context affects a language model, but the resulting numbers are easy to overinterpret. We study the common case in which an evaluator teacher-forces one fixed reference and compares its likelihood across serialized inputs. The score is a candidate-prefix log-ratio: it identifies a change in support for one specified token event, not a change in predictive entropy, a decoded answer, or a preference between alternatives. We prove two practical failure results. Normalizing the shift by query-only surprisal is ill-conditioned near a confident baseline, while clipping erases every negative magnitude. In a conflict condition, scoring only the original reference cannot identify whether the model favors the replacement. We turn these observations into a four-level claim guide, a seven-item reporting checklist, and an explicit-token reference scorer with boundary tests. An audit of an archived three-checkpoint question-answering study then shows how apparently strong behavioral conclusions narrow once the scored event and evidence record are made explicit.

First page preview of From Likelihood Shifts to Defensible Claims
Read the full paperOpen the PDF in a full-screen viewer for clear, comfortable reading.

Open PDF full screen