A formal and practical audit of fixed-target likelihood measures: what they identify, where common normalizations fail, and what evidence is needed for defensible RAG claims.
Likelihood shifts are increasingly used to ask whether context affects a language model, but the resulting numbers are easy to overinterpret. We study the common case in which an evaluator teacher-forces one fixed reference and compares its likelihood across serialized inputs. The score is a candidate-prefix log-ratio: it identifies a change in support for one specified token event, not a change in predictive entropy, a decoded answer, or a preference between alternatives. We prove two practical failure results. Normalizing the shift by query-only surprisal is ill-conditioned near a confident baseline, while clipping erases every negative magnitude. In a conflict condition, scoring only the original reference cannot identify whether the model favors the replacement. We turn these observations into a four-level claim guide, a seven-item reporting checklist, and an explicit-token reference scorer with boundary tests. An audit of an archived three-checkpoint question-answering study then shows how apparently strong behavioral conclusions narrow once the scored event and evidence record are made explicit.
