Phantom evidence: why generative AI can make scientific results look true when they are not
Researchers warn that modern generative artificial intelligence (AI) can produce outputs that look highly convincing but carry little real evidence. The paper calls this gap “phantom evidence.” It shows how a system’s ability to produce persuasive-looking artifacts can be much larger in our imagination than in what the system actually reaches. That mismatch makes single convincing results poor evidence unless we control for how often similar convincing outputs appear when the claimed effect is absent.
The authors formalize the problem with two simple ideas. T stands for the claim that an artifact actually achieves its target (for example, a reported experimental effect is real). C stands for the artifact looking convincing to a human judge. The paper measures evidence with a likelihood ratio: how much more likely C is if T is true than if T is false. Phantom evidence is tied to the gap between the nominal set of possible outputs N (what we think could happen) and the small effective set keff the system actually produces. The authors pack that gap into a single quantity, written as Δ = log(N/keff).
At a high level the danger comes from the denominator of the likelihood ratio. If generative AI can cheaply produce many convincing fakes, the probability that something looks convincing when the target is absent rises. Meanwhile the chance that a genuine target looks convincing often saturates near 1. That makes the ratio collapse toward 1 and the observation carry almost no evidence. The paper uses familiar examples to show the pattern: the horse “Clever Hans” that read unconscious human cues; large language model abstracts that fooled human reviewers (one study found about 32% of ChatGPT-generated abstracts were mistaken for real ones); and photorealistic image reconstructions that look right even when the original signal was weak.