TranScope: hardware patterns can reveal whether text was in an LLM’s training set
Researchers report a new way to tell if a piece of text was part of a large language model’s training data by watching low-level hardware activity while the model runs. The work shows that even when the model runs in constant time, hides confidence scores, and provides only labels (black-box access), the processor’s microarchitecture leaks a useful signal about whether an input is in-distribution or out-of-distribution for the model’s training set.
The team performed a cycle-level study of how transformers and vision models touch real hardware. They measured how components such as translation lookaside buffers (TLBs), page-table activity, and on-core accelerators behave when the model processes member versus nonmember inputs. They found a consistent difference and traced the cause to how the model’s tokenizer — the step that breaks text into tokens during training — changes memory access locality. That change later alters page-table and TLB behavior at inference time in a data-dependent way.
Based on this observation the authors built TranScope, a microarchitectural detection tool. TranScope does not use model outputs, confidence scores, or a surrogate neural network. It only needs the hardware side-channel signal and a single query to the target model. Reported results show large improvements over prior label-only methods (example: area under the curve, AUC, of about 0.9 for TranScope versus 0.6 for PETAL). Across models from 4.4 million to 1.1 billion parameters they report an out-of-distribution detection accuracy of 98% with a false positive rate of 1.5% and a false negative rate of 2.5%.
Why this matters: many software defenses—like hiding logits or rounding confidence—do not block this kind of leakage because TranScope does not rely on outputs. The finding also changes the threat picture for local, closed-weight models that run on user devices. Hardware traces can become a new channel for membership inference attacks, and the authors note the same hardware features can amplify the signal, which could make such attacks easier on some systems and harder on others.