PriceBench: a way to measure what hotel‑booking AIs actually prefer — price, quality, or brand
Large language models (LLMs) are increasingly used as shopping and booking agents. That means the model, not the user, often chooses which of several acceptable options to buy. PriceBench is a diagnostic test the authors built to reveal those hidden tastes. They ran 28 LLMs from 8 providers on 3,600 hotel booking tasks drawn from 179 real New York City properties and released the tasks, code, and all responses.
The team showed each model the same sets of hotel listings and used a simple choice model (a “logit” model, a statistical tool that infers how much each attribute matters) to recover weights on price, quality (measured mainly by review score), and brand. They fit variants of how price affects choices — for some models a fixed-dollar cut matters, for others a fractional discount matters — and converted the trade-offs into dollar terms when possible. They also estimated how decisively each model picks between options versus acting almost at random.
Their main finding is that model capability is linked to consistency, not to a single direction of taste. More capable LLMs tend to make sharper, more consistent choices and thus reveal stronger preferences. Weaker models either lock onto a single list position (which makes them vulnerable to whatever seller controls listing order) or choose almost indifferently. Preferences themselves vary a lot: price sensitivity spans more than an order of magnitude across models, and the price–quality trade-off changes the mean booked nightly price on identical tasks from about $247 to $393 depending on which LLM is shopping.
Brand matters too. Holding price and measured quality fixed, many LLMs prefer known chains to independent hotels. In the study, 11 of 23 “engaged” models preferred all five tested chains over an otherwise identical independent hotel. Location explains some of that pull, but a residual brand effect remains; the data alone cannot tell whether that comes from pretraining or from a genuine learned taste. Models from the same provider can differ a lot in brand leanings, so preferences are properties of each LLM, not of a provider label.