ScholarCatalyst: a new benchmark that asks which prior papers actually inspired later computer‑science research
Scientists are often better than machines at spotting the single prior idea that helps a new project move forward. ScholarCatalyst is a new benchmark designed to study that skill. It asks whether a system can find the earlier papers that authors say truly advanced their completed work.
The team built ScholarCatalyst by asking 184 lead authors of 207 recent computer‑science papers to label which candidate prior papers did or could have helped their projects. Each chosen paper comes with an author’s written rationale. The authors’ judgments become the ground truth for a retrieval task: given an initial research question, can a system retrieve the papers that the real authors found useful, using only the literature that existed when the project began?
The paper tests several retrieval approaches. “Embedding retrieval” is a common method that turns text into vectors and finds nearest neighbors. “Agentic search” is a system that plans and calls tools, including the same embedding retriever, as part of a multi‑step search process. On the benchmark, embedding retrieval reached 0.48 Recall@20, while agentic search scored 0.42 Recall@20. Recall@20 measures how often a needed paper appears among the top 20 results. An agent built on Claude Fable 5.1 reached 0.51 Recall@20, though the authors note this model may have seen some of the completed papers during its training.
These results suggest current methods do not yet match expert intuition for finding the single prior work that sparks progress. The dataset and results point to a gap between tool‑driven search and the kind of broad, insight‑based retrieval practiced by researchers. The paper argues that new training techniques are needed to give models this expert sense of which prior idea truly matters.