Study of 407 leaked system prompts finds they act like configuration files, not moral statements
This paper looks at 407 system prompts—the hidden instructions sent to large language models—and finds they mostly read like operational configuration, not public value statements. After merging four community collections covering 62 vendors, the authors show that most words in these prompts tell the model how to use tools, follow protocols, or format output. Only a small share of words look like explicit safety or ethical policy.
The researchers combined four independently maintained repositories: CL4R1T4S, System‑Prompts, system‑prompts‑and‑models‑of‑ai‑tools, and system_prompts_leaks, collected between 2026-06-17 and 2026-07-11. They flagged 29 near‑duplicate clusters covering 66 files so they would not silently double‑count copies. Files were hand‑labeled into six archetypes (for example: 155 chat assistants and 108 coding agents) and into three value postures (236 operational, 156 caution, 15 truth‑seeking). They also used simple, auditable text detectors and three different similarity measures—word overlap, TF–IDF scores, and sentence embeddings—to study reuse and change over time.
Quantitatively, a simple block‑level classifier assigned roughly 58% of classified words to tool or protocol instructions and about 5% to safety policy. The strictest rule lines in prompts focus much more on controlling tool use and file safety than on guarding against harmful content—by about an 11:1 margin in the observed corpus. Literal copying of text across vendors was rare and concentrated in a few cross‑vendor pairs. Prompts also show measurable maintenance debt: version chains for some vendors turn over thousands of words between releases, suggesting frequent edits and possible “prompt rot.” The paper reports concrete file lengths, for example a mean of 5,630 words for chat assistant prompts and 3,878 words for coding agents.