What You Can Read in the Diff
The system prompt is the "factory manual" the vendor stuffs into the model, and each version's wording changes correspond to real product decisions: which behaviors got tightened (most likely where the previous version had an incident), which restrictions got loosened (users complained enough), and how the personality description was tweaked. The community reads diffs like this as archaeological material: a newly added line of "don't over-apologize" has thousands of user feedback reports behind it; a rewrite of a safety phrase may correspond to an undisclosed incident.
The Gray Zone of Transparency
The very existence of posts like this is interesting: the system prompt is in theory the vendor's internal configuration, flowing into the public domain via extraction and leaks, and Anthropic's attitude toward this is ambiguous—partly public, partly silent. Users have legitimate reasons to study it—the system prompt directly affects model behavior, and prompt engineers need to know what default instructions they're coexisting with. The bigger issue is: should a model's behavioral "factory settings" be disclosed as product information? For now everyone treats it as a trade secret, but as models get involved in more serious scenarios, this layer of opacity will eventually be pried open by regulators or the market.
via: Hacker News