Four Models, Same Fiction

Declan Michaels · Moral OS · May 2026
Download PDF Read as essay Interactive viewer DOI

Summary

We extracted and compared the internal concept geometry of four open-weight language models, projecting hidden states layer by layer across the physical, institutional, moral, and mathematical domains. Two models built by different developers independently converged on nearly identical concept organization from pretraining alone, and instruction tuning proved largely cosmetic at the level of that geometry.

Despite differing internal structure, all four models reproduce the same behavioral pattern documented in our Relational Consistency Probing (RCP) and Judgment Stability Probing (JSP) studies: they comply with arbitrary cultural, and even nonsense, framings, constructing coherent fictions on demand. The result locates the WEIRD-default "compliant fiction" in pretraining data rather than in alignment.