I designed a black-box protocol that measures how language models' similarity judgments shift under cultural framing. The protocol uses only API access. I pre-registered the experiment on the Open Science Framework and ran it against five deployed models: Claude Sonnet 4, GPT-4o, Gemini 2.5 Flash, Grok, and Llama 3.3 70B. The inventory covered 18 concepts across physical, institutional, and moral domains under seven framings.
The pre-registered hypothesis failed. Institutional drift exceeded moral drift in all four interpretable models. The pre-registered permutation test was structurally underpowered at the chosen inventory size. Post-hoc analysis traced both failures to specific design errors that the v2 protocol corrects.
What survived: the physical control domain held with moderate-to-large effect sizes (Hedges' g = 0.54 to 5.15), confirming the protocol discriminates culturally invariant concepts from culturally loaded ones. Four distinct nonsense-compliance profiles emerged across the models. Every model except one constructed coherent moral frameworks from an irrelevant weather preamble. Three of four interpretable models default to a WEIRD-individualist judgment position under neutral prompting.
This paper reports the protocol, the pilot data, and the design errors as a package. The protocol, data, and analysis code are open-source.