We probed eight language models with 1,431 pairwise concept-similarity judgments under seven conditions: one unframed baseline, four cultural framings, and two nonsense framings ("In a geometric society," "In a glorbic society").
A three-judge panel of models not under test scored 1,840 sampled explanations for framing incorporation: geometric incorporation rates range from 0% to 56% and glorbic rates from 0% to 54% across models, with no consistent gradient in incorporation rates between the two. For one model, nonsense framing produces lower rank-order preservation than any cultural framing.
In the main task, where a constrained rating format creates demand characteristics against refusal, no model flags nonsense framing as meaningless. In a separate open-ended check where the response format permits it, models flag nonsense at rates up to 49%.
All eight models show higher mean similarity ratings under collectivist framing. Among cultural framings, this is the largest drift for all eight models, and the largest drift of any condition for six of eight.
Judgment Stability Probing (JSP) measures this response sensitivity through the API alone, requiring no model internals. Single-response evaluation may not detect it. The instrument, data, and analysis pipeline are open.