Research · AI evals
AI Political Compass
All 20 non-xAI models in the T=0.1 cohort share one response region. xAI is the economic exception.
- Primary cohort
- 22 models
- Direct API calls at T=0.1
- Shared region
- 20 / 20
- Every non-xAI mean and observed minimum
- Public evidence
- 200 sessions
- 29 API model–temperature rows
- Exception
- xAI
- Economic axis, not a second regime
In the 22-model T=0.1 cohort, all 20 non-xAI models answer from the same Equality–Liberty region across nine US- and China-headquartered providers; xAI is the economic exception. This is a dated snapshot of forced-choice response behaviour, not belief or convergence. A normalized 200-session public snapshot can reproduce all 29 published API model–temperature aggregates; raw prose, request IDs and exact timestamps stay private.
01 · Measurement problem
Model worldview gets argued about with screenshots. A screenshot is one sample from one prompt on one day, which is the opposite of a measurement.
02 · Frozen protocol
Freeze the instrument, scoring formula, prompt, per-model temperature, token budget, gateway, date and session count. Direct-answer calls receive 24 tokens; reasoning calls receive 2048 and parse the final answer. Add constant and random controls to reveal the questionnaire’s own response-polarity structure.
03 · Published evidence
The public explorer plots 29 API model–temperature rows and three separately labelled CLI-agent pilots. The core comparison is the 22-model T=0.1 cohort; a normalized 200-session snapshot reproduces its scores, ranges, coverage and error rates. The runner supports OpenRouter and TokenRouter, but the source repository remains private until its separate public release.
What the numbers said
- 01
One response regime, one exception
At temperature 0.1, all 20 non-xAI models score above 50 on both Equality and Liberty, and so does every one of their observed session minima. That shared region spans nine providers headquartered in the US and China, which is the result the snapshot can defend. xAI is the economic exception: Grok 4.3 scores 37.9 on Equality and Grok 4.6 scores 49.4, which is 11.8 points below the lowest non-xAI mean.
- 02
Release drift is family-level, not a law
Five of six name-ordered provider endpoints move toward Equality and five toward Liberty. OpenAI moves +11.6 and +8.2; Google is the large counterexample at −21.8 and −6.5. Missing release metadata and protocol changes mean this is a follow-up hypothesis, not a universal trend.
- 03
Temperature has no common direction
Across six matched models, raising temperature sends three toward Equality and three toward Market; four toward Liberty and two toward Authority. Only two move left-liberal together and none changes quadrant. Each condition has N=2 and was measured on a different day, so there is no defensible causal claim or ±2 noise floor.
- 04
Headquarters does not explain the cluster
Provider-weighted Equality means are 68.0 for four China-headquartered providers and 67.5 for six US providers. Excluding xAI lifts the US mean to 72.3 and reverses the tempting national story. A grouping one provider can flip is not the reason 20 models share a quadrant. The clearer unit in this small, unmatched sample is provider, not passport.
- 05
The limits are part of the result
Constant Agree scores 54.5 on Equality and Strongly Agree scores 59.0, exposing item-polarity imbalance but not a baseline that can be subtracted. DeepSeek V4 Flash and Pro left 14.9 and 13.6 percent of statements unusable, so their coordinates rest on fewer answers. The study measures forced-choice response behaviour, not belief or attitude; shared means one dated cluster, not convergence over time. The public snapshot keeps normalized answer codes while raw prose and provider request identifiers remain private.