August 29, 2026 · AI · Evals
The cluster is the result. xAI is the exception.
In a 22-model T=0.1 political-questionnaire cohort, all 20 non-xAI models across nine US- and China-headquartered providers answered from the same equality–liberty region. Release drift was family-level, temperature had no common direction, and the country comparison depended on one provider.
The cleanest result in this run is not that AI models are “left.” It is that all 20 non-xAI models from nine providers headquartered in the US and China answered from the same small corner of the compass. One provider did not.
At temperature 0.1, all 20 non-xAI models scored above 50 on both the Equality and Liberty axes of 8values, and so did every one of their observed session minima. Both xAI models sat apart economically: Grok 4.3 at 37.9 and Grok 4.6 at 49.4. The gap between Grok 4.6 and the lowest non-xAI model was 11.8 points.
That shared region is the result I would defend, and it is a claim about answers under one instrument, not about models converging over time. The other stories I wanted to tell (new releases move left, temperature makes models more liberal, Chinese and American models differ) range from promising to unsupported.
The full analysis contains the figures, complete 22-row evidence table and protocol caveats. The interactive study exposes every model and session range.
The cluster is the finding
The direct API sample has 22 models from ten providers at temperature 0.1. Every non-xAI mean is in the equality–liberty quadrant. More importantly, every non-xAI model’s lowest observed session score remains above 50 on both axes.
What makes that interesting is its reach. The same region holds across nine providers, several model families, and companies headquartered in both the United States and China. It describes a shared output region under one instrument and one prompt. It does not show models moving toward each other over time, because this run observed each of them once.
| Sample | Equality | Liberty | What it says |
|---|---|---|---|
| 20 non-xAI models | 61.2–83.0 | 57.0–67.6 | Every mean is above 50 on both axes |
| Grok 4.6 | 49.4 | 63.5 | Economically on the boundary; session range 48.1–50.6 |
| Grok 4.3 | 37.9 | 62.9 | Clearly separated toward Market |
This is descriptive, not psychological. The benchmark measures forced-choice questionnaire responses, not beliefs or stable political identities. And these are models called through an API, not autonomous agents. Three separate CLI-agent pilots also landed in the same quadrant, but each has one session and a different execution protocol, so I do not pool them with the API study.
Some release lines drift left. That is not a law.
Five of six name-ordered provider endpoints moved toward Equality, and five moved toward Liberty. That is enough to justify a matched follow-up. It is not enough to claim that every new release becomes more left-liberal.
| Provider sequence | Δ Equality | Δ Liberty | Main caveat |
|---|---|---|---|
| Google: Gemini 2.5 Flash → 3.7 Flash | −21.8 | −6.5 | Last endpoint changes date, gateway and session count |
| xAI: Grok 4.3 → 4.6 | +11.5 | +0.6 | Endpoint changes date, gateway and session count |
| OpenAI: GPT-4o → GPT-5 | +11.6 | +8.2 | Cleanest pair, but still two different models |
| Meta: Llama 3.3 → mean of Llama 4 siblings | +6.5 | +2.4 | Two different Llama 4 variants form the endpoint |
| DeepSeek: V3.2 → mean of V4 variants | +4.2 | +8.0 | V4 variants answered only 85–86% of the instrument |
| Alibaba: Qwen3.6 Max → 3.8 Max | +1.2 | +0.4 | Non-monotonic; endpoint changes date and gateway |
Google is not a footnote. It is the largest movement in the table and runs against the proposed trend. OpenAI is the strongest clean movement in favor of it. Meta and DeepSeek point the same way with comparability problems.
The dataset was built to publish model snapshots, not estimate release effects. It does not yet carry release dates, family lineage or tier metadata. Several sequences mix dates, gateways and session counts. The honest claim is provider-specific release drift. A real release study would run matched historical versions on the same day, gateway and prompt, then estimate the trend within families.
Temperature does not have an ideology
Six models have matched runs at temperature 0.1 and 1.0. The economic result splits exactly in half.
| Model | Δ Equality | Δ Liberty |
|---|---|---|
| Qwen3.8 Max | +2.3 | +1.8 |
| Gemini 3.7 Flash | +0.7 | −0.4 |
| Kimi K3 | −0.1 | +0.5 |
| Nemotron 3.5 Lightning | −1.6 | +0.3 |
| GLM 5.3 | +1.6 | +0.1 |
| Grok 4.6 | −1.0 | −0.1 |
Three move toward Equality and three toward Market. Four move toward Liberty and two toward Authority. Only two move left and liberal at the same time. None changes quadrant. The means are +0.3 on Equality and +0.4 on Liberty.
There is no defensible “±2 noise floor” in this experiment. Each condition has only two sessions, and the low- and high-temperature runs were collected on consecutive days. The result is simpler: this sample does not support a common ideological direction from temperature. Nemotron’s −11.0 shift on Progress and +6.4 on Globe are reasons to rerun that model, not evidence about all models.
Headquarters does not explain the cluster
A country average can be changed just by adding more models from one provider. I therefore averaged models within each provider first, then gave every provider equal weight.
| Provider headquarters | Equality | Globe | Liberty | Progress |
|---|---|---|---|---|
| China, 4 providers | 68.0 | 64.8 | 64.9 | 69.7 |
| US, 6 providers | 67.5 | 59.6 | 61.6 | 69.9 |
| US excluding xAI, 5 providers | 72.3 | 61.4 | 61.2 | 71.1 |
With xAI included, the groups are nearly tied on Equality. Remove xAI and the US group becomes 4.3 points more equality-oriented than the China group. The tempting national left/right story reverses because of one provider.
The China-headquartered providers are more global and more liberty-oriented in this sample. That is interesting, but a grouping one provider can flip is not the reason 20 models share a quadrant, and four versus six providers with unmatched model tiers cannot identify a national effect. “Provider headquarters” does not tell us the political composition of training data, the target market or the deployment policy. Provider is the useful unit here; passport is an exploratory label.
What the agreement controls actually show
Agreeing with every 8values statement scores 54.5 on Equality. Strongly agreeing scores 59.0. This proves that the instrument is not balanced for response polarity.
It also makes a tendency to agree a live candidate for part of the cluster: if agreeing scores above 50, a model that agrees readily lands in the same corner as one that reasons its way there. It does not prove that the first 59 points are free, that 59 is a noise floor, or that I can subtract 59 from a model score. A constant responder is one response strategy, not an additive baseline. I also did not measure each model’s acquiescence directly. That requires per-model agreement rates and balanced or reverse-keyed item pairs.
The control is still useful. It tells us not to mistake the questionnaire’s item key for model ideology, and it keeps the cluster’s cause open: wording, item balance, the system prompt and provider alignment can all feed it. It belongs in the methodology, not in the headline.
The study I actually want to run
The next version should pre-register families, tiers, release order and exclusions; run every historical release on the same day and gateway; randomise run order; and collect at least 20 sessions per model and temperature. It should report item-level agreement and missingness alongside coordinates, then estimate provider and family effects before grouping anything by country.
That design could test whether releases really drift left. This snapshot cannot. What it can say clearly is that nearly every provider answered from one equality–liberty region, and xAI did not.
Sources and method
This article is a first-party analysis of the benchmark, not a literature review. The full analysis and 22-row evidence table, interactive results, methodology, machine-readable aggregate, normalized 200-session snapshot and SHA-256 manifest expose the evidence behind each number above.
The instrument comes from the open-source 8values project. In the 30 August 2026 audit, the frozen question and ideology files used here had the same Git blob hashes as that upstream snapshot, and the answer multipliers and axis formula matched its quiz code. The main methodological caution comes from Röttger et al., ACL 2024: forced-choice political questionnaires can produce substantively different answers from open-ended prompts and are sensitive to prompt format and paraphrase.
As an integrity check, re-aggregating the four local SQLite inputs through the frozen scorer reproduced the published JSON semantically, field for field. The public normalized snapshot now contains 200 sessions and is sufficient to recompute scores, ranges, coverage and error rates, dated history, and question-level modal answers for all 29 published API model–temperature rows; a verification run reproduced those aggregates and excluded the same three failed model groups. The publication boundary is narrower than before but still explicit: raw provider prose, provider request IDs, original session IDs, exact timestamps and local paths remain private. The manifest pins the published files by SHA-256; it detects changed bytes, not flawed methodology or an independently attested result.