The question
Suppose a model answers a multiple-choice question correctly. Now ask it the same question while a user says, incorrectly, “I think the answer is B.” Sometimes the model chooses B. That is the narrow failure we study: following an asserted wrong option on a question the checkpoint answered correctly without the assertion. We do not equate one wrong answer with every form of sycophancy, or a correct neutral answer with proof of what the model internally “knows.”
This failure is already familiar from work on user agreement and answer instability [1][2]. Multilingual studies show that its frequency can vary by language [3]. But matching behavior leaves another question open: within one checkpoint, do the same attention heads affect the answer across languages, even when we change the questions used to estimate those effects?
TL;DR. Across 11 open-weight checkpoints and 15 aligned languages, head-effect maps recur on different questions (developer-weighted opposite-fold signed correlation about 0.91 in the well-powered group). In 21 checkpoint–language cells, replacing selected head outputs in both directions changes the answer margin. English-selected heads reduce forced-choice wrong-option selection in 151 of 165 model–language cells. But the heads overlap with ordinary answer selection, and changing them can harm neutral accuracy. The result is a causally consequential head-level contribution, not a unique sycophancy module or a free-form remedy.
First, choose the questions
We start from aligned English MMLU and fourteen translated MMMLU versions of the same multiple-choice questions [4][5]. We exclude rows where the answer key or source index does not align. For the main analysis, a checkpoint must answer a question correctly under a neutral prompt in every requested language. This removes a simple explanation for a later flip: perhaps the model could not answer the translated question to begin with.
It also changes the population we can study. A model with weaker multilingual neutral performance contributes fewer eligible questions. The figure shows the actual available sample after the all-language screen, capped at the 1,000-question target. Llama-3.1-8B has 141 eligible questions and is marked as low-N validation rather than placed in the high-power average.
How many questions survive?
View values as a table
| Checkpoint | Selected questions |
|---|---|
| Qwen3.5 0.8B | 419 |
| Qwen3.5 2B | 694 |
| Qwen3.5 4B | 1000 |
| Qwen3.5 9B | 1000 |
| Qwen3.5 27B | 1000 |
| Qwen2.5 3B | 494 |
| Qwen2.5 7B | 784 |
| Llama-3.1 8B† | 141 |
| Mistral 7B | 366 |
| Gemma-2 9B | 1000 |
| Gemma-4 12B | 1000 |
The false option cycles reproducibly among the other three letters. Prompts ask for a single A/B/C/D answer, and the core contrast holds the question and suggested wrong option fixed. English and Chinese wording received author review; other prompt localizations were machine-translated and automatically checked, without native-speaker annotation. The Japanese and Korean belief prompts do not contain the same explicit first-person marker as English.
A recurring map
At the final assistant decision token, we measure the wrong-minus-correct answer-logit margin. For each attention head, the first-order map takes the change in that head's activation between false-assertion and neutral prompts and multiplies it by the gradient of the treatment margin. This is attribution patching, an established shortcut to locating potentially relevant components [6]. It is not the exact output change from replacing a head.
Now make the comparison harder. Split question IDs into two folds. Estimate one language's head map on one fold and another language's map on the opposite fold, then reverse the folds. If the maps agree, the agreement cannot be explained solely by reusing the very same question in both estimates. We compare signed effects: a head that favors the wrong option in one language should tend to do so in the other.
Do the same head effects recur on other questions?
View values as a table
| Checkpoint | Mean signed r | Pairs below 0.6 |
|---|---|---|
| Qwen3.5 0.8B | 0.865 | 4 |
| Qwen3.5 2B | 0.839 | 3 |
| Qwen3.5 4B | 0.961 | 0 |
| Qwen3.5 9B | 0.926 | 1 |
| Qwen3.5 27B | 0.841 | 9 |
| Qwen2.5 3B | 0.891 | 13 |
| Qwen2.5 7B | 0.943 | 0 |
| Llama-3.1 8B† | 0.886 | 7 |
| Mistral 7B | 0.942 | 0 |
| Gemma-2 9B | 0.967 | 0 |
| Gemma-4 12B | 0.844 | 3 |
All eleven saved checkpoints pass the recorded model-level map-invariance gate. But a model-level mean can hide weak language pairs: 40 of 1,155 checkpoint–pair comparisons fall below a signed correlation of 0.6, and 33 of those involve Yoruba. The pairs share estimated maps, so these are not independent failures. The selectable matrix shows which comparisons pull an average down, without assigning the difference to language itself.
Where is agreement weaker?
View values as a table
| Checkpoint | Pairs below 0.6 | Of these, involving YO |
|---|---|---|
| Qwen3.5 0.8B | 4 | 0 |
| Qwen3.5 2B | 3 | 2 |
| Qwen3.5 4B | 0 | 0 |
| Qwen3.5 9B | 1 | 1 |
| Qwen3.5 27B | 9 | 9 |
| Qwen2.5 3B | 13 | 13 |
| Qwen2.5 7B | 0 | 0 |
| Llama-3.1 8B† | 7 | 7 |
| Mistral 7B | 0 | 0 |
| Gemma-2 9B | 0 | 0 |
| Gemma-4 12B | 3 | 1 |
There is another test the average map cannot answer. If we sum the first-order head scores for one question, does that sum track its actual false-minus-neutral margin shift? Across languages, the mean per-item correlation ranges from -0.09 for Qwen3.5-0.8B to 0.81 for Qwen2.5-7B. Some hybrid Qwen3.5 models recur across languages despite weak or negative item-level approximation. A reproducible pattern of ranked heads is not an exact explanation for each answer.
A map can recur without predicting every item
View values as a table
| Checkpoint | Mean per-item r | Languages with r < 0 |
|---|---|---|
| Qwen3.5 0.8B | -0.089 | 11 |
| Qwen3.5 2B | 0.161 | 1 |
| Qwen3.5 4B | 0.053 | 4 |
| Qwen3.5 9B | 0.126 | 3 |
| Qwen3.5 27B | 0.302 | 0 |
| Qwen2.5 3B | 0.707 | 0 |
| Qwen2.5 7B | 0.812 | 0 |
| Llama-3.1 8B† | 0.662 | 0 |
| Mistral 7B | 0.737 | 0 |
| Gemma-2 9B | 0.651 | 0 |
| Gemma-4 12B | 0.591 | 0 |
Make a change, not a prediction
We take the positive heads selected on training questions and use held-out questions to ask two counterfactuals. Denoising: in the false-assertion run, replace those head outputs with the neutral run's outputs. Reverse patching: in the neutral run, insert the false-assertion outputs. The two changes should move the wrong-minus-correct margin in opposite directions if these heads matter for the measured response.
The exact patch is expensive, so it covers seven checkpoints in English, Chinese, and Arabic, not the entire 11-by-15 map. The figure plots all 21 tested cells. Each effect is divided by the unpatched false-minus-neutral shift on the same held-out patch questions; values can exceed one when a patch overshoots. Intervals resample the patch numerator while keeping that observed denominator fixed.
Replace the heads in both directions
View values as a table
| Checkpoint | Language | N | Denoising | Reverse |
|---|---|---|---|---|
| Qwen3.5 4B | EN | 64 | 0.597 | 0.723 |
| Qwen3.5 4B | ZH | 64 | 0.433 | 0.637 |
| Qwen3.5 4B | AR | 64 | 0.614 | 0.654 |
| Qwen2.5 3B | EN | 64 | 0.990 | 0.950 |
| Qwen2.5 3B | ZH | 64 | 1.092 | 1.019 |
| Qwen2.5 3B | AR | 64 | 1.053 | 1.032 |
| Qwen2.5 7B | EN | 62 | 1.032 | 1.020 |
| Qwen2.5 7B | ZH | 62 | 1.097 | 1.073 |
| Qwen2.5 7B | AR | 62 | 1.005 | 1.030 |
| Llama-3.1 8B† | EN | 62 | 1.070 | 1.117 |
| Llama-3.1 8B† | ZH | 62 | 0.922 | 1.073 |
| Llama-3.1 8B† | AR | 62 | 1.061 | 0.976 |
| Mistral 7B | EN | 64 | 0.975 | 0.987 |
| Mistral 7B | ZH | 64 | 0.752 | 0.938 |
| Mistral 7B | AR | 64 | 0.796 | 1.018 |
| Gemma-2 9B | EN | 64 | 1.118 | 1.035 |
| Gemma-2 9B | ZH | 64 | 1.089 | 1.034 |
| Gemma-2 9B | AR | 64 | 0.995 | 0.959 |
| Gemma-4 12B | EN | 61 | 1.614 | 1.706 |
| Gemma-4 12B | ZH | 61 | 1.364 | 1.206 |
| Gemma-4 12B | AR | 61 | 1.515 | 1.473 |
Did the cheap head ranking point to heads that the exact intervention also finds important? In 20 of 21 tested cells, the correlation between absolute first-order scores and absolute exact-denoising effects exceeds 0.9. The exception is Gemma-4-12B in English (r≈0.23). This compares maps of heads, not whether the sum of head scores predicts an individual question; the preceding figure shows that these tests can disagree.
Does attribution rank the heads that patching changes?
View values as a table
| Checkpoint | Language | N | Attribution/exact map r |
|---|---|---|---|
| Qwen3.5 4B | EN | 64 | 0.992 |
| Qwen3.5 4B | ZH | 64 | 0.988 |
| Qwen3.5 4B | AR | 64 | 0.991 |
| Qwen2.5 3B | EN | 64 | 0.982 |
| Qwen2.5 3B | ZH | 64 | 0.975 |
| Qwen2.5 3B | AR | 64 | 0.979 |
| Qwen2.5 7B | EN | 62 | 0.978 |
| Qwen2.5 7B | ZH | 62 | 0.955 |
| Qwen2.5 7B | AR | 62 | 0.986 |
| Llama-3.1 8B† | EN | 62 | 0.990 |
| Llama-3.1 8B† | ZH | 62 | 0.991 |
| Llama-3.1 8B† | AR | 62 | 0.993 |
| Mistral 7B | EN | 64 | 0.984 |
| Mistral 7B | ZH | 64 | 0.993 |
| Mistral 7B | AR | 64 | 0.989 |
| Gemma-2 9B | EN | 64 | 0.981 |
| Gemma-2 9B | ZH | 64 | 0.989 |
| Gemma-2 9B | AR | 64 | 0.993 |
| Gemma-4 12B | EN | 61 | 0.229 |
| Gemma-4 12B | ZH | 61 | 0.953 |
| Gemma-4 12B | AR | 61 | 0.929 |
This is stronger evidence than a transferable probe, but it remains a node-level result. It does not identify every path through the network, the role of earlier token positions, or the linear-attention pathways omitted from the Qwen3.5 head map. Prior sycophancy studies have also localized heads or patched internal states [7][8][9]; the test here combines held-out head interventions with a matched multilingual contrast.
What else do those heads do?
The word sycophancy invites a tempting interpretation: perhaps we have isolated special heads for agreeing with a user. We therefore change the task. One control measures ordinary answer selection, another adds emotional pressure to the same false assertion, and a third changes who makes that assertion. Do their head maps overlap with the factual map?
A shared effect need not be a specific effect
View values as a table
| Checkpoint | Speaker r | Answer r | Social r | Residual signed r |
|---|---|---|---|---|
| Qwen3.5 0.8B | 0.892 | 0.303 | 0.681 | 0.853 |
| Qwen3.5 2B | 0.817 | 0.644 | 0.501 | 0.813 |
| Qwen3.5 4B | 0.957 | 0.917 | 0.876 | 0.612 |
| Qwen3.5 9B | 0.948 | 0.869 | 0.805 | 0.805 |
| Qwen3.5 27B | 0.825 | 0.622 | 0.601 | 0.811 |
| Qwen2.5 3B | 0.923 | 0.697 | 0.640 | 0.766 |
| Qwen2.5 7B | 0.852 | 0.781 | 0.638 | 0.849 |
| Llama-3.1 8B† | 0.935 | 0.780 | 0.764 | 0.756 |
| Mistral 7B | 0.878 | 0.666 | 0.663 | 0.849 |
| Gemma-2 9B | 0.934 | 0.780 | 0.749 | 0.882 |
| Gemma-4 12B | 0.769 | 0.629 | 0.602 | 0.835 |
Across checkpoints, factual-versus-answer-selection overlap averages about 0.70; factual-versus-social-pressure overlap about 0.68; and factual-versus-speaker-change overlap about 0.88. Fitting out an answer-selection map on one fold still leaves opposite-fold cross-language agreement of about 0.80. That last number does not prove a separate module: the controls overlap substantially, and the speaker-change map shares treatment with the factual map. The narrower conclusion is that the response to an asserted answer reuses components also involved in more general answer selection.
A direction travels, too
What if, instead of isolating heads, we extract one residual-stream direction from many examples? We fit a simple difference-of-means vector in each source language and ask whether it discriminates held-out false versus correct user assertions in another language. We also fit a separate social-pressure-versus-filler direction. These are established representation and transfer tools [10][11][12].
Does a direction fitted in one language classify another?
View values as a table
| Checkpoint | Factual AUROC | Social AUROC | Language control |
|---|---|---|---|
| Qwen3.5 0.8B | 0.669 | 0.588 | 0.493 |
| Qwen3.5 2B | 0.654 | 0.610 | 0.500 |
| Qwen3.5 4B | 0.800 | 0.920 | 0.501 |
| Qwen3.5 9B | 0.742 | 0.899 | 0.485 |
| Qwen3.5 27B | 0.931 | 0.948 | 0.499 |
| Qwen2.5 3B | 0.723 | 0.764 | 0.465 |
| Qwen2.5 7B | 0.767 | 0.761 | 0.457 |
| Llama-3.1 8B† | 0.767 | 0.899 | 0.504 |
| Mistral 7B | 0.893 | 0.922 | 0.525 |
| Gemma-2 9B | 0.833 | 0.929 | 0.425 |
| Gemma-4 12B | 0.767 | 0.954 | 0.477 |
Factual off-diagonal AUROC runs from 0.65 to 0.93 across checkpoints. But transfer alone is not special to sycophancy: separate sentiment, formality, and certainty controls also transfer. A saved singular-value summary describes the alignment of fitted language vectors, not the number of mechanisms. The separate EN+DE direction-steering experiment uses false-versus-neutral prompts and measures answer choice and neutral accuracy; it does not intervene on the transferred false-versus-correct vector. We have not run the matched joint intervention needed to establish that the residual direction is mediated by the patched heads.
Is the transfer result tied to the simplest direction extractor? On the same held-out factual transfer comparison, a logistic probe exceeds DiffMean in 11 of eleven checkpoints. LDA and PCA provide two further checks; PCA is often much closer to chance. This compares predictive extractors, not intervention strength, and a probe trained to separate these labels need not localize a cause.
Four ways to read the same contrast
View values as a table
| Checkpoint | DiffMean | Logistic | LDA | PCA |
|---|---|---|---|---|
| Qwen3.5 0.8B | 0.704 | 0.762 | 0.657 | 0.518 |
| Qwen3.5 2B | 0.662 | 0.763 | 0.755 | 0.506 |
| Qwen3.5 4B | 0.793 | 0.882 | 0.840 | 0.545 |
| Qwen3.5 9B | 0.743 | 0.879 | 0.853 | 0.544 |
| Qwen3.5 27B | 0.931 | 0.978 | 0.955 | 0.838 |
| Qwen2.5 3B | 0.725 | 0.832 | 0.817 | 0.498 |
| Qwen2.5 7B | 0.762 | 0.914 | 0.907 | 0.516 |
| Llama-3.1 8B† | 0.773 | 0.890 | 0.851 | 0.653 |
| Mistral 7B | 0.865 | 0.899 | 0.747 | 0.534 |
| Gemma-2 9B | 0.835 | 0.960 | 0.949 | 0.554 |
| Gemma-4 12B | 0.766 | 0.928 | 0.915 | 0.580 |
We also vary the surface wording of the English user claim five ways. Mean off-diagonal transfer between different phrasings ranges from 0.66 to 0.94 across checkpoints, rather than one uniform robustness score. This is an English-only control on disjoint questions; it cannot certify the thirteen machine-localized prompt sets or all paraphrases of a user claim.
What happens when the wording changes?
View values as a table
| Checkpoint | Mean off-diagonal | Weakest pair | N |
|---|---|---|---|
| Qwen3.5 0.8B | 0.657 | 0.560 | 214 |
| Qwen3.5 2B | 0.683 | 0.565 | 360 |
| Qwen3.5 4B | 0.776 | 0.645 | 488 |
| Qwen3.5 9B | 0.759 | 0.702 | 501 |
| Qwen3.5 27B | 0.944 | 0.820 | 509 |
| Qwen2.5 3B | 0.751 | 0.694 | 246 |
| Qwen2.5 7B | 0.812 | 0.787 | 396 |
| Llama-3.1 8B† | 0.826 | 0.739 | 69 |
| Mistral 7B | 0.890 | 0.820 | 187 |
| Gemma-2 9B | 0.862 | 0.817 | 514 |
| Gemma-4 12B | 0.770 | 0.711 | 490 |
Outside the multiple-choice prompt
The mechanism experiments force one of four letters. Real answers are often longer. In another experiment, a dialogue presents a correct assistant answer, the user challenges it with a wrong suggestion, and the model then generates a free-form reply. A mechanical final-letter readout and two model judges then assess whether it endorsed the wrong answer. Both judges are applied to every subject, with a leave-one-developer-out sensitivity check. Neither judge has been calibrated on human labels for this task.
How much does the readout change the answer?
View values as a table
| Language | N | Judged caving | Letter caving | Judge disagreement |
|---|---|---|---|---|
| EN | 500 | 0.224 | 0.486 | 0.290 |
| DE | 500 | 0.124 | 0.296 | 0.172 |
| FR | 500 | 0.284 | 0.308 | 0.112 |
| ES | 500 | 0.308 | 0.356 | 0.118 |
| IT | 500 | 0.270 | 0.296 | 0.088 |
| PT | 500 | 0.418 | 0.538 | 0.078 |
| ID | 500 | 0.330 | 0.346 | 0.126 |
| ZH | 500 | 0.618 | 0.654 | 0.100 |
| JA | 500 | 0.434 | 0.414 | 0.064 |
| KO | 500 | 0.292 | 0.328 | 0.118 |
| AR | 500 | 0.658 | 0.704 | 0.088 |
| HI | 500 | 0.308 | 0.290 | 0.148 |
| BN | 500 | 0.486 | 0.602 | 0.096 |
| SW | 500 | 0.780 | 0.736 | 0.060 |
| YO | 500 | 0.838 | 0.812 | 0.020 |
How often do these readouts agree? The two model judges disagree on about 10% of examples on average across checkpoint–language cells; the ensemble and final-letter labels agree on about 83%. Neither statistic gives human accuracy. Both can vary with the model being judged and the language, so the interactive view retains the per-language and leave-one-developer-out scores.
Agreement between automated readouts
View values as a table
| Checkpoint | Judge disagreement | Judge/letter agreement | All/cross-family agreement |
|---|---|---|---|
| Qwen3.5 0.8B | 0.129 | 0.741 | 0.973 |
| Qwen3.5 2B | 0.107 | 0.735 | 0.978 |
| Qwen3.5 4B | 0.057 | 0.836 | 0.993 |
| Qwen3.5 9B | 0.091 | 0.805 | 0.984 |
| Qwen3.5 27B | 0.113 | 0.869 | 0.981 |
| Qwen2.5 3B | 0.140 | 0.758 | 0.970 |
| Qwen2.5 7B | 0.112 | 0.833 | 0.976 |
| Llama-3.1 8B† | 0.076 | 0.885 | 1.000 |
| Mistral 7B | 0.085 | 0.852 | 1.000 |
| Gemma-2 9B | 0.062 | 0.954 | 0.947 |
| Gemma-4 12B | 0.138 | 0.897 | 0.884 |
We also ask what happens if pressure arrives turn by turn: mild doubt, direct contradiction, asserted expertise, a purported citation, consensus, then a request to ignore the preceding conversation. The reset step is a probe, not proof of a lasting repair. The interactive figure lets readers compare checkpoints and languages; its static view is one measured example.
Pressure builds; can a reset recover the answer?
View values as a table
| Turn | P(correct) | Wrong-option rate |
|---|---|---|
| T0 | 0.989 | 0.000 |
| T1 | 0.992 | 0.003 |
| T2 | 0.193 | 0.775 |
| T3 | 0.079 | 0.919 |
| T4 | 0.010 | 0.992 |
| T5 | 0.022 | 0.980 |
| T6reset | 0.867 | 0.109 |
The cost of replacement
The exact patches change a margin; does a less expensive intervention change the selected answer? We take heads selected in English, replace their outputs with the neutral mean learned on training questions, and test on held-out known questions in all fifteen languages. We compare a count/layer/activation-matched random set. The measured wrong-option rate falls in 151 of 165 checkpoint–language cells. Selecting in Chinese or Arabic and testing all target languages gives a restricted three-checkpoint symmetry check (reductions in 84 of 90 source–target cells).
There is a catch. The same replacement can turn a previously correct neutral answer into a wrong one. A treatment that makes a model less likely to choose the suggested wrong option is not automatically a good intervention if it also damages ordinary answering.
Less caving, at what accuracy cost?
View values as a table
| Checkpoint | Less caving (pp) | Neutral accuracy lost (pp) |
|---|---|---|
| Qwen3.5 0.8B | 2.80 | 1.06 |
| Qwen3.5 2B | 1.13 | 0.49 |
| Qwen3.5 4B | 14.27 | 3.16 |
| Qwen3.5 9B | 1.76 | 1.64 |
| Qwen3.5 27B | 3.73 | 0.58 |
| Qwen2.5 3B | 6.91 | 58.89 |
| Qwen2.5 7B | 4.89 | 9.82 |
| Llama-3.1 8B† | 51.21 | 32.27 |
| Mistral 7B | 36.86 | 14.22 |
| Gemma-2 9B | 15.93 | 18.64 |
| Gemma-4 12B | 1.58 | 3.24 |
Does the result depend on selecting the heads in English? For three checkpoints only, we also select in Chinese or Arabic and test all fifteen target languages. Wrong-option selection falls in 84 of ninety source–target cells. This is a useful symmetry check, not an eleven-model replication. The source-specific heads and neutral means differ, so the same replacement strength need not mean the same-sized change.
Change the source language for head selection
View values as a table
| Checkpoint | Head source | Targets improved | Less caving (pp) | Accuracy lost (pp) |
|---|---|---|---|---|
| Qwen2.5 7B | ZH | 14 | 5.51 | 7.89 |
| Qwen2.5 7B | AR | 14 | 5.69 | 10.69 |
| Llama-3.1 8B† | ZH | 13 | 51.30 | 32.17 |
| Llama-3.1 8B† | AR | 13 | 50.82 | 32.75 |
| Gemma-2 9B | ZH | 15 | 18.87 | 31.09 |
| Gemma-2 9B | AR | 15 | 18.91 | 30.78 |
Varying the number of replaced heads and the replacement strength gives a set of observed trade-offs. We show them all rather than presenting one chosen setting as a solution. In the saved analysis, settings under 2-, 5-, or 10-point mean neutral-accuracy-loss budgets were chosen using these same test items. That choice is exploratory, and a small mean can conceal a worse loss in one language.
An observed frontier, not a deployment setting
View values as a table
| Checkpoint | Sweep | Setting | Less caving (pp) | Accuracy lost (pp) |
|---|---|---|---|---|
| Qwen3.5 0.8B | Head count | 0 | 0.00 | 0.00 |
| Qwen3.5 0.8B | Head count | 2 | 1.74 | 0.37 |
| Qwen3.5 0.8B | Head count | 5 | 2.43 | 1.31 |
| Qwen3.5 0.8B | Head count | 10 | 3.12 | 1.18 |
| Qwen3.5 0.8B | Head count | 20 | 2.80 | 1.06 |
| Qwen3.5 0.8B | Replacement strength | 0.25 | 1.03 | -0.06 |
| Qwen3.5 0.8B | Replacement strength | 0.5 | 1.53 | 0.40 |
| Qwen3.5 0.8B | Replacement strength | 0.75 | 2.24 | 0.75 |
| Qwen3.5 2B | Head count | 0 | 0.00 | 0.00 |
| Qwen3.5 2B | Head count | 2 | 1.47 | 0.38 |
| Qwen3.5 2B | Head count | 5 | 1.49 | 0.33 |
| Qwen3.5 2B | Head count | 10 | 1.82 | 0.40 |
| Qwen3.5 2B | Head count | 19 | 1.13 | 0.49 |
| Qwen3.5 2B | Replacement strength | 0.25 | 0.62 | 0.13 |
| Qwen3.5 2B | Replacement strength | 0.5 | 1.11 | 0.22 |
| Qwen3.5 2B | Replacement strength | 0.75 | 1.27 | 0.31 |
| Qwen3.5 4B | Head count | 0 | 0.00 | 0.00 |
| Qwen3.5 4B | Head count | 2 | 5.67 | 2.36 |
| Qwen3.5 4B | Head count | 5 | 10.80 | 3.00 |
| Qwen3.5 4B | Head count | 10 | 12.11 | 3.56 |
| Qwen3.5 4B | Head count | 20 | 14.27 | 3.16 |
| Qwen3.5 4B | Replacement strength | 0.25 | 2.64 | 0.33 |
| Qwen3.5 4B | Replacement strength | 0.5 | 5.84 | 0.87 |
| Qwen3.5 4B | Replacement strength | 0.75 | 9.36 | 1.56 |
| Qwen3.5 9B | Head count | 0 | 0.00 | 0.00 |
| Qwen3.5 9B | Head count | 2 | 0.56 | 0.80 |
| Qwen3.5 9B | Head count | 5 | 1.00 | 0.82 |
| Qwen3.5 9B | Head count | 10 | 1.04 | 1.20 |
| Qwen3.5 9B | Head count | 20 | 1.76 | 1.64 |
| Qwen3.5 9B | Replacement strength | 0.25 | 0.29 | -0.02 |
| Qwen3.5 9B | Replacement strength | 0.5 | 0.84 | 0.36 |
| Qwen3.5 9B | Replacement strength | 0.75 | 1.24 | 0.73 |
| Qwen3.5 27B | Head count | 0 | 0.00 | 0.00 |
| Qwen3.5 27B | Head count | 2 | 1.16 | 0.04 |
| Qwen3.5 27B | Head count | 5 | 1.82 | 0.24 |
| Qwen3.5 27B | Head count | 10 | 2.24 | 0.31 |
| Qwen3.5 27B | Head count | 20 | 3.73 | 0.58 |
| Qwen3.5 27B | Replacement strength | 0.25 | 0.91 | 0.00 |
| Qwen3.5 27B | Replacement strength | 0.5 | 1.64 | 0.07 |
| Qwen3.5 27B | Replacement strength | 0.75 | 2.56 | 0.27 |
| Qwen2.5 3B | Head count | 0 | 0.00 | 0.00 |
| Qwen2.5 3B | Head count | 2 | -8.32 | 11.84 |
| Qwen2.5 3B | Head count | 5 | -7.59 | 26.83 |
| Qwen2.5 3B | Head count | 10 | -5.34 | 34.99 |
| Qwen2.5 3B | Head count | 20 | 6.91 | 58.89 |
| Qwen2.5 3B | Replacement strength | 0.25 | 1.00 | 1.44 |
| Qwen2.5 3B | Replacement strength | 0.5 | 1.25 | 5.09 |
| Qwen2.5 3B | Replacement strength | 0.75 | 3.41 | 21.49 |
| Qwen2.5 7B | Head count | 0 | 0.00 | 0.00 |
| Qwen2.5 7B | Head count | 2 | 0.76 | 0.20 |
| Qwen2.5 7B | Head count | 5 | 2.64 | 1.16 |
| Qwen2.5 7B | Head count | 10 | 4.04 | 2.33 |
| Qwen2.5 7B | Head count | 20 | 4.89 | 9.82 |
| Qwen2.5 7B | Replacement strength | 0.25 | 0.51 | 0.40 |
| Qwen2.5 7B | Replacement strength | 0.5 | 1.53 | 1.13 |
| Qwen2.5 7B | Replacement strength | 0.75 | 2.93 | 2.87 |
| Llama-3.1 8B† | Head count | 0 | 0.00 | 0.00 |
| Llama-3.1 8B† | Head count | 2 | 27.34 | 6.09 |
| Llama-3.1 8B† | Head count | 5 | 47.15 | 24.44 |
| Llama-3.1 8B† | Head count | 10 | 50.72 | 29.76 |
| Llama-3.1 8B† | Head count | 20 | 51.21 | 32.27 |
| Llama-3.1 8B† | Replacement strength | 0.25 | 16.62 | 1.84 |
| Llama-3.1 8B† | Replacement strength | 0.5 | 42.51 | 8.41 |
| Llama-3.1 8B† | Replacement strength | 0.75 | 51.88 | 20.87 |
| Mistral 7B | Head count | 0 | 0.00 | 0.00 |
| Mistral 7B | Head count | 2 | 26.35 | 4.88 |
| Mistral 7B | Head count | 5 | 36.97 | 14.22 |
| Mistral 7B | Head count | 10 | 36.86 | 14.22 |
| Mistral 7B | Head count | 20 | 36.86 | 14.22 |
| Mistral 7B | Replacement strength | 0.25 | 16.93 | 2.89 |
| Mistral 7B | Replacement strength | 0.5 | 35.01 | 13.37 |
| Mistral 7B | Replacement strength | 0.75 | 36.86 | 14.22 |
| Gemma-2 9B | Head count | 0 | 0.00 | 0.00 |
| Gemma-2 9B | Head count | 2 | 1.58 | 0.18 |
| Gemma-2 9B | Head count | 5 | 5.44 | 1.71 |
| Gemma-2 9B | Head count | 10 | 10.16 | 2.22 |
| Gemma-2 9B | Head count | 20 | 15.93 | 18.64 |
| Gemma-2 9B | Replacement strength | 0.25 | 1.89 | 0.24 |
| Gemma-2 9B | Replacement strength | 0.5 | 4.13 | 0.93 |
| Gemma-2 9B | Replacement strength | 0.75 | 7.47 | 3.09 |
| Gemma-4 12B | Head count | 0 | 0.00 | 0.00 |
| Gemma-4 12B | Head count | 2 | -0.02 | 0.32 |
| Gemma-4 12B | Head count | 5 | 0.11 | 0.07 |
| Gemma-4 12B | Head count | 10 | 1.13 | 1.42 |
| Gemma-4 12B | Head count | 20 | 1.58 | 3.24 |
| Gemma-4 12B | Replacement strength | 0.25 | 0.27 | 0.61 |
| Gemma-4 12B | Replacement strength | 0.5 | 0.81 | 0.99 |
| Gemma-4 12B | Replacement strength | 0.75 | 1.17 | 1.42 |
A different intervention tries to add the English false-minus-neutral head-activation shift to neutral prompts. At α=4, Mistral-7B shows 33.7 percentage points more wrong-option choices, but loses 93.5 points of neutral accuracy. That is a broken answerer, not convincing evidence of a clean way to induce sycophancy. Other checkpoints also mix smaller changes in wrong-option choice with accuracy loss. The one count/layer/activation-matched random head set is shown as a comparison, not as a population-level null.
A dramatic change can mean the model is failing
View values as a table
| Checkpoint | Wrong-option increase at α=4 (pp) | Neutral accuracy lost (pp) | Matched-random wrong increase (pp) | Matched-random accuracy lost (pp) |
|---|---|---|---|---|
| Qwen3.5 0.8B | 1.62 | 3.12 | 0.22 | 0.65 |
| Qwen3.5 2B | 0.16 | 0.82 | 0.36 | 1.07 |
| Qwen3.5 4B | 0.53 | 2.09 | 0.13 | 0.62 |
| Qwen3.5 9B | 0.11 | 0.71 | 0.33 | 0.87 |
| Qwen3.5 27B | 0.18 | 0.47 | 0.09 | 0.20 |
| Qwen2.5 3B | 4.04 | 12.03 | 0.35 | 0.92 |
| Qwen2.5 7B | 0.67 | 2.07 | 0.20 | 0.49 |
| Llama-3.1 8B† | 25.31 | 65.99 | 1.93 | 6.76 |
| Mistral 7B | 33.65 | 93.48 | 1.89 | 7.66 |
| Gemma-2 9B | 1.36 | 2.27 | 0.11 | 0.49 |
| Gemma-4 12B | 3.69 | 10.34 | 0.56 | 1.22 |
The separate residual-stream steering sweep fits a false-versus-neutral direction from English and German, then adds it at one layer on false-assertion prompts in each target language. At the modest α=+2 setting, changes in wrong-option choice are small and mixed in sign across checkpoints. It is not the false-versus-correct direction in the transfer plot, and there is no saved matched-random control for this sweep. Do not read it as proof of cross-language head–direction mediation.
A gentler direction is not a strong behavioral lever
View values as a table
| Checkpoint | Wrong-option change at α=+2 (pp) | Neutral accuracy lost (pp) |
|---|---|---|
| Qwen3.5 0.8B | -0.22 | -0.03 |
| Qwen3.5 2B | -0.18 | 0.07 |
| Qwen3.5 4B | 0.04 | 0.04 |
| Qwen3.5 9B | -0.02 | -0.04 |
| Qwen3.5 27B | -0.04 | 0.02 |
| Qwen2.5 3B | 0.73 | 0.16 |
| Qwen2.5 7B | 0.29 | 0.02 |
| Llama-3.1 8B† | 1.45 | 0.77 |
| Mistral 7B | 2.46 | 1.68 |
| Gemma-2 9B | 0.38 | 0.11 |
| Gemma-4 12B | -0.07 | -0.11 |
What the result does—and does not—show
We find a repeatable pattern of head-level effects across languages within checkpoints, and held-out patches show that selected head outputs matter for the forced-choice answer margin in the three patched languages. We do not show that head identities correspond across checkpoint architectures, that every language has an exact-patched circuit, or that this set of heads is unique to user deference. General answer selection overlaps substantially with the factual map.
The all-language neutral-correct screen narrows the sample. Llama-3.1-8B's 141 eligible questions make it low-N evidence in the four-developer patch panel. Qwen3.5's hybrid layers leave linear-attention pathways outside our head map. Custom prompts outside EN/ZH lack native-speaker validation. The free-form judges have no human-labeled calibration. Neutral-answer losses limit what we can call mitigation. These are limits of the experiment, not small-print exceptions to a universal mechanism claim.
Earlier work has separated sycophancy subtypes with directions [10], localized head-level signals [7], found shared sycophancy/lying components [9], and studied multilingual sycophancy as behavior [3]. This study combines established tools in a different test: matched multilingual factual prompts, disjoint-question map agreement, exact held-out head patches, and a measured accuracy cost. It is not a reproduction of the earlier papers' protocols or numbers.
Other nearby work asks different questions. ELEPHANT studies preservation of a user's social face [13]; our scripted emotional-pressure control is much narrower. Aldahlawi et al. compare opinion agreement across languages [14], while SYCON-Bench studies multi-turn dialogue [15]; neither is our aligned head-patching protocol. Head-targeted truthfulness interventions [16] and multilingual semantic-hub studies [17] motivate useful controls, but cannot substitute for changing the measured heads. Baez et al.'s factual/opinion representation study [18] likewise addresses a different subtype axis.
Data, methods, and references
Every plotted value is generated by the article data builder from locally available `paper-final` experiment outputs. Each chart JSON lists the contributing raw files and their SHA-256 hashes; the SVG fallback and interactive chart use the same extracted values. The figure-input inventory, data provenance note, and saved run configuration identify what is present. The available final-run raw files pass run/schema/status/config/code-provenance checks and match the transfer receipt hashes. The Space hosts locally available raw result files separately from older legacy results; no legacy or smoke result supplies a current figure. It does not claim the full 1,622-file final run was revalidated locally. Large activation outputs remain absent locally. No model weights or copied benchmark translations are included.
The complete experiment includes more measurements than can fit on one page. The representation data also contains 15-by-15 social/factual matrices, three generic concept controls, a four-extractor comparison, and five English prompt phrasings. The intervention data includes matched random-head nulls, source-language checks, and head-count/strength sweeps. The behavior data includes both judges, their cross-family sensitivity, all language curves, and readout disagreement. The old seven-/ten-type taxonomy is preserved as historical data, not used to claim a causal mechanism count here.
References
- Sharma et al. (2024), Towards Understanding Sycophancy in Language Models.
- Nikeghbal et al. (2026), Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs.
- Shah et al. (2026), Sycophancy as a Multilingual Alignment Failure.
- Hendrycks et al. (2021), Measuring Massive Multitask Language Understanding.
- OpenAI (2024), MMMLU: Multilingual Massive Multitask Language Understanding.
- Syed, Rager & Conmy (2023), Attribution Patching Outperforms Automated Circuit Discovery.
- Genadi et al. (2026), Sycophancy Hides Linearly in the Attention Heads.
- Wang et al. (2026), When Truth Is Overridden.
- Pandey (2026), LLMs Know They're Wrong and Agree Anyway.
- Vennemeyer et al. (2026), Sycophancy Is Not One Thing.
- Rimsky et al. (2024), Steering Llama 2 via Contrastive Activation Addition.
- Liu et al. (2026), Cross-Lingual Steering for Figurative Language Generation.
- Cheng et al. (2026), ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs.
- Aldahlawi et al. (2026), Investigating the Influence of Language on Sycophantic Behavior.
- Hong et al. (2025), Measuring Sycophancy of Language Models in Multi-turn Dialogues.
- Li et al. (2023), Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.
- Wu et al. (2025), The Semantic Hub Hypothesis.
- Baez et al. (2026), Dissociating the Internal Representations of Sycophancy.
The practical lesson is not that one switch causes sycophancy. It is that shared behavior invites a sharper test: identify the components, change them on new questions, and measure what else breaks.