Sycophancy · research noteEleven checkpoints · fifteen languages · one narrower question

An illustrated experiment report · October 2026

When a wrong answer
crosses languages

Do different-language versions of the same false user claim recruit the same parts of a language model? We trace the answer from matched questions to attention heads, then test the heads rather than trusting the map.

Heatmap of 18 English-selected attention-head effects for Qwen2.5 7B across 15 languages; the same head coordinates appear in each row.
One checkpoint, one set of head coordinates. The columns are 18 heads selected by their positive English training-fold attribution for this display; they are not necessarily the per-layer-capped 20-head patch set. Rows show the measured signed effect in each language. Color is an estimated contribution to a wrong-minus-correct margin, not a patched effect or a mapping between architectures. Data and source hashes.
StudyShared attention-head effects in multilingual factual sycophancy
ScopePublic research note · forced-choice mechanism study
Reading timeAbout 25 minutes

The question

Suppose a model answers a multiple-choice question correctly. Now ask it the same question while a user says, incorrectly, “I think the answer is B.” Sometimes the model chooses B. That is the narrow failure we study: following an asserted wrong option on a question the checkpoint answered correctly without the assertion. We do not equate one wrong answer with every form of sycophancy, or a correct neutral answer with proof of what the model internally “knows.”

This failure is already familiar from work on user agreement and answer instability [1][2]. Multilingual studies show that its frequency can vary by language [3]. But matching behavior leaves another question open: within one checkpoint, do the same attention heads affect the answer across languages, even when we change the questions used to estimate those effects?

TL;DR. Across 11 open-weight checkpoints and 15 aligned languages, head-effect maps recur on different questions (developer-weighted opposite-fold signed correlation about 0.91 in the well-powered group). In 21 checkpoint–language cells, replacing selected head outputs in both directions changes the answer margin. English-selected heads reduce forced-choice wrong-option selection in 151 of 165 model–language cells. But the heads overlap with ordinary answer selection, and changing them can harm neutral accuracy. The result is a causally consequential head-level contribution, not a unique sycophancy module or a free-form remedy.

First, choose the questions

We start from aligned English MMLU and fourteen translated MMMLU versions of the same multiple-choice questions [4][5]. We exclude rows where the answer key or source index does not align. For the main analysis, a checkpoint must answer a question correctly under a neutral prompt in every requested language. This removes a simple explanation for a later flip: perhaps the model could not answer the translated question to begin with.

It also changes the population we can study. A model with weaker multilingual neutral performance contributes fewer eligible questions. The figure shows the actual available sample after the all-language screen, capped at the 1,000-question target. Llama-3.1-8B has 141 eligible questions and is marked as low-N validation rather than placed in the high-power average.

How many questions survive?

Horizontal bars for 11 checkpoints showing the number of questions answered correctly in all 15 languages; the Llama-3.1 8B row has 141.
Eligibility is not the whole benchmark. The dashed 300-item line marks the high-power reporting threshold; smaller pools retain every eligible question. These are model-specific, aligned, neutral-correct questions, not representative samples of all questions. Raw-derived data.
View values as a table
CheckpointSelected questions
Qwen3.5 0.8B419
Qwen3.5 2B694
Qwen3.5 4B1000
Qwen3.5 9B1000
Qwen3.5 27B1000
Qwen2.5 3B494
Qwen2.5 7B784
Llama-3.1 8B†141
Mistral 7B366
Gemma-2 9B1000
Gemma-4 12B1000

The false option cycles reproducibly among the other three letters. Prompts ask for a single A/B/C/D answer, and the core contrast holds the question and suggested wrong option fixed. English and Chinese wording received author review; other prompt localizations were machine-translated and automatically checked, without native-speaker annotation. The Japanese and Korean belief prompts do not contain the same explicit first-person marker as English.

A recurring map

At the final assistant decision token, we measure the wrong-minus-correct answer-logit margin. For each attention head, the first-order map takes the change in that head's activation between false-assertion and neutral prompts and multiplies it by the gradient of the treatment margin. This is attribution patching, an established shortcut to locating potentially relevant components [6]. It is not the exact output change from replacing a head.

Now make the comparison harder. Split question IDs into two folds. Estimate one language's head map on one fold and another language's map on the opposite fold, then reverse the folds. If the maps agree, the agreement cannot be explained solely by reusing the very same question in both estimates. We compare signed effects: a head that favors the wrong option in one language should tend to do so in the other.

Do the same head effects recur on other questions?

Bar chart of opposite-fold signed head-map correlations across 15 languages for each of 11 checkpoints.
Strong on average, not uniform in every pair. Each bar is the mean over 105 language pairs within one checkpoint, using opposite question folds. The dagger/low-N status applies to Llama-3.1-8B. This does not compare head identity across architectures. Raw-derived pairs, matrices, and source hashes.
View values as a table
CheckpointMean signed rPairs below 0.6
Qwen3.5 0.8B0.8654
Qwen3.5 2B0.8393
Qwen3.5 4B0.9610
Qwen3.5 9B0.9261
Qwen3.5 27B0.8419
Qwen2.5 3B0.89113
Qwen2.5 7B0.9430
Llama-3.1 8B†0.8867
Mistral 7B0.9420
Gemma-2 9B0.9670
Gemma-4 12B0.8443

All eleven saved checkpoints pass the recorded model-level map-invariance gate. But a model-level mean can hide weak language pairs: 40 of 1,155 checkpoint–pair comparisons fall below a signed correlation of 0.6, and 33 of those involve Yoruba. The pairs share estimated maps, so these are not independent failures. The selectable matrix shows which comparisons pull an average down, without assigning the difference to language itself.

Where is agreement weaker?

Fifteen-by-fifteen opposite-fold signed head-map correlations in Qwen2.5 7B; the diagonal is unmeasured by design.
Different questions on opposite folds. The static matrix shows Qwen2.5-7B; select any checkpoint interactively. Each off-diagonal cell averages the two fold directions. The diagonal is deliberately blank, not zero. Forty of 1,155 pair cells are below the descriptive 0.6 threshold; weak YO pairs may also reflect prompt translation and selection. Raw-derived pair matrices.
View values as a table
CheckpointPairs below 0.6Of these, involving YO
Qwen3.5 0.8B40
Qwen3.5 2B32
Qwen3.5 4B00
Qwen3.5 9B11
Qwen3.5 27B99
Qwen2.5 3B1313
Qwen2.5 7B00
Llama-3.1 8B†77
Mistral 7B00
Gemma-2 9B00
Gemma-4 12B31

There is another test the average map cannot answer. If we sum the first-order head scores for one question, does that sum track its actual false-minus-neutral margin shift? Across languages, the mean per-item correlation ranges from -0.09 for Qwen3.5-0.8B to 0.81 for Qwen2.5-7B. Some hybrid Qwen3.5 models recur across languages despite weak or negative item-level approximation. A reproducible pattern of ranked heads is not an exact explanation for each answer.

A map can recur without predicting every item

Checkpoint means of the per-item correlation between summed first-order attribution and the observed answer-margin shift, averaged over fifteen languages.
First-order approximation, not a patch. Static bars show factual false-versus-neutral correlations averaged over fifteen languages; selecting a checkpoint shows its individual-language values. The contrast selector exposes three separately measured controls (speaker change, ordinary answer selection, and emotional pressure). Their target margins and prompts differ, so do not pool these four correlation distributions or call their differences separate mechanisms. A negative r means the summed estimate moves opposite the observed shift across items. Low-N Llama is marked †. All four valid item-level raw result families and hashes.
View values as a table
CheckpointMean per-item rLanguages with r < 0
Qwen3.5 0.8B-0.08911
Qwen3.5 2B0.1611
Qwen3.5 4B0.0534
Qwen3.5 9B0.1263
Qwen3.5 27B0.3020
Qwen2.5 3B0.7070
Qwen2.5 7B0.8120
Llama-3.1 8B†0.6620
Mistral 7B0.7370
Gemma-2 9B0.6510
Gemma-4 12B0.5910

Make a change, not a prediction

We take the positive heads selected on training questions and use held-out questions to ask two counterfactuals. Denoising: in the false-assertion run, replace those head outputs with the neutral run's outputs. Reverse patching: in the neutral run, insert the false-assertion outputs. The two changes should move the wrong-minus-correct margin in opposite directions if these heads matter for the measured response.

The exact patch is expensive, so it covers seven checkpoints in English, Chinese, and Arabic, not the entire 11-by-15 map. The figure plots all 21 tested cells. Each effect is divided by the unpatched false-minus-neutral shift on the same held-out patch questions; values can exceed one when a patch overshoots. Intervals resample the patch numerator while keeping that observed denominator fixed.

Replace the heads in both directions

Denoising and reverse-patching estimates with conditional item-bootstrap intervals for 21 checkpoint-language cells.
The held-out selected set has effects in both directions. Blue circles restore the margin in the false-assertion prompt; orange squares induce it in the neutral prompt. These are final-decision-token head-output patches in EN, ZH, and AR only; an untested cell is not a failed cell. Raw-derived patch effects.
View values as a table
CheckpointLanguageNDenoisingReverse
Qwen3.5 4BEN640.5970.723
Qwen3.5 4BZH640.4330.637
Qwen3.5 4BAR640.6140.654
Qwen2.5 3BEN640.9900.950
Qwen2.5 3BZH641.0921.019
Qwen2.5 3BAR641.0531.032
Qwen2.5 7BEN621.0321.020
Qwen2.5 7BZH621.0971.073
Qwen2.5 7BAR621.0051.030
Llama-3.1 8B†EN621.0701.117
Llama-3.1 8B†ZH620.9221.073
Llama-3.1 8B†AR621.0610.976
Mistral 7BEN640.9750.987
Mistral 7BZH640.7520.938
Mistral 7BAR640.7961.018
Gemma-2 9BEN641.1181.035
Gemma-2 9BZH641.0891.034
Gemma-2 9BAR640.9950.959
Gemma-4 12BEN611.6141.706
Gemma-4 12BZH611.3641.206
Gemma-4 12BAR611.5151.473

Did the cheap head ranking point to heads that the exact intervention also finds important? In 20 of 21 tested cells, the correlation between absolute first-order scores and absolute exact-denoising effects exceeds 0.9. The exception is Gemma-4-12B in English (r≈0.23). This compares maps of heads, not whether the sum of head scores predicts an individual question; the preceding figure shows that these tests can disagree.

Does attribution rank the heads that patching changes?

Absolute attribution-versus-exact-denoising head-map correlations for 21 checkpoint-language cells, with one low outlier in Gemma-4 English.
A ranking check, not per-item faithfulness. Attribution uses fold-0 false-versus-neutral head maps; exact denoising uses held-out patch items in EN/ZH/AR. Absolute correlations discard sign and do not measure whether one head or one pathway uniquely causes the behavior. The dotted 0.9 line is descriptive, not a significance test. Raw-derived 21-cell comparison.
View values as a table
CheckpointLanguageNAttribution/exact map r
Qwen3.5 4BEN640.992
Qwen3.5 4BZH640.988
Qwen3.5 4BAR640.991
Qwen2.5 3BEN640.982
Qwen2.5 3BZH640.975
Qwen2.5 3BAR640.979
Qwen2.5 7BEN620.978
Qwen2.5 7BZH620.955
Qwen2.5 7BAR620.986
Llama-3.1 8B†EN620.990
Llama-3.1 8B†ZH620.991
Llama-3.1 8B†AR620.993
Mistral 7BEN640.984
Mistral 7BZH640.993
Mistral 7BAR640.989
Gemma-2 9BEN640.981
Gemma-2 9BZH640.989
Gemma-2 9BAR640.993
Gemma-4 12BEN610.229
Gemma-4 12BZH610.953
Gemma-4 12BAR610.929

This is stronger evidence than a transferable probe, but it remains a node-level result. It does not identify every path through the network, the role of earlier token positions, or the linear-attention pathways omitted from the Qwen3.5 head map. Prior sycophancy studies have also localized heads or patched internal states [7][8][9]; the test here combines held-out head interventions with a matched multilingual contrast.

What else do those heads do?

The word sycophancy invites a tempting interpretation: perhaps we have isolated special heads for agreeing with a user. We therefore change the task. One control measures ordinary answer selection, another adds emotional pressure to the same false assertion, and a third changes who makes that assertion. Do their head maps overlap with the factual map?

A shared effect need not be a specific effect

Grouped horizontal bars compare three within-language map overlaps for eleven checkpoints.
Three controls overlap the factual map. Bars are within-language correlations of absolute head maps, averaged across 15 languages in each checkpoint. The speaker-change contrast uses the same user-role treatment as the factual map: it is user versus third-person, not third-person versus neutral. The cross-language residual statistic below is a different measurement. Raw-derived overlaps.
View values as a table
CheckpointSpeaker rAnswer rSocial rResidual signed r
Qwen3.5 0.8B0.8920.3030.6810.853
Qwen3.5 2B0.8170.6440.5010.813
Qwen3.5 4B0.9570.9170.8760.612
Qwen3.5 9B0.9480.8690.8050.805
Qwen3.5 27B0.8250.6220.6010.811
Qwen2.5 3B0.9230.6970.6400.766
Qwen2.5 7B0.8520.7810.6380.849
Llama-3.1 8B†0.9350.7800.7640.756
Mistral 7B0.8780.6660.6630.849
Gemma-2 9B0.9340.7800.7490.882
Gemma-4 12B0.7690.6290.6020.835

Across checkpoints, factual-versus-answer-selection overlap averages about 0.70; factual-versus-social-pressure overlap about 0.68; and factual-versus-speaker-change overlap about 0.88. Fitting out an answer-selection map on one fold still leaves opposite-fold cross-language agreement of about 0.80. That last number does not prove a separate module: the controls overlap substantially, and the speaker-change map shares treatment with the factual map. The narrower conclusion is that the response to an asserted answer reuses components also involved in more general answer selection.

A direction travels, too

What if, instead of isolating heads, we extract one residual-stream direction from many examples? We fit a simple difference-of-means vector in each source language and ask whether it discriminates held-out false versus correct user assertions in another language. We also fit a separate social-pressure-versus-filler direction. These are established representation and transfer tools [10][11][12].

Does a direction fitted in one language classify another?

Dot plot of held-out off-diagonal AUROC for factual and social contrasts in eleven checkpoints.
Prediction transfers, with two different prompt contrasts. AUROC is measured on held-out questions and target languages; it is not a causal steering outcome. The social contrast is emotional pressure versus matched filler on a false assertion. Neither contrast is identical to the false-versus-neutral head map. Raw transfer matrices and concept controls.
View values as a table
CheckpointFactual AUROCSocial AUROCLanguage control
Qwen3.5 0.8B0.6690.5880.493
Qwen3.5 2B0.6540.6100.500
Qwen3.5 4B0.8000.9200.501
Qwen3.5 9B0.7420.8990.485
Qwen3.5 27B0.9310.9480.499
Qwen2.5 3B0.7230.7640.465
Qwen2.5 7B0.7670.7610.457
Llama-3.1 8B†0.7670.8990.504
Mistral 7B0.8930.9220.525
Gemma-2 9B0.8330.9290.425
Gemma-4 12B0.7670.9540.477

Factual off-diagonal AUROC runs from 0.65 to 0.93 across checkpoints. But transfer alone is not special to sycophancy: separate sentiment, formality, and certainty controls also transfer. A saved singular-value summary describes the alignment of fitted language vectors, not the number of mechanisms. The separate EN+DE direction-steering experiment uses false-versus-neutral prompts and measures answer choice and neutral accuracy; it does not intervene on the transferred false-versus-correct vector. We have not run the matched joint intervention needed to establish that the residual direction is mediated by the patched heads.

Is the transfer result tied to the simplest direction extractor? On the same held-out factual transfer comparison, a logistic probe exceeds DiffMean in 11 of eleven checkpoints. LDA and PCA provide two further checks; PCA is often much closer to chance. This compares predictive extractors, not intervention strength, and a probe trained to separate these labels need not localize a cause.

Four ways to read the same contrast

Four small-multiple plots of held-out factual cross-language AUROC for DiffMean, logistic regression, LDA, and PCA across eleven checkpoints.
Predictive performance depends on the extractor. The static four panels share an AUROC scale and held-out question folds; the interactive view selects one extractor. No score here is a causal estimate. Raw method comparison.
View values as a table
CheckpointDiffMeanLogisticLDAPCA
Qwen3.5 0.8B0.7040.7620.6570.518
Qwen3.5 2B0.6620.7630.7550.506
Qwen3.5 4B0.7930.8820.8400.545
Qwen3.5 9B0.7430.8790.8530.544
Qwen3.5 27B0.9310.9780.9550.838
Qwen2.5 3B0.7250.8320.8170.498
Qwen2.5 7B0.7620.9140.9070.516
Llama-3.1 8B†0.7730.8900.8510.653
Mistral 7B0.8650.8990.7470.534
Gemma-2 9B0.8350.9600.9490.554
Gemma-4 12B0.7660.9280.9150.580

We also vary the surface wording of the English user claim five ways. Mean off-diagonal transfer between different phrasings ranges from 0.66 to 0.94 across checkpoints, rather than one uniform robustness score. This is an English-only control on disjoint questions; it cannot certify the thirteen machine-localized prompt sets or all paraphrases of a user claim.

What happens when the wording changes?

Per-checkpoint mean and weakest held-out English cross-phrasing AUROC over five false-answer phrasings.
Mean and weakest pair both matter. Each checkpoint has twenty cross-template directions fitted on one item fold and evaluated on another; the diagonal of each five-by-five matrix is excluded from the displayed summary. This is a prediction test, not proof that a causal head set is unchanged by paraphrasing. Raw five-phrasing matrices.
View values as a table
CheckpointMean off-diagonalWeakest pairN
Qwen3.5 0.8B0.6570.560214
Qwen3.5 2B0.6830.565360
Qwen3.5 4B0.7760.645488
Qwen3.5 9B0.7590.702501
Qwen3.5 27B0.9440.820509
Qwen2.5 3B0.7510.694246
Qwen2.5 7B0.8120.787396
Llama-3.1 8B†0.8260.73969
Mistral 7B0.8900.820187
Gemma-2 9B0.8620.817514
Gemma-4 12B0.7700.711490

Outside the multiple-choice prompt

The mechanism experiments force one of four letters. Real answers are often longer. In another experiment, a dialogue presents a correct assistant answer, the user challenges it with a wrong suggestion, and the model then generates a free-form reply. A mechanical final-letter readout and two model judges then assess whether it endorsed the wrong answer. Both judges are applied to every subject, with a leave-one-developer-out sensitivity check. Neither judge has been calibrated on human labels for this task.

How much does the readout change the answer?

For three English checkpoint examples, compare free-form caving rates from a two-model judge ensemble and from a final-letter readout.
These are diagnostic labels, not ground truth. The static view shows three English checkpoints; the interactive chart can show any saved checkpoint and language. Judge disagreement and cross-family estimates are available in the table and data. No head replacement is tested on these free-form replies. Raw-derived behavior data.
View values as a table
LanguageNJudged cavingLetter cavingJudge disagreement
EN5000.2240.4860.290
DE5000.1240.2960.172
FR5000.2840.3080.112
ES5000.3080.3560.118
IT5000.2700.2960.088
PT5000.4180.5380.078
ID5000.3300.3460.126
ZH5000.6180.6540.100
JA5000.4340.4140.064
KO5000.2920.3280.118
AR5000.6580.7040.088
HI5000.3080.2900.148
BN5000.4860.6020.096
SW5000.7800.7360.060
YO5000.8380.8120.020

How often do these readouts agree? The two model judges disagree on about 10% of examples on average across checkpoint–language cells; the ensemble and final-letter labels agree on about 83%. Neither statistic gives human accuracy. Both can vary with the model being judged and the language, so the interactive view retains the per-language and leave-one-developer-out scores.

Agreement between automated readouts

Two vertically separated charts show average model-judge disagreement and judge-versus-letter agreement for eleven checkpoints.
Agreement is not a truth label. The static panels have separately labeled fraction scales: top is two-judge disagreement, bottom is ensemble versus mechanical-letter agreement. The interactive view selects either measure and can inspect each language. Neither is calibrated against human annotations. Raw-derived judge panels.
View values as a table
CheckpointJudge disagreementJudge/letter agreementAll/cross-family agreement
Qwen3.5 0.8B0.1290.7410.973
Qwen3.5 2B0.1070.7350.978
Qwen3.5 4B0.0570.8360.993
Qwen3.5 9B0.0910.8050.984
Qwen3.5 27B0.1130.8690.981
Qwen2.5 3B0.1400.7580.970
Qwen2.5 7B0.1120.8330.976
Llama-3.1 8B†0.0760.8851.000
Mistral 7B0.0850.8521.000
Gemma-2 9B0.0620.9540.947
Gemma-4 12B0.1380.8970.884

We also ask what happens if pressure arrives turn by turn: mild doubt, direct contradiction, asserted expertise, a purported citation, consensus, then a request to ignore the preceding conversation. The reset step is a probe, not proof of a lasting repair. The interactive figure lets readers compare checkpoints and languages; its static view is one measured example.

Pressure builds; can a reset recover the answer?

Qwen2.5 7B English probability of the correct option over the initial turn, five challenges, and a reset request.
One conversational ladder, not a natural conversation sample. T0 is the initial question, T1–T5 are increasing scripted pressure, and T6 requests a context reset. The plotted probability comes from the four option logits; the full saved curves cover 11 checkpoints and 15 languages. Raw-derived curves.
View values as a table
TurnP(correct)Wrong-option rate
T00.9890.000
T10.9920.003
T20.1930.775
T30.0790.919
T40.0100.992
T50.0220.980
T6reset0.8670.109

The cost of replacement

The exact patches change a margin; does a less expensive intervention change the selected answer? We take heads selected in English, replace their outputs with the neutral mean learned on training questions, and test on held-out known questions in all fifteen languages. We compare a count/layer/activation-matched random set. The measured wrong-option rate falls in 151 of 165 checkpoint–language cells. Selecting in Chinese or Arabic and testing all target languages gives a restricted three-checkpoint symmetry check (reductions in 84 of 90 source–target cells).

There is a catch. The same replacement can turn a previously correct neutral answer into a wrong one. A treatment that makes a model less likely to choose the suggested wrong option is not automatically a good intervention if it also damages ordinary answering.

Less caving, at what accuracy cost?

Each checkpoint is a point comparing average reduction in wrong-option selection with loss of neutral accuracy over fifteen languages.
Both axes come from the same held-out, neutral-correct multiple-choice items. Each point averages languages equally within one checkpoint. Lower neutral accuracy is a cost, not a general utility measure. Matched random-head results and per-language values are in the linked data. Raw-derived intervention data.
View values as a table
CheckpointLess caving (pp)Neutral accuracy lost (pp)
Qwen3.5 0.8B2.801.06
Qwen3.5 2B1.130.49
Qwen3.5 4B14.273.16
Qwen3.5 9B1.761.64
Qwen3.5 27B3.730.58
Qwen2.5 3B6.9158.89
Qwen2.5 7B4.899.82
Llama-3.1 8B†51.2132.27
Mistral 7B36.8614.22
Gemma-2 9B15.9318.64
Gemma-4 12B1.583.24

Does the result depend on selecting the heads in English? For three checkpoints only, we also select in Chinese or Arabic and test all fifteen target languages. Wrong-option selection falls in 84 of ninety source–target cells. This is a useful symmetry check, not an eleven-model replication. The source-specific heads and neutral means differ, so the same replacement strength need not mean the same-sized change.

Change the source language for head selection

Six source-model intervention cells compare average wrong-answer reduction with neutral-accuracy loss over fifteen targets.
Restricted source-symmetry panel. ZH- and AR-selected sets were tested only on Qwen2.5-7B, Gemma-2-9B, and low-N Llama-3.1-8B. Axes are equal-language averages in percentage points on held-out, previously known questions. Equal α does not ensure equal activation changes. Raw source-specific measurements.
View values as a table
CheckpointHead sourceTargets improvedLess caving (pp)Accuracy lost (pp)
Qwen2.5 7BZH145.517.89
Qwen2.5 7BAR145.6910.69
Llama-3.1 8B†ZH1351.3032.17
Llama-3.1 8B†AR1350.8232.75
Gemma-2 9BZH1518.8731.09
Gemma-2 9BAR1518.9130.78

Varying the number of replaced heads and the replacement strength gives a set of observed trade-offs. We show them all rather than presenting one chosen setting as a solution. In the saved analysis, settings under 2-, 5-, or 10-point mean neutral-accuracy-loss budgets were chosen using these same test items. That choice is exploratory, and a small mean can conceal a worse loss in one language.

An observed frontier, not a deployment setting

Scatter of head-count and replacement-strength settings by neutral accuracy loss and reduction in wrong-option selection.
Each mark is a measured setting. Both axes are percentage-point changes averaged equally over all 15 target languages for a checkpoint. The dashed line is a five-point mean-loss reference, not a validated safety threshold. No free-form utility or deployment validation was run. Raw-derived settings and worst-language costs.
View values as a table
CheckpointSweepSettingLess caving (pp)Accuracy lost (pp)
Qwen3.5 0.8BHead count00.000.00
Qwen3.5 0.8BHead count21.740.37
Qwen3.5 0.8BHead count52.431.31
Qwen3.5 0.8BHead count103.121.18
Qwen3.5 0.8BHead count202.801.06
Qwen3.5 0.8BReplacement strength0.251.03-0.06
Qwen3.5 0.8BReplacement strength0.51.530.40
Qwen3.5 0.8BReplacement strength0.752.240.75
Qwen3.5 2BHead count00.000.00
Qwen3.5 2BHead count21.470.38
Qwen3.5 2BHead count51.490.33
Qwen3.5 2BHead count101.820.40
Qwen3.5 2BHead count191.130.49
Qwen3.5 2BReplacement strength0.250.620.13
Qwen3.5 2BReplacement strength0.51.110.22
Qwen3.5 2BReplacement strength0.751.270.31
Qwen3.5 4BHead count00.000.00
Qwen3.5 4BHead count25.672.36
Qwen3.5 4BHead count510.803.00
Qwen3.5 4BHead count1012.113.56
Qwen3.5 4BHead count2014.273.16
Qwen3.5 4BReplacement strength0.252.640.33
Qwen3.5 4BReplacement strength0.55.840.87
Qwen3.5 4BReplacement strength0.759.361.56
Qwen3.5 9BHead count00.000.00
Qwen3.5 9BHead count20.560.80
Qwen3.5 9BHead count51.000.82
Qwen3.5 9BHead count101.041.20
Qwen3.5 9BHead count201.761.64
Qwen3.5 9BReplacement strength0.250.29-0.02
Qwen3.5 9BReplacement strength0.50.840.36
Qwen3.5 9BReplacement strength0.751.240.73
Qwen3.5 27BHead count00.000.00
Qwen3.5 27BHead count21.160.04
Qwen3.5 27BHead count51.820.24
Qwen3.5 27BHead count102.240.31
Qwen3.5 27BHead count203.730.58
Qwen3.5 27BReplacement strength0.250.910.00
Qwen3.5 27BReplacement strength0.51.640.07
Qwen3.5 27BReplacement strength0.752.560.27
Qwen2.5 3BHead count00.000.00
Qwen2.5 3BHead count2-8.3211.84
Qwen2.5 3BHead count5-7.5926.83
Qwen2.5 3BHead count10-5.3434.99
Qwen2.5 3BHead count206.9158.89
Qwen2.5 3BReplacement strength0.251.001.44
Qwen2.5 3BReplacement strength0.51.255.09
Qwen2.5 3BReplacement strength0.753.4121.49
Qwen2.5 7BHead count00.000.00
Qwen2.5 7BHead count20.760.20
Qwen2.5 7BHead count52.641.16
Qwen2.5 7BHead count104.042.33
Qwen2.5 7BHead count204.899.82
Qwen2.5 7BReplacement strength0.250.510.40
Qwen2.5 7BReplacement strength0.51.531.13
Qwen2.5 7BReplacement strength0.752.932.87
Llama-3.1 8B†Head count00.000.00
Llama-3.1 8B†Head count227.346.09
Llama-3.1 8B†Head count547.1524.44
Llama-3.1 8B†Head count1050.7229.76
Llama-3.1 8B†Head count2051.2132.27
Llama-3.1 8B†Replacement strength0.2516.621.84
Llama-3.1 8B†Replacement strength0.542.518.41
Llama-3.1 8B†Replacement strength0.7551.8820.87
Mistral 7BHead count00.000.00
Mistral 7BHead count226.354.88
Mistral 7BHead count536.9714.22
Mistral 7BHead count1036.8614.22
Mistral 7BHead count2036.8614.22
Mistral 7BReplacement strength0.2516.932.89
Mistral 7BReplacement strength0.535.0113.37
Mistral 7BReplacement strength0.7536.8614.22
Gemma-2 9BHead count00.000.00
Gemma-2 9BHead count21.580.18
Gemma-2 9BHead count55.441.71
Gemma-2 9BHead count1010.162.22
Gemma-2 9BHead count2015.9318.64
Gemma-2 9BReplacement strength0.251.890.24
Gemma-2 9BReplacement strength0.54.130.93
Gemma-2 9BReplacement strength0.757.473.09
Gemma-4 12BHead count00.000.00
Gemma-4 12BHead count2-0.020.32
Gemma-4 12BHead count50.110.07
Gemma-4 12BHead count101.131.42
Gemma-4 12BHead count201.583.24
Gemma-4 12BReplacement strength0.250.270.61
Gemma-4 12BReplacement strength0.50.810.99
Gemma-4 12BReplacement strength0.751.171.42

A different intervention tries to add the English false-minus-neutral head-activation shift to neutral prompts. At α=4, Mistral-7B shows 33.7 percentage points more wrong-option choices, but loses 93.5 points of neutral accuracy. That is a broken answerer, not convincing evidence of a clean way to induce sycophancy. Other checkpoints also mix smaller changes in wrong-option choice with accuracy loss. The one count/layer/activation-matched random head set is shown as a comparison, not as a population-level null.

A dramatic change can mean the model is failing

Eleven model-level head-induction points compare wrong-option increase on neutral prompts with neutral-accuracy loss; extreme Mistral and Llama effects have large losses.
Head induction is a different experiment from exact reverse patching. English-fitted mean shifts are added at selected head outputs on held-out neutral questions. Static marks average fifteen target languages equally at α=4; the interactive view can show another dose or one checkpoint's target languages. Outlined squares are one matched random set where measured. Large positive changes in wrong-choice rate coincide with loss of previously correct neutral answers. Raw induction doses and controls.
View values as a table
CheckpointWrong-option increase at α=4 (pp)Neutral accuracy lost (pp)Matched-random wrong increase (pp)Matched-random accuracy lost (pp)
Qwen3.5 0.8B1.623.120.220.65
Qwen3.5 2B0.160.820.361.07
Qwen3.5 4B0.532.090.130.62
Qwen3.5 9B0.110.710.330.87
Qwen3.5 27B0.180.470.090.20
Qwen2.5 3B4.0412.030.350.92
Qwen2.5 7B0.672.070.200.49
Llama-3.1 8B†25.3165.991.936.76
Mistral 7B33.6593.481.897.66
Gemma-2 9B1.362.270.110.49
Gemma-4 12B3.6910.340.561.22

The separate residual-stream steering sweep fits a false-versus-neutral direction from English and German, then adds it at one layer on false-assertion prompts in each target language. At the modest α=+2 setting, changes in wrong-option choice are small and mixed in sign across checkpoints. It is not the false-versus-correct direction in the transfer plot, and there is no saved matched-random control for this sweep. Do not read it as proof of cross-language head–direction mediation.

A gentler direction is not a strong behavioral lever

Eleven model-level residual-steering points compare change in false-assertion wrong-answer selection at alpha plus two against neutral accuracy loss.
One layer, a separate contrast, a separate scale. Static marks show α=+2 equal-language means; the interactive controls can show other doses or one checkpoint's target languages. Positive/negative values indicate a change in choosing the user's false answer, not a general free-form effect. Head induction and residual steering use different sites and dose definitions and should not be pooled. Raw steering values.
View values as a table
CheckpointWrong-option change at α=+2 (pp)Neutral accuracy lost (pp)
Qwen3.5 0.8B-0.22-0.03
Qwen3.5 2B-0.180.07
Qwen3.5 4B0.040.04
Qwen3.5 9B-0.02-0.04
Qwen3.5 27B-0.040.02
Qwen2.5 3B0.730.16
Qwen2.5 7B0.290.02
Llama-3.1 8B†1.450.77
Mistral 7B2.461.68
Gemma-2 9B0.380.11
Gemma-4 12B-0.07-0.11

What the result does—and does not—show

We find a repeatable pattern of head-level effects across languages within checkpoints, and held-out patches show that selected head outputs matter for the forced-choice answer margin in the three patched languages. We do not show that head identities correspond across checkpoint architectures, that every language has an exact-patched circuit, or that this set of heads is unique to user deference. General answer selection overlaps substantially with the factual map.

The all-language neutral-correct screen narrows the sample. Llama-3.1-8B's 141 eligible questions make it low-N evidence in the four-developer patch panel. Qwen3.5's hybrid layers leave linear-attention pathways outside our head map. Custom prompts outside EN/ZH lack native-speaker validation. The free-form judges have no human-labeled calibration. Neutral-answer losses limit what we can call mitigation. These are limits of the experiment, not small-print exceptions to a universal mechanism claim.

Earlier work has separated sycophancy subtypes with directions [10], localized head-level signals [7], found shared sycophancy/lying components [9], and studied multilingual sycophancy as behavior [3]. This study combines established tools in a different test: matched multilingual factual prompts, disjoint-question map agreement, exact held-out head patches, and a measured accuracy cost. It is not a reproduction of the earlier papers' protocols or numbers.

Other nearby work asks different questions. ELEPHANT studies preservation of a user's social face [13]; our scripted emotional-pressure control is much narrower. Aldahlawi et al. compare opinion agreement across languages [14], while SYCON-Bench studies multi-turn dialogue [15]; neither is our aligned head-patching protocol. Head-targeted truthfulness interventions [16] and multilingual semantic-hub studies [17] motivate useful controls, but cannot substitute for changing the measured heads. Baez et al.'s factual/opinion representation study [18] likewise addresses a different subtype axis.

Data, methods, and references

Every plotted value is generated by the article data builder from locally available `paper-final` experiment outputs. Each chart JSON lists the contributing raw files and their SHA-256 hashes; the SVG fallback and interactive chart use the same extracted values. The figure-input inventory, data provenance note, and saved run configuration identify what is present. The available final-run raw files pass run/schema/status/config/code-provenance checks and match the transfer receipt hashes. The Space hosts locally available raw result files separately from older legacy results; no legacy or smoke result supplies a current figure. It does not claim the full 1,622-file final run was revalidated locally. Large activation outputs remain absent locally. No model weights or copied benchmark translations are included.

The complete experiment includes more measurements than can fit on one page. The representation data also contains 15-by-15 social/factual matrices, three generic concept controls, a four-extractor comparison, and five English prompt phrasings. The intervention data includes matched random-head nulls, source-language checks, and head-count/strength sweeps. The behavior data includes both judges, their cross-family sensitivity, all language curves, and readout disagreement. The old seven-/ten-type taxonomy is preserved as historical data, not used to claim a causal mechanism count here.

References

  1. Sharma et al. (2024), Towards Understanding Sycophancy in Language Models.
  2. Nikeghbal et al. (2026), Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs.
  3. Shah et al. (2026), Sycophancy as a Multilingual Alignment Failure.
  4. Hendrycks et al. (2021), Measuring Massive Multitask Language Understanding.
  5. OpenAI (2024), MMMLU: Multilingual Massive Multitask Language Understanding.
  6. Syed, Rager & Conmy (2023), Attribution Patching Outperforms Automated Circuit Discovery.
  7. Genadi et al. (2026), Sycophancy Hides Linearly in the Attention Heads.
  8. Wang et al. (2026), When Truth Is Overridden.
  9. Pandey (2026), LLMs Know They're Wrong and Agree Anyway.
  10. Vennemeyer et al. (2026), Sycophancy Is Not One Thing.
  11. Rimsky et al. (2024), Steering Llama 2 via Contrastive Activation Addition.
  12. Liu et al. (2026), Cross-Lingual Steering for Figurative Language Generation.
  13. Cheng et al. (2026), ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs.
  14. Aldahlawi et al. (2026), Investigating the Influence of Language on Sycophantic Behavior.
  15. Hong et al. (2025), Measuring Sycophancy of Language Models in Multi-turn Dialogues.
  16. Li et al. (2023), Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.
  17. Wu et al. (2025), The Semantic Hub Hypothesis.
  18. Baez et al. (2026), Dissociating the Internal Representations of Sycophancy.

The practical lesson is not that one switch causes sycophancy. It is that shared behavior invites a sharper test: identify the components, change them on new questions, and measure what else breaks.