# Data and figure provenance

## What the public article uses

The article's chart builder reads **all 1,457 locally available, saved `paper-final` raw result files**. Each current-run result family feeds at least one figure or diagnostic dataset, but not every item is shown as a separate mark. Every one of these files passed schema/run/complete-status, configuration-hash and accepted-code-hash checks, and matched the HCC transfer receipt's SHA-256. Each chart JSON records its own raw inputs and SHA-256 values; `assets/data/source_manifest.json` lists the 1,457-file union. Figure SVGs are generated from those values by `scripts/build_article_data.py`, and the interactive version reads the same chart JSON. The saved `derived/summary.json` checks selected recomputations, never supplies a plotted value. Passing file and run-provenance checks does not replace scientific validation of each contrast or a complete replay from model activations.

The archive was transferred from HCC with per-file checksums. The historical validation counted **1,622 required raw files** for the eleven-checkpoint, fifteen-language run, including 165 large per-language activation NPZ files (`acts2_*`) that are not present on this machine. The Space contains the **1,457 locally available final-run raw files**, not an independent complete raw replay. `data/provenance/file_inventory.json` contains each published file's byte count and SHA-256, plus the names of the 165 missing files. The original complete copies of these activation files remain on HCC. Another 195 legacy activation files are absent locally; no claim about a complete legacy raw archive is made.

The 857 locally available `data/raw/legacy/*.json` files are historical outputs from an earlier paper design. The old seven-/ten-type taxonomy's heterogeneous prompts do not justify a causal mechanism count. The article does not fold legacy numbers into the current eleven-checkpoint analysis. Logs, the Hugging Face dataset/model cache, compiled manuscript, and original benchmark rows are excluded.

## Dataset, models, and measurement limits

- English questions come from the pinned `cais/mmlu` test split; the fourteen translated versions come from pinned `openai/MMMLU` configs. See `data/provenance/revisions.json` for the model/dataset commits and `data/provenance/manifest.json` for the saved run configuration. The original GPU experiment source remains in a separate local research checkout, while this Space contains the complete article-building code and all locally available figure inputs.
- Models are instruction-tuned, with thinking modes disabled for matched comparisons. Qwen3.5 head maps cover its softmax-attention layers, not its linear-attention pathways.
- The headline pool consists of questions answered correctly under a neutral prompt in **all fifteen languages** for that checkpoint. The four response options and the suggested wrong letter are matched across conditions. This is a selected subset, not representative multilingual accuracy.
- Exact denoising/reverse patches are held out and cover **seven checkpoints, three languages**. Choice-level neutral-mean head replacement uses an English-selected set and all fifteen target languages; two alternate source languages are tested on three checkpoints only.
- Free-form caving uses two automatic model judges plus a letter readout. No human-labeled judge calibration was run. Figure annotations keep this separate from forced-choice head interventions.
- The original prompts outside English and Chinese were machine-translated and automatically checked, not validated by native-language annotators.
- The intervention trade-off settings are selected on the **same** held-out multiple-choice items used to display the trade-off. This is exploratory, not general utility or deployment validation.
- Head-induction doses and residual-stream steering use different sites, prompt contrasts, and scales. A large rise in wrong-option selection accompanied by a collapse in neutral accuracy is not evidence of a usable induction or mitigation. The residual-steering sweep has no saved matched-random baseline.
- Item-level first-order attribution faithfulness and head-ranking agreement with exact patches are separate statistics. A high map correlation does not show that each item's margin change is well approximated. Prompt-phrasing robustness is measured in English only.

## Rights and credits

The article's structure was inspired by [FineWeb's public research post](https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1), but does not reuse its text, images, stylesheet, JavaScript, or Distill bundle. Matplotlib produces SVG fallback figures; a locally bundled [Plotly.js](https://plotly.com/javascript/) enhances charts. Plotly.js has its own [MIT license and copyright notice](assets/vendor/PLOTLY_LICENSE.txt), separate from the unlicensed research materials. MMLU, MMMLU, model checkpoints, and cited research have their own upstream provenance and rights. This Space's metadata intentionally declares no blanket license for original research text or measurements. It does not bundle original MMLU/MMMLU question rows or model weights.

For literature citations, see the article's References section and the original manuscript bibliography in the research project. Earlier methods include difference-of-means directions, attribution patching, activation patching, and attention-head intervention; their use here is an adapted multilingual study, not a claim to have invented those techniques.
