Under peer review. The full dataset will be available upon publication.
  1. 1

    Unlearning in one language does not close the other-language routes.

    Methoda→a removed %ℓ→a removed %ℓ→ℓ removed %other languages ≥90% removed %worst language: access left %
    SimNPO97.291.081.830.049.2
    GradDiff99.796.4110.571.126.2
    RMU100.291.163.518.075.4

    Here a is the fine-tuning language and ℓ any other language (question→answer); removal = 100·(KF − KU)/(KF − KO) with KF, KO and KU the fine-tuned, retain-only-oracle and unlearned answer strengths, access left is 100 minus removal, and values above 100 mean suppression below the retain-only oracle.

Likelihood summaries over 50 forget facts per cell; eligible languages weighted equally within a fine-tuning language, languages within a family, families equally.

Every question language, grouped by writing system

one block per language · hover for its detail

The Language Grid

One square per (model, question language). Hover or focus a square to read it out below; click or press Enter to pin the full comparison into the table further down.

Hover a square to read out its question language
 
 
 

measured
hatched: below the learning threshold (fine-tuned answer strength < 0.10, or gain over the retain-only oracle < 0.05) — removal ratios are unreliable there; exact values remain in the tooltip
does not apply to this row (blank)

Pinned square

Click any square in the grid to pin every measurement for that fine-tuning language and question language here — all three answer languages, both fact sets, nothing hover-only.

Question form: paraphrases and translated paraphrases

The same 100 facts are asked four ways, with the English answer held fixed: the canonical English question, an English paraphrase, the translated canonical question, and the translated paraphrase. Numbers are teacher-forced gold-answer likelihood. Unlearned rows are SimNPO endpoints trained on three-language source subsets (three subsets × three training orders per model family).

Knowledge gained by fine-tuning, by question form

Acquisition gain = K(fine-tuned) − K(base) on the 0–1 answer-strength scale, English gold answer. For the English rows each language column restricts to the facts matched in that language, so English values vary slightly by column. The five-language mean weights the languages equally.


Download data/expression_form_language.parquet
ModelQuestion formArabicHausaHindiNepaliSwahiliFive-language mean
AyaEnglish canonical0.9260.9260.9260.9260.9260.926
AyaEnglish paraphrase0.8900.8910.8930.8910.8920.892
AyaTranslated canonical0.7990.7900.7310.7750.8180.783
AyaTranslated paraphrase0.7810.7730.7250.7530.7950.766
QwenEnglish canonical0.8940.8930.8930.8930.8930.893
QwenEnglish paraphrase0.8550.8580.8580.8570.8560.857
QwenTranslated canonical0.5600.5040.3520.4730.4820.474
QwenTranslated paraphrase0.5140.4880.3960.4510.4630.462
LlamaEnglish canonical0.8590.8580.8580.8580.8580.859
LlamaEnglish paraphrase0.8260.8260.8260.8250.8250.826
LlamaTranslated canonical0.4930.4760.3640.4240.4650.445
LlamaTranslated paraphrase0.4530.4700.4140.4150.4520.441

Access remaining after SimNPO, by question form (%)

Remaining access = 100 · (K(unlearned) − K(oracle)) / (K(fine-tuned) − K(oracle)), unclipped: 0 is the retain-only oracle, 100 the fine-tuned model. Averages the nine SimNPO control endpoints per family. A fact counts in a column when the fine-tuned K ≥ 0.10 and fine-tuned − oracle ≥ 0.05 in all four forms.


Download data/expression_form_language.parquet
ModelQuestion formArabicHausaHindiNepaliSwahiliFive-language mean
AyaEnglish canonical14.810.010.210.59.911.1
AyaEnglish paraphrase13.99.79.69.99.410.5
AyaTranslated canonical8.33.29.44.93.95.9
AyaTranslated paraphrase7.44.55.20.92.94.2
QwenEnglish canonical32.432.232.232.235.632.9
QwenEnglish paraphrase32.733.232.731.435.133.0
QwenTranslated canonical29.224.837.327.329.629.6
QwenTranslated paraphrase32.926.928.927.031.629.5
LlamaEnglish canonical20.820.427.823.020.622.5
LlamaEnglish paraphrase22.622.527.124.522.023.7
LlamaTranslated canonical28.827.831.322.924.927.1
LlamaTranslated paraphrase29.526.427.923.128.227.0

English paraphrase types ranked

A win is higher gold-answer likelihood on a shared question; ties count ½. Win % averages matched facts within author, authors within each pair of types, the eight other types within a family, then the three families equally. Before = the fine-tuned models; after = the mean of the nine SimNPO control endpoints. QA = facts with a paraphrase of that type. Format is unranked (six facts).


Download data/expression_type_ranking.csv
RankParaphrase typeQABefore win %After win %Rank after, by model
AyaQwenLlama
1Modal verb2877.767.6111
2Subordination6354.558.8332
3Sentence modality10053.758.1223
4Same-polarity habitual9152.147.4456
5Same-polarity contextual9253.045.7874
6Syntax/discourse structure8937.344.2548
7Synthetic/analytic6145.343.4795
8Voice (diathesis)2638.843.0669
9Word order8037.641.7987
—Format6—————

Paraphrase types

Likelihood summaries; author-clustered bootstrap intervals are in the data files.

About this dashboard — what it shows and how to read it

K = mean over the 50 facts of a split of exp(−mean answer-token negative log-likelihood).

Definitions

On this pageWhat it means
Fine-tuning languageThe single language whose translated TOFU facts a model was fine-tuned on.
Question languageThe language a probe question is asked in; 174 of them, one per column.
Unlearned modelA fine-tuned model after unlearning with the method selected above.
Unlearning languageThe language the unlearning itself was carried out in.
Retain-only oracleTrained identically but never shown the forget facts — the target unlearning aims at.
Base model (no fine-tuning)The untouched checkpoint, before any fine-tuning.
Facts targeted for forgettingThe 50 facts unlearning is asked to remove.
Facts kept50 facts unlearning must leave intact.
Answer strengthMean likelihood of the stored answer over the 50 facts of a split.
Knowledge gained vs base modelK(fine-tuned) − K(base).
Knowledge gained vs retain-only oracleK(fine-tuned) − K(retain-only oracle).
Removed (vs base model)Share of everything fine-tuning added that unlearning took back.
Removed (vs retain-only oracle)Share of the fact-specific knowledge that unlearning took back.
Answer language Q→Q / Q→EN / Q→AWhich language the model is required to answer in.