- 1
Unlearning in one language does not close the other-language routes.
Method a→a removed % ℓ→a removed % ℓ→ℓ removed % other languages ≥90% removed % worst language: access left % SimNPO 97.2 91.0 81.8 30.0 49.2 GradDiff 99.7 96.4 110.5 71.1 26.2 RMU 100.2 91.1 63.5 18.0 75.4 Here a is the fine-tuning language and ℓ any other language (question→answer); removal = 100·(KF − KU)/(KF − KO) with KF, KO and KU the fine-tuned, retain-only-oracle and unlearned answer strengths, access left is 100 minus removal, and values above 100 mean suppression below the retain-only oracle.
Likelihood summaries over 50 forget facts per cell; eligible languages weighted equally within a fine-tuning language, languages within a family, families equally.
Every question language, grouped by writing system
one block per language · hover for its detailThe Language Grid
One square per (model, question language). Hover or focus a square to read it out below; click or press Enter to pin the full comparison into the table further down.
Pinned square
Click any square in the grid to pin every measurement for that fine-tuning language and question language here — all three answer languages, both fact sets, nothing hover-only.
Question form: paraphrases and translated paraphrases
The same 100 facts are asked four ways, with the English answer held fixed: the canonical English question, an English paraphrase, the translated canonical question, and the translated paraphrase. Numbers are teacher-forced gold-answer likelihood. Unlearned rows are SimNPO endpoints trained on three-language source subsets (three subsets × three training orders per model family).
Knowledge gained by fine-tuning, by question form
Acquisition gain = K(fine-tuned) − K(base) on the 0–1 answer-strength scale, English gold answer. For the English rows each language column restricts to the facts matched in that language, so English values vary slightly by column. The five-language mean weights the languages equally.
| Model | Question form | Arabic | Hausa | Hindi | Nepali | Swahili | Five-language mean |
|---|---|---|---|---|---|---|---|
| Aya | English canonical | 0.926 | 0.926 | 0.926 | 0.926 | 0.926 | 0.926 |
| Aya | English paraphrase | 0.890 | 0.891 | 0.893 | 0.891 | 0.892 | 0.892 |
| Aya | Translated canonical | 0.799 | 0.790 | 0.731 | 0.775 | 0.818 | 0.783 |
| Aya | Translated paraphrase | 0.781 | 0.773 | 0.725 | 0.753 | 0.795 | 0.766 |
| Qwen | English canonical | 0.894 | 0.893 | 0.893 | 0.893 | 0.893 | 0.893 |
| Qwen | English paraphrase | 0.855 | 0.858 | 0.858 | 0.857 | 0.856 | 0.857 |
| Qwen | Translated canonical | 0.560 | 0.504 | 0.352 | 0.473 | 0.482 | 0.474 |
| Qwen | Translated paraphrase | 0.514 | 0.488 | 0.396 | 0.451 | 0.463 | 0.462 |
| Llama | English canonical | 0.859 | 0.858 | 0.858 | 0.858 | 0.858 | 0.859 |
| Llama | English paraphrase | 0.826 | 0.826 | 0.826 | 0.825 | 0.825 | 0.826 |
| Llama | Translated canonical | 0.493 | 0.476 | 0.364 | 0.424 | 0.465 | 0.445 |
| Llama | Translated paraphrase | 0.453 | 0.470 | 0.414 | 0.415 | 0.452 | 0.441 |
Access remaining after SimNPO, by question form (%)
Remaining access = 100 · (K(unlearned) − K(oracle)) / (K(fine-tuned) − K(oracle)), unclipped: 0 is the retain-only oracle, 100 the fine-tuned model. Averages the nine SimNPO control endpoints per family. A fact counts in a column when the fine-tuned K ≥ 0.10 and fine-tuned − oracle ≥ 0.05 in all four forms.
| Model | Question form | Arabic | Hausa | Hindi | Nepali | Swahili | Five-language mean |
|---|---|---|---|---|---|---|---|
| Aya | English canonical | 14.8 | 10.0 | 10.2 | 10.5 | 9.9 | 11.1 |
| Aya | English paraphrase | 13.9 | 9.7 | 9.6 | 9.9 | 9.4 | 10.5 |
| Aya | Translated canonical | 8.3 | 3.2 | 9.4 | 4.9 | 3.9 | 5.9 |
| Aya | Translated paraphrase | 7.4 | 4.5 | 5.2 | 0.9 | 2.9 | 4.2 |
| Qwen | English canonical | 32.4 | 32.2 | 32.2 | 32.2 | 35.6 | 32.9 |
| Qwen | English paraphrase | 32.7 | 33.2 | 32.7 | 31.4 | 35.1 | 33.0 |
| Qwen | Translated canonical | 29.2 | 24.8 | 37.3 | 27.3 | 29.6 | 29.6 |
| Qwen | Translated paraphrase | 32.9 | 26.9 | 28.9 | 27.0 | 31.6 | 29.5 |
| Llama | English canonical | 20.8 | 20.4 | 27.8 | 23.0 | 20.6 | 22.5 |
| Llama | English paraphrase | 22.6 | 22.5 | 27.1 | 24.5 | 22.0 | 23.7 |
| Llama | Translated canonical | 28.8 | 27.8 | 31.3 | 22.9 | 24.9 | 27.1 |
| Llama | Translated paraphrase | 29.5 | 26.4 | 27.9 | 23.1 | 28.2 | 27.0 |
English paraphrase types ranked
A win is higher gold-answer likelihood on a shared question; ties count ½. Win % averages matched facts within author, authors within each pair of types, the eight other types within a family, then the three families equally. Before = the fine-tuned models; after = the mean of the nine SimNPO control endpoints. QA = facts with a paraphrase of that type. Format is unranked (six facts).
| Rank | Paraphrase type | QA | Before win % | After win % | Rank after, by model | ||
|---|---|---|---|---|---|---|---|
| Aya | Qwen | Llama | |||||
| 1 | Modal verb | 28 | 77.7 | 67.6 | 1 | 1 | 1 |
| 2 | Subordination | 63 | 54.5 | 58.8 | 3 | 3 | 2 |
| 3 | Sentence modality | 100 | 53.7 | 58.1 | 2 | 2 | 3 |
| 4 | Same-polarity habitual | 91 | 52.1 | 47.4 | 4 | 5 | 6 |
| 5 | Same-polarity contextual | 92 | 53.0 | 45.7 | 8 | 7 | 4 |
| 6 | Syntax/discourse structure | 89 | 37.3 | 44.2 | 5 | 4 | 8 |
| 7 | Synthetic/analytic | 61 | 45.3 | 43.4 | 7 | 9 | 5 |
| 8 | Voice (diathesis) | 26 | 38.8 | 43.0 | 6 | 6 | 9 |
| 9 | Word order | 80 | 37.6 | 41.7 | 9 | 8 | 7 |
| — | Format | 6 | — | — | — | — | — |
Paraphrase types
- Modal verb: expresses the same proposition with a different modal construction.
- Subordination: reorganizes information across main, subordinate or embedded clauses.
- Sentence modality: changes sentence form (declarative, interrogative, imperative).
- Same-polarity habitual: conventional synonym substitution.
- Same-polarity contextual: context-specific synonym substitution.
- Syntax/discourse structure: reorganizes information across clauses or sentences.
- Synthetic/analytic: switches between one word and an equivalent multi-word construction.
- Voice (diathesis): changes how semantic roles map to syntax, e.g. active to passive.
- Word order: reorders words, phrases, clauses or information.
- Format: represents the same information in a different textual format.
Likelihood summaries; author-clustered bootstrap intervals are in the data files.
About this dashboard — what it shows and how to read it
K = mean over the 50 facts of a split of exp(−mean answer-token negative log-likelihood).
Definitions
| On this page | What it means |
|---|---|
| Fine-tuning language | The single language whose translated TOFU facts a model was fine-tuned on. |
| Question language | The language a probe question is asked in; 174 of them, one per column. |
| Unlearned model | A fine-tuned model after unlearning with the method selected above. |
| Unlearning language | The language the unlearning itself was carried out in. |
| Retain-only oracle | Trained identically but never shown the forget facts — the target unlearning aims at. |
| Base model (no fine-tuning) | The untouched checkpoint, before any fine-tuning. |
| Facts targeted for forgetting | The 50 facts unlearning is asked to remove. |
| Facts kept | 50 facts unlearning must leave intact. |
| Answer strength | Mean likelihood of the stored answer over the 50 facts of a split. |
| Knowledge gained vs base model | K(fine-tuned) − K(base). |
| Knowledge gained vs retain-only oracle | K(fine-tuned) − K(retain-only oracle). |
| Removed (vs base model) | Share of everything fine-tuning added that unlearning took back. |
| Removed (vs retain-only oracle) | Share of the fact-specific knowledge that unlearning took back. |
| Answer language Q→Q / Q→EN / Q→A | Which language the model is required to answer in. |