Dictionary-family generalization
Radixor learns transformations from lexical families rather than storing a closed word-to-root answer list. This experiment measures how those transformations transfer to dictionary families that were not used to build the Java trie.
For every one of the 143 default models, complete dictionary rows are placed in a
frozen pseudorandom order. Exact-size, nested prefixes retain 10% through 100% of
the rows for training. Five predeclared splits are evaluated against the complete
dictionary; the primary Unseen columns exclude a held-out occurrence whenever its
normalized surface form also appeared in training. This prevents a duplicated form
from being presented as unseen evidence.
The experiment measures within-resource dictionary-family generalization. It does not claim performance on new domains, misspellings, arbitrary compounds, or other out-of-distribution text. See the methodology and limitations.
All-language summary
Each cell is the language-macro mean across languages with a defined denominator, so large dictionaries do not dominate small ones. Changed-form exactness is the most demanding measure because it excludes words whose expected root is already the input token.
| Training rows | Unseen all exact | Unseen changed exact | Unseen root preserved |
|---|---|---|---|
| 100% | n/a | n/a | n/a |
| 90% | 52.61% | 44.54% | 95.46% |
| 80% | 51.73% | 43.77% | 95.11% |
| 70% | 51.33% | 43.33% | 94.83% |
| 60% | 50.55% | 42.55% | 94.65% |
| 50% | 50.01% | 41.99% | 94.15% |
| 40% | 49.46% | 41.40% | 93.90% |
| 30% | 48.48% | 40.32% | 93.64% |
| 20% | 47.04% | 38.80% | 93.69% |
| 10% | 44.68% | 36.22% | 93.20% |
The language range is material and is therefore not replaced by the macro mean. At 10% knowledge, median unseen changed-form exactness ranges from 0.000% for Akan to 94.807% for Bashkir. At 90%, it ranges from 0.000% for Akan to 95.477% for Crimean Tatar. The largest 10%–90% change is +59.536 pp for Cree. These are descriptive within-resource results, not a causal ranking of language, script, dictionary size, or regularity. Every language's evidence-derived endpoint conclusion is also published on its separate language benchmark page, rather than inferred from this macro mean.
Complete evidence
The lossless generalization appendix preserves all 143 languages × 10 training levels, including 100% rows, five-seed medians and ranges, model identity, raw-counter links, and campaign provenance. The language catalog links the evidence-derived conclusion for each model.