Built-in languages and default models
Use this page to choose a language alias, Java enum, default model ID, and aggregate package. The exhaustive artifact provenance and checksums live in the generated model catalog; executable loading examples live in Model Selection and Loading.
“Registered language” means that the repository maintains a dictionary and a
stable Java runtime mapping for that language. Packaging differs by runtime:
Java keeps the core dictionary-free and resolves individual model artifacts,
while the Java and Python radixor-models-standard aggregates install the same
reviewed 31 standard dictionaries.
The Java language enum carries language identity, writing direction, a legacy resource-directory name, and the stable default model ID. A Java model descriptor carries the independently versioned model identity and resource. Python accepts the short alias or the same full model ID. See Model Selection and Loading for Java and Python Usage and API for Python.
Registered languages and models
The Java and Python standard aggregates contain the 31 IDs in models/standard-model-projects.properties. Every other active model is individually published and belongs to the Java-only extended aggregate; filtered alternatives use their separate opt-in Java aggregate.
| Language | Java enum | Model ID | Topology role | Package | Relative dictionary size |
|---|---|---|---|---|---|
| Adyghe | ADY |
ady-default |
standalone |
extended |
(20,347 distinct forms) |
| Afrikaans | AF_ZA |
af-za-default |
standalone |
extended |
(297,154 distinct forms) |
| Aimele | AIL |
ail-default |
standalone |
extended |
(3,061 distinct forms) |
| Akan | AK |
aka-default |
standalone |
extended |
(1,890 distinct forms) |
| Albanian | SQ_AL |
sq-al-default |
standalone |
extended |
(10,218 distinct forms) |
| Alsatian | GSW |
gsw-default |
standalone |
extended |
(1,186 distinct forms) |
| Amharic | AM_ET |
am-et-default |
standalone |
extended |
(41,308 distinct forms) |
| Amuzgo | AZG |
azg-default |
standalone |
extended |
(9,371 distinct forms) |
| Ancient Greek | GRC |
grc-default |
standalone |
extended |
(26,386 distinct forms) |
| Anglo-Norman | XNO |
xno-default |
standalone |
extended |
(186 distinct forms) |
| Arabic | AR |
ar-default |
standalone |
standard |
(452,974 distinct forms) |
| Armenian | HY_AM |
hy-am-default |
standalone |
standard |
(246,576 distinct forms) |
| Ashaninka | CNI |
cni-default |
standalone |
extended |
(10,095 distinct forms) |
| Assamese | AS_IN |
as-in-default |
standalone |
extended |
(67,733 distinct forms) |
| Asturian | AST |
ast-default |
standalone |
extended |
(21,589 distinct forms) |
| Aymara | AYM |
aym-default |
standalone |
extended |
(308,884 distinct forms) |
| Azerbaijani | AZ_AZ |
az-az-default |
standalone |
extended |
(6,669 distinct forms) |
| Bashkir | BAK |
bak-default |
standalone |
extended |
(10,723 distinct forms) |
| Belarusian | BE_BY |
be-by-default |
standalone |
extended |
(19,527 distinct forms) |
| Bengali | BN_BD |
bn-bd-default |
standalone |
extended |
(2,847 distinct forms) |
| Bininj Kun-wok | GUP |
gup-default |
standalone |
extended |
(370 distinct forms) |
| Braj | BRA |
bra-default |
standalone |
extended |
(2,028 distinct forms) |
| Breton | BRE |
bre-default |
standalone |
extended |
(1,940 distinct forms) |
| Bulgarian | BG_BG |
bg-bg-default |
standalone |
extended |
(45,806 distinct forms) |
| Catalan | CA_ES |
ca-es-default |
standalone |
standard |
(130,366 distinct forms) |
| Cebuano | CEB |
ceb-default |
standalone |
extended |
(430 distinct forms) |
| Chichewa | NY_MW |
ny-mw-default |
standalone |
extended |
(3,178 distinct forms) |
| Chichicapan Zapotec | ZPV |
zpv-default |
standalone |
extended |
(818 distinct forms) |
| Chukchi | CKT |
ckt-default |
standalone |
extended |
(302 distinct forms) |
| Church Slavonic | CHU |
chu-default |
standalone |
extended |
(1,652 distinct forms) |
| Classical Armenian | XCL |
xcl-default |
standalone |
extended |
(59,215 distinct forms) |
| Classical Syriac | SYC |
syc-default |
standalone |
extended |
(27,529 distinct forms) |
| Congo Swahili | SWC |
swc-default |
standalone |
extended |
(11,121 distinct forms) |
| Copala Triqui | CPA |
cpa-default |
standalone |
extended |
(2,741 distinct forms) |
| Cornish | COR |
cor-default |
standalone |
extended |
(162 distinct forms) |
| Cree | CRE |
cre-default |
standalone |
extended |
(1,076 distinct forms) |
| Crimean Tatar | CRH |
crh-default |
standalone |
extended |
(7,195 distinct forms) |
| Czech | CS_CZ |
cs-cz-default |
default |
standard |
(51,401 distinct forms) |
| Dakota | DAK |
dak-default |
standalone |
extended |
(3,332 distinct forms) |
| Danish | DA_DK |
da-dk-default |
default |
standard |
(27,921 distinct forms) |
| Dutch | NL_NL |
nl-nl-default |
default |
standard |
(26,201 distinct forms) |
| Eastern Chatino | CLY |
cly-default |
standalone |
extended |
(2,035 distinct forms) |
| Egyptian Arabic | ARZ |
arz-default |
standalone |
extended |
(17,937 distinct forms) |
| English | US_UK |
us-uk-default |
default |
standard |
(591,946 distinct forms) |
| Estonian | ET_EE |
et-ee-default |
standalone |
standard |
(24,811 distinct forms) |
| Evenki | EVN |
evn-default |
standalone |
extended |
(13,549 distinct forms) |
| Faroese | FO_FO |
fo-fo-default |
standalone |
extended |
(31,366 distinct forms) |
| Finnish | FI_FI |
fi-fi-default |
default |
standard |
(1,788,784 distinct forms) |
| French | FR_FR |
fr-fr-default |
default |
standard |
(404,011 distinct forms) |
| Friulian | FUR |
fur-default |
standalone |
extended |
(5,007 distinct forms) |
| Ga | GAA |
gaa-default |
standalone |
extended |
(469 distinct forms) |
| Galolen | GAL |
gal-default |
standalone |
extended |
(25,436 distinct forms) |
| German | DE_DE |
de-de-default |
default |
standard |
(277,266 distinct forms) |
| Gothic | GOT |
got-default |
standalone |
extended |
(134,332 distinct forms) |
| Greek | EL_GR |
el-gr-default |
standalone |
standard |
(76,869 distinct forms) |
| Gulf Arabic | AFB |
afb-default |
standalone |
extended |
(29,344 distinct forms) |
| Haida | HAI |
hai-default |
standalone |
extended |
(5,385 distinct forms) |
| Hebrew | HE_IL |
he-il-default |
default |
standard |
(57,658 distinct forms) |
| Hiligaynon | HIL |
hil-default |
standalone |
extended |
(393 distinct forms) |
| Hsilimo | HSI |
hsi-default |
standalone |
extended |
(158 distinct forms) |
| Hungarian | HU_HU |
hu-hu-default |
default |
standard |
(910,688 distinct forms) |
| Icelandic | IS_IS |
is-is-default |
standalone |
extended |
(52,197 distinct forms) |
| Indonesian | ID_ID |
id-id-default |
standalone |
standard |
(21,296 distinct forms) |
| Ingrian | IZH |
izh-default |
standalone |
extended |
(1,026 distinct forms) |
| Irish | GA_IE |
ga-ie-default |
standalone |
standard |
(24,035 distinct forms) |
| Italian | IT_IT |
it-it-default |
default |
standard |
(324,366 distinct forms) |
| Itelmen | ITL |
itl-default |
standalone |
extended |
(3,659 distinct forms) |
| Japanese | JA_JP |
ja-jp-default |
standalone |
extended |
(10,848 distinct forms) |
| Kabardian | KBD |
kbd-default |
standalone |
extended |
(3,054 distinct forms) |
| Kalaallisut | KL_GL |
kl-gl-default |
standalone |
extended |
(321 distinct forms) |
| Kannada | KN_IN |
kn-in-default |
standalone |
extended |
(3,802 distinct forms) |
| Karelian | KRL |
krl-default |
standalone |
extended |
(566 distinct forms) |
| Kashubian | CSB |
csb-default |
standalone |
extended |
(350 distinct forms) |
| Kazakh | KK_KZ |
kk-kz-default |
standalone |
extended |
(35,223 distinct forms) |
| Khakas | KJH |
kjh-default |
standalone |
extended |
(1,172 distinct forms) |
| Khaling | KLR |
klr-default |
standalone |
extended |
(54,602 distinct forms) |
| Kodi | KOD |
kod-default |
standalone |
extended |
(524 distinct forms) |
| Kongo | KON |
kon-default |
standalone |
extended |
(557 distinct forms) |
| Kyrgyz | KY_KG |
ky-kg-default |
standalone |
extended |
(2,997 distinct forms) |
| Ladin | LLD |
lld-default |
standalone |
extended |
(4,819 distinct forms) |
| Latin | LA |
la-default |
standalone |
extended |
(486,375 distinct forms) |
| Latvian | LV_LV |
lv-lv-default |
standalone |
extended |
(75,492 distinct forms) |
| Lingala | LIN |
lin-default |
standalone |
extended |
(230 distinct forms) |
| Lithuanian | LT_LT |
lt-lt-default |
standalone |
standard |
(28,889 distinct forms) |
| Livonian | LIV |
liv-default |
standalone |
extended |
(2,861 distinct forms) |
| Low German | NDS |
nds-default |
standalone |
extended |
(2,495 distinct forms) |
| Lower Sorbian | DSB |
dsb-default |
standalone |
extended |
(12,039 distinct forms) |
| Luganda | LG_UG |
lg-ug-default |
standalone |
extended |
(4,673 distinct forms) |
| Macedonian | MK_MK |
mk-mk-default |
standalone |
extended |
(135,754 distinct forms) |
| Magahi | MAG |
mag-default |
standalone |
extended |
(1,448 distinct forms) |
| Malagasy | MG_MG |
mg-mg-default |
standalone |
extended |
(636 distinct forms) |
| Maltese | MT_MT |
mt-mt-default |
standalone |
extended |
(1,491 distinct forms) |
| Manx | GV_IM |
gv-im-default |
standalone |
extended |
(15 distinct forms) |
| Maori | MI_NZ |
mi-nz-default |
standalone |
extended |
(207 distinct forms) |
| Mapudungun | ARN |
arn-default |
standalone |
extended |
(548 distinct forms) |
| Middle French | FRM |
frm-default |
standalone |
extended |
(27,102 distinct forms) |
| Middle High German | GMH |
gmh-default |
standalone |
extended |
(384 distinct forms) |
| Middle Low German | GML |
gml-default |
standalone |
extended |
(624 distinct forms) |
| Mongolian | MN_MN |
mn-mn-default |
standalone |
extended |
(17,231 distinct forms) |
| Murrinh-Patha | MWF |
mwf-default |
standalone |
extended |
(592 distinct forms) |
| Navajo | NAV |
nav-default |
standalone |
extended |
(11,046 distinct forms) |
| Neapolitan | NAP |
nap-default |
standalone |
extended |
(1,497 distinct forms) |
| North Frisian | FRR |
frr-default |
standalone |
extended |
(364 distinct forms) |
| Northern Sami | SME |
sme-default |
standalone |
extended |
(54,435 distinct forms) |
| Norwegian Bokmål | NB_NO |
nb-no-default |
default |
standard |
(73,170 distinct forms) |
| Norwegian Nynorsk | NN_NO |
nn-no-default |
default |
standard |
(16,937 distinct forms) |
| Old English | ANG |
ang-default |
standalone |
extended |
(65,202 distinct forms) |
| Old French | FRO |
fro-default |
standalone |
extended |
(94,338 distinct forms) |
| Old High German | GOH |
goh-default |
standalone |
extended |
(5,267 distinct forms) |
| Old Irish | SGA |
sga-default |
standalone |
extended |
(934 distinct forms) |
| Old Norse | NON |
non-default |
standalone |
extended |
(46,089 distinct forms) |
| Old Saxon | OSX |
osx-default |
standalone |
extended |
(12,361 distinct forms) |
| Oodham | OOD |
ood-default |
standalone |
extended |
(1,159 distinct forms) |
| Otomi (ote) | OTE |
ote-default |
standalone |
extended |
(3,417 distinct forms) |
| Pashto | PS_AF |
ps-af-default |
standalone |
extended |
(2,945 distinct forms) |
| Persian | FA_IR |
fa-ir-default |
default |
standard |
(3,544 distinct forms) |
| Polish | PL_PL |
pl-pl-polimorf |
optional |
extended |
(4,668,685 distinct forms) |
| Polish | PL_PL |
pl-pl-unimorph |
default |
standard |
(120,867 distinct forms) |
| Portuguese | PT_PT |
pt-pt-default |
default |
standard |
(211,091 distinct forms) |
| Quechua | QUE |
que-default |
standalone |
extended |
(147,914 distinct forms) |
| Romanian | RO_RO |
ro-ro-default |
standalone |
standard |
(48,565 distinct forms) |
| Russian | RU_RU |
ru-ru-default |
default |
standard |
(759,333 distinct forms) |
| Seneca | SEE |
see-default |
standalone |
extended |
(5,318 distinct forms) |
| Serbo-Croatian | HBS |
hbs-default |
standalone |
extended |
(272,634 distinct forms) |
| Shipibo-Conibo | SHP |
shp-default |
standalone |
extended |
(7,729 distinct forms) |
| Shona | SN_ZW |
sn-zw-default |
standalone |
extended |
(2,640 distinct forms) |
| Sotho, Southern | ST_ZA |
st-za-default |
standalone |
standard |
(416 distinct forms) |
| Southern Kurdish | SDH |
sdh-default |
standalone |
extended |
(165 distinct forms) |
| Spanish | ES_ES |
es-es-default |
default |
standard |
(849,661 distinct forms) |
| Swedish | SV_SE |
sv-se-default |
default |
standard |
(95,181 distinct forms) |
| Tagalog | TL_PH |
tl-ph-default |
standalone |
extended |
(2,510 distinct forms) |
| Turkish | TR_TR |
tr-tr-default |
standalone |
standard |
(222,553 distinct forms) |
| Ukrainian | UK_UA |
uk-ua-default |
default |
standard |
(14,150 distinct forms) |
| Uyghur | UG_CN |
ug-cn-default |
standalone |
extended |
(6,368 distinct forms) |
| Uzbek | UZ_UZ |
uz-uz-default |
standalone |
extended |
(9,835 distinct forms) |
| Voro | VRO |
vro-default |
standalone |
extended |
(412 distinct forms) |
| Western Highland Chatino | CTP |
ctp-default |
standalone |
extended |
(2,597 distinct forms) |
| Xibe | SJO |
sjo-default |
standalone |
extended |
(3,151 distinct forms) |
| Yanesha’ | AME |
ame-default |
standalone |
extended |
(2,635 distinct forms) |
| Yiddish | YI |
yi-default |
default |
standard |
(3,532 distinct forms) |
| Yoloxóchitl Mixtec | XTY |
xty-default |
standalone |
extended |
(2,754 distinct forms) |
| Zarma | DJE |
dje-default |
standalone |
extended |
(81 distinct forms) |
| Zenzontepec Chatino | CZN |
czn-default |
standalone |
extended |
(1,864 distinct forms) |
| Zulu | ZU_ZA |
zu-za-default |
standalone |
extended |
(32,384 distinct forms) |
The maintained table deliberately avoids duplicating mutable provenance fields. Those values come from model module metadata and the generated model catalog.
The Polish dual-model case
PL_PL represents Polish. It is not an alias for either source dictionary.
loadCompiled(Language.PL_PL, ...)resolvespl-pl-unimorph.registry.require("pl-pl-polimorf")resolves the optional PoliMorf model.StemmerPatchTrieLoader.loadCompiled("pl-pl-polimorf", true, reductionMode)constructs its compiled trie explicitly; complete construction is verified with a dedicated 6 GiB test heap.- Both artifacts may be present and loaded independently.
- Adding PoliMorf does not change the language default.
- Radixor does not merge their dictionaries or outputs automatically.
UniMorph and PoliMorf have different lexical sources and provenance. Applications should compare outputs with application-specific regression tests before changing an explicit model choice.
In Python, Stemmer("pl") selects pl-pl-unimorph. The standard Python data
package does not include PoliMorf; applications that need it must compile and
load it explicitly as a trusted custom model. As in Java, it never changes the
Polish default implicitly.
Dependency patterns
Resolve the placeholders below from the Maven Central artifact page and the model catalog, respectively.
Minimal English:
dependencies {
implementation 'org.egothor:radixor:<latest-java-version>'
runtimeOnly 'org.egothor:radixor-model-us-uk-default:<compatible-model-version>'
}
All 31 standard-aggregate defaults:
dependencies {
implementation 'org.egothor:radixor:<latest-java-version>'
runtimeOnly 'org.egothor:radixor-models-standard:<compatible-catalog-version>'
}
The standard pack is metadata-only and excludes optional PoliMorf.
Every individual model artifact carries its own provenance and licensing material. UniMorph models carry model-specific audited licenses and notices because their official language repositories identify different lexical sources, contributors, and applicable terms. Each notice preserves upstream attribution and records the Radixor transformations and Leo Galambos contribution statement. Legacy imports disclose when an exact historical revision was not recorded; this is a reproducibility limitation, not a claim that the source or license is unknown. The active new-model set includes CC BY-SA 3.0, CC BY-SA 4.0, CC BY 4.0, and LGPLLR material; unsupported evidence is quarantined rather than assigned a generic license.
Loading a language default
final FrequencyTrie<CompiledPatchCommand> trie =
StemmerPatchTrieLoader.loadCompiled(
StemmerPatchTrieLoader.Language.US_UK,
true,
ReductionMode.MERGE_SUBTREES_WITH_EQUIVALENT_RANKED_GET_ALL_RESULTS);
The call discovers the default descriptor from the runtime classpath, verifies its compressed resource, parses the GZip UTF-8 dictionary, and constructs a read-only trie. A missing default throws StemmerModelNotFoundException; there is no arbitrary fallback.
Writing direction
Arabic (including Gulf and Egyptian Arabic), Persian, Hebrew, Pashto, Southern Kurdish, Classical Syriac, Uyghur, Urdu, and Yiddish declare right-to-left writing metadata for presentation. Writing direction does not reorder characters in a Java String: all built-in natural-language models therefore use backward traversal from the stored sequence end, where suffixes remain located. Explicit forward traversal is reserved for deliberately prefix-oriented custom data. The selected traversal must remain aligned across dictionary parsing, trie lookup, patch generation, persistence, and application; model identity and writing direction are separate concerns.
Custom and persisted alternatives
Registered model artifacts are a convenient reproducible baseline. Applications may instead load caller-owned textual dictionaries or persist compiled .radixor.gz tries. Those paths are distinct from model artifact discovery:
- a model
stemmer.gzis a compressed textual dictionary plus descriptor/index metadata; - a
.radixor.gzcreated by the binary writer is a persisted compiled trie; - a source dictionary is upstream input, not automatically a valid model artifact.
See Dictionary Format, CLI Compilation, and Stemmer Models.
Benchmark interpretation
Benchmark rows must identify the Radixor model ID used. Default rows use the default IDs above. Optional Polish PoliMorf comparisons must be labeled pl-pl-polimorf; they are not interchangeable with the historical default Polish row. Continue with Benchmarking and Reproducibility.