Skip to content

Built-in languages and default models

Use this page to choose a language alias, Java enum, default model ID, and aggregate package. The exhaustive artifact provenance and checksums live in the generated model catalog; executable loading examples live in Model Selection and Loading.

“Registered language” means that the repository maintains a dictionary and a stable Java runtime mapping for that language. Packaging differs by runtime: Java keeps the core dictionary-free and resolves individual model artifacts, while the Java and Python radixor-models-standard aggregates install the same reviewed 31 standard dictionaries.

The Java language enum carries language identity, writing direction, a legacy resource-directory name, and the stable default model ID. A Java model descriptor carries the independently versioned model identity and resource. Python accepts the short alias or the same full model ID. See Model Selection and Loading for Java and Python Usage and API for Python.

Registered languages and models

The Java and Python standard aggregates contain the 31 IDs in models/standard-model-projects.properties. Every other active model is individually published and belongs to the Java-only extended aggregate; filtered alternatives use their separate opt-in Java aggregate.

Language Java enum Model ID Topology role Package Relative dictionary size
Adyghe ADY ady-default standalone extended ★★★★☆ (20,347 distinct forms)
Afrikaans AF_ZA af-za-default standalone extended ★★★★★ (297,154 distinct forms)
Aimele AIL ail-default standalone extended ★★☆☆☆ (3,061 distinct forms)
Akan AK aka-default standalone extended ★★☆☆☆ (1,890 distinct forms)
Albanian SQ_AL sq-al-default standalone extended ★★★☆☆ (10,218 distinct forms)
Alsatian GSW gsw-default standalone extended ★★☆☆☆ (1,186 distinct forms)
Amharic AM_ET am-et-default standalone extended ★★★★☆ (41,308 distinct forms)
Amuzgo AZG azg-default standalone extended ★★★☆☆ (9,371 distinct forms)
Ancient Greek GRC grc-default standalone extended ★★★★☆ (26,386 distinct forms)
Anglo-Norman XNO xno-default standalone extended ★☆☆☆☆ (186 distinct forms)
Arabic AR ar-default standalone standard ★★★★★ (452,974 distinct forms)
Armenian HY_AM hy-am-default standalone standard ★★★★★ (246,576 distinct forms)
Ashaninka CNI cni-default standalone extended ★★★☆☆ (10,095 distinct forms)
Assamese AS_IN as-in-default standalone extended ★★★★★ (67,733 distinct forms)
Asturian AST ast-default standalone extended ★★★★☆ (21,589 distinct forms)
Aymara AYM aym-default standalone extended ★★★★★ (308,884 distinct forms)
Azerbaijani AZ_AZ az-az-default standalone extended ★★★☆☆ (6,669 distinct forms)
Bashkir BAK bak-default standalone extended ★★★☆☆ (10,723 distinct forms)
Belarusian BE_BY be-by-default standalone extended ★★★★☆ (19,527 distinct forms)
Bengali BN_BD bn-bd-default standalone extended ★★☆☆☆ (2,847 distinct forms)
Bininj Kun-wok GUP gup-default standalone extended ★☆☆☆☆ (370 distinct forms)
Braj BRA bra-default standalone extended ★★☆☆☆ (2,028 distinct forms)
Breton BRE bre-default standalone extended ★★☆☆☆ (1,940 distinct forms)
Bulgarian BG_BG bg-bg-default standalone extended ★★★★☆ (45,806 distinct forms)
Catalan CA_ES ca-es-default standalone standard ★★★★★ (130,366 distinct forms)
Cebuano CEB ceb-default standalone extended ★☆☆☆☆ (430 distinct forms)
Chichewa NY_MW ny-mw-default standalone extended ★★★☆☆ (3,178 distinct forms)
Chichicapan Zapotec ZPV zpv-default standalone extended ★★☆☆☆ (818 distinct forms)
Chukchi CKT ckt-default standalone extended ★☆☆☆☆ (302 distinct forms)
Church Slavonic CHU chu-default standalone extended ★★☆☆☆ (1,652 distinct forms)
Classical Armenian XCL xcl-default standalone extended ★★★★☆ (59,215 distinct forms)
Classical Syriac SYC syc-default standalone extended ★★★★☆ (27,529 distinct forms)
Congo Swahili SWC swc-default standalone extended ★★★☆☆ (11,121 distinct forms)
Copala Triqui CPA cpa-default standalone extended ★★☆☆☆ (2,741 distinct forms)
Cornish COR cor-default standalone extended ★☆☆☆☆ (162 distinct forms)
Cree CRE cre-default standalone extended ★★☆☆☆ (1,076 distinct forms)
Crimean Tatar CRH crh-default standalone extended ★★★☆☆ (7,195 distinct forms)
Czech CS_CZ cs-cz-default default standard ★★★★☆ (51,401 distinct forms)
Dakota DAK dak-default standalone extended ★★★☆☆ (3,332 distinct forms)
Danish DA_DK da-dk-default default standard ★★★★☆ (27,921 distinct forms)
Dutch NL_NL nl-nl-default default standard ★★★★☆ (26,201 distinct forms)
Eastern Chatino CLY cly-default standalone extended ★★☆☆☆ (2,035 distinct forms)
Egyptian Arabic ARZ arz-default standalone extended ★★★★☆ (17,937 distinct forms)
English US_UK us-uk-default default standard ★★★★★ (591,946 distinct forms)
Estonian ET_EE et-ee-default standalone standard ★★★★☆ (24,811 distinct forms)
Evenki EVN evn-default standalone extended ★★★☆☆ (13,549 distinct forms)
Faroese FO_FO fo-fo-default standalone extended ★★★★☆ (31,366 distinct forms)
Finnish FI_FI fi-fi-default default standard ★★★★★ (1,788,784 distinct forms)
French FR_FR fr-fr-default default standard ★★★★★ (404,011 distinct forms)
Friulian FUR fur-default standalone extended ★★★☆☆ (5,007 distinct forms)
Ga GAA gaa-default standalone extended ★☆☆☆☆ (469 distinct forms)
Galolen GAL gal-default standalone extended ★★★★☆ (25,436 distinct forms)
German DE_DE de-de-default default standard ★★★★★ (277,266 distinct forms)
Gothic GOT got-default standalone extended ★★★★★ (134,332 distinct forms)
Greek EL_GR el-gr-default standalone standard ★★★★★ (76,869 distinct forms)
Gulf Arabic AFB afb-default standalone extended ★★★★☆ (29,344 distinct forms)
Haida HAI hai-default standalone extended ★★★☆☆ (5,385 distinct forms)
Hebrew HE_IL he-il-default default standard ★★★★☆ (57,658 distinct forms)
Hiligaynon HIL hil-default standalone extended ★☆☆☆☆ (393 distinct forms)
Hsilimo HSI hsi-default standalone extended ★☆☆☆☆ (158 distinct forms)
Hungarian HU_HU hu-hu-default default standard ★★★★★ (910,688 distinct forms)
Icelandic IS_IS is-is-default standalone extended ★★★★☆ (52,197 distinct forms)
Indonesian ID_ID id-id-default standalone standard ★★★★☆ (21,296 distinct forms)
Ingrian IZH izh-default standalone extended ★★☆☆☆ (1,026 distinct forms)
Irish GA_IE ga-ie-default standalone standard ★★★★☆ (24,035 distinct forms)
Italian IT_IT it-it-default default standard ★★★★★ (324,366 distinct forms)
Itelmen ITL itl-default standalone extended ★★★☆☆ (3,659 distinct forms)
Japanese JA_JP ja-jp-default standalone extended ★★★☆☆ (10,848 distinct forms)
Kabardian KBD kbd-default standalone extended ★★☆☆☆ (3,054 distinct forms)
Kalaallisut KL_GL kl-gl-default standalone extended ★☆☆☆☆ (321 distinct forms)
Kannada KN_IN kn-in-default standalone extended ★★★☆☆ (3,802 distinct forms)
Karelian KRL krl-default standalone extended ★☆☆☆☆ (566 distinct forms)
Kashubian CSB csb-default standalone extended ★☆☆☆☆ (350 distinct forms)
Kazakh KK_KZ kk-kz-default standalone extended ★★★★☆ (35,223 distinct forms)
Khakas KJH kjh-default standalone extended ★★☆☆☆ (1,172 distinct forms)
Khaling KLR klr-default standalone extended ★★★★☆ (54,602 distinct forms)
Kodi KOD kod-default standalone extended ★☆☆☆☆ (524 distinct forms)
Kongo KON kon-default standalone extended ★☆☆☆☆ (557 distinct forms)
Kyrgyz KY_KG ky-kg-default standalone extended ★★☆☆☆ (2,997 distinct forms)
Ladin LLD lld-default standalone extended ★★★☆☆ (4,819 distinct forms)
Latin LA la-default standalone extended ★★★★★ (486,375 distinct forms)
Latvian LV_LV lv-lv-default standalone extended ★★★★★ (75,492 distinct forms)
Lingala LIN lin-default standalone extended ★☆☆☆☆ (230 distinct forms)
Lithuanian LT_LT lt-lt-default standalone standard ★★★★☆ (28,889 distinct forms)
Livonian LIV liv-default standalone extended ★★☆☆☆ (2,861 distinct forms)
Low German NDS nds-default standalone extended ★★☆☆☆ (2,495 distinct forms)
Lower Sorbian DSB dsb-default standalone extended ★★★☆☆ (12,039 distinct forms)
Luganda LG_UG lg-ug-default standalone extended ★★★☆☆ (4,673 distinct forms)
Macedonian MK_MK mk-mk-default standalone extended ★★★★★ (135,754 distinct forms)
Magahi MAG mag-default standalone extended ★★☆☆☆ (1,448 distinct forms)
Malagasy MG_MG mg-mg-default standalone extended ★☆☆☆☆ (636 distinct forms)
Maltese MT_MT mt-mt-default standalone extended ★★☆☆☆ (1,491 distinct forms)
Manx GV_IM gv-im-default standalone extended ★☆☆☆☆ (15 distinct forms)
Maori MI_NZ mi-nz-default standalone extended ★☆☆☆☆ (207 distinct forms)
Mapudungun ARN arn-default standalone extended ★☆☆☆☆ (548 distinct forms)
Middle French FRM frm-default standalone extended ★★★★☆ (27,102 distinct forms)
Middle High German GMH gmh-default standalone extended ★☆☆☆☆ (384 distinct forms)
Middle Low German GML gml-default standalone extended ★☆☆☆☆ (624 distinct forms)
Mongolian MN_MN mn-mn-default standalone extended ★★★★☆ (17,231 distinct forms)
Murrinh-Patha MWF mwf-default standalone extended ★☆☆☆☆ (592 distinct forms)
Navajo NAV nav-default standalone extended ★★★☆☆ (11,046 distinct forms)
Neapolitan NAP nap-default standalone extended ★★☆☆☆ (1,497 distinct forms)
North Frisian FRR frr-default standalone extended ★☆☆☆☆ (364 distinct forms)
Northern Sami SME sme-default standalone extended ★★★★☆ (54,435 distinct forms)
Norwegian Bokmål NB_NO nb-no-default default standard ★★★★★ (73,170 distinct forms)
Norwegian Nynorsk NN_NO nn-no-default default standard ★★★★☆ (16,937 distinct forms)
Old English ANG ang-default standalone extended ★★★★☆ (65,202 distinct forms)
Old French FRO fro-default standalone extended ★★★★★ (94,338 distinct forms)
Old High German GOH goh-default standalone extended ★★★☆☆ (5,267 distinct forms)
Old Irish SGA sga-default standalone extended ★★☆☆☆ (934 distinct forms)
Old Norse NON non-default standalone extended ★★★★☆ (46,089 distinct forms)
Old Saxon OSX osx-default standalone extended ★★★☆☆ (12,361 distinct forms)
Oodham OOD ood-default standalone extended ★★☆☆☆ (1,159 distinct forms)
Otomi (ote) OTE ote-default standalone extended ★★★☆☆ (3,417 distinct forms)
Pashto PS_AF ps-af-default standalone extended ★★☆☆☆ (2,945 distinct forms)
Persian FA_IR fa-ir-default default standard ★★★☆☆ (3,544 distinct forms)
Polish PL_PL pl-pl-polimorf optional extended ★★★★★ (4,668,685 distinct forms)
Polish PL_PL pl-pl-unimorph default standard ★★★★★ (120,867 distinct forms)
Portuguese PT_PT pt-pt-default default standard ★★★★★ (211,091 distinct forms)
Quechua QUE que-default standalone extended ★★★★★ (147,914 distinct forms)
Romanian RO_RO ro-ro-default standalone standard ★★★★☆ (48,565 distinct forms)
Russian RU_RU ru-ru-default default standard ★★★★★ (759,333 distinct forms)
Seneca SEE see-default standalone extended ★★★☆☆ (5,318 distinct forms)
Serbo-Croatian HBS hbs-default standalone extended ★★★★★ (272,634 distinct forms)
Shipibo-Conibo SHP shp-default standalone extended ★★★☆☆ (7,729 distinct forms)
Shona SN_ZW sn-zw-default standalone extended ★★☆☆☆ (2,640 distinct forms)
Sotho, Southern ST_ZA st-za-default standalone standard ★☆☆☆☆ (416 distinct forms)
Southern Kurdish SDH sdh-default standalone extended ★☆☆☆☆ (165 distinct forms)
Spanish ES_ES es-es-default default standard ★★★★★ (849,661 distinct forms)
Swedish SV_SE sv-se-default default standard ★★★★★ (95,181 distinct forms)
Tagalog TL_PH tl-ph-default standalone extended ★★☆☆☆ (2,510 distinct forms)
Turkish TR_TR tr-tr-default standalone standard ★★★★★ (222,553 distinct forms)
Ukrainian UK_UA uk-ua-default default standard ★★★★☆ (14,150 distinct forms)
Uyghur UG_CN ug-cn-default standalone extended ★★★☆☆ (6,368 distinct forms)
Uzbek UZ_UZ uz-uz-default standalone extended ★★★☆☆ (9,835 distinct forms)
Voro VRO vro-default standalone extended ★☆☆☆☆ (412 distinct forms)
Western Highland Chatino CTP ctp-default standalone extended ★★☆☆☆ (2,597 distinct forms)
Xibe SJO sjo-default standalone extended ★★★☆☆ (3,151 distinct forms)
Yanesha’ AME ame-default standalone extended ★★☆☆☆ (2,635 distinct forms)
Yiddish YI yi-default default standard ★★★☆☆ (3,532 distinct forms)
Yoloxóchitl Mixtec XTY xty-default standalone extended ★★☆☆☆ (2,754 distinct forms)
Zarma DJE dje-default standalone extended ★☆☆☆☆ (81 distinct forms)
Zenzontepec Chatino CZN czn-default standalone extended ★★☆☆☆ (1,864 distinct forms)
Zulu ZU_ZA zu-za-default standalone extended ★★★★☆ (32,384 distinct forms)

The maintained table deliberately avoids duplicating mutable provenance fields. Those values come from model module metadata and the generated model catalog.

The Polish dual-model case

PL_PL represents Polish. It is not an alias for either source dictionary.

  • loadCompiled(Language.PL_PL, ...) resolves pl-pl-unimorph.
  • registry.require("pl-pl-polimorf") resolves the optional PoliMorf model.
  • StemmerPatchTrieLoader.loadCompiled("pl-pl-polimorf", true, reductionMode) constructs its compiled trie explicitly; complete construction is verified with a dedicated 6 GiB test heap.
  • Both artifacts may be present and loaded independently.
  • Adding PoliMorf does not change the language default.
  • Radixor does not merge their dictionaries or outputs automatically.

UniMorph and PoliMorf have different lexical sources and provenance. Applications should compare outputs with application-specific regression tests before changing an explicit model choice.

In Python, Stemmer("pl") selects pl-pl-unimorph. The standard Python data package does not include PoliMorf; applications that need it must compile and load it explicitly as a trusted custom model. As in Java, it never changes the Polish default implicitly.

Dependency patterns

Resolve the placeholders below from the Maven Central artifact page and the model catalog, respectively.

Minimal English:

dependencies {
    implementation 'org.egothor:radixor:<latest-java-version>'
    runtimeOnly 'org.egothor:radixor-model-us-uk-default:<compatible-model-version>'
}

All 31 standard-aggregate defaults:

dependencies {
    implementation 'org.egothor:radixor:<latest-java-version>'
    runtimeOnly 'org.egothor:radixor-models-standard:<compatible-catalog-version>'
}

The standard pack is metadata-only and excludes optional PoliMorf.

Every individual model artifact carries its own provenance and licensing material. UniMorph models carry model-specific audited licenses and notices because their official language repositories identify different lexical sources, contributors, and applicable terms. Each notice preserves upstream attribution and records the Radixor transformations and Leo Galambos contribution statement. Legacy imports disclose when an exact historical revision was not recorded; this is a reproducibility limitation, not a claim that the source or license is unknown. The active new-model set includes CC BY-SA 3.0, CC BY-SA 4.0, CC BY 4.0, and LGPLLR material; unsupported evidence is quarantined rather than assigned a generic license.

Loading a language default

final FrequencyTrie<CompiledPatchCommand> trie =
        StemmerPatchTrieLoader.loadCompiled(
                StemmerPatchTrieLoader.Language.US_UK,
                true,
                ReductionMode.MERGE_SUBTREES_WITH_EQUIVALENT_RANKED_GET_ALL_RESULTS);

The call discovers the default descriptor from the runtime classpath, verifies its compressed resource, parses the GZip UTF-8 dictionary, and constructs a read-only trie. A missing default throws StemmerModelNotFoundException; there is no arbitrary fallback.

Writing direction

Arabic (including Gulf and Egyptian Arabic), Persian, Hebrew, Pashto, Southern Kurdish, Classical Syriac, Uyghur, Urdu, and Yiddish declare right-to-left writing metadata for presentation. Writing direction does not reorder characters in a Java String: all built-in natural-language models therefore use backward traversal from the stored sequence end, where suffixes remain located. Explicit forward traversal is reserved for deliberately prefix-oriented custom data. The selected traversal must remain aligned across dictionary parsing, trie lookup, patch generation, persistence, and application; model identity and writing direction are separate concerns.

Custom and persisted alternatives

Registered model artifacts are a convenient reproducible baseline. Applications may instead load caller-owned textual dictionaries or persist compiled .radixor.gz tries. Those paths are distinct from model artifact discovery:

  • a model stemmer.gz is a compressed textual dictionary plus descriptor/index metadata;
  • a .radixor.gz created by the binary writer is a persisted compiled trie;
  • a source dictionary is upstream input, not automatically a valid model artifact.

See Dictionary Format, CLI Compilation, and Stemmer Models.

Benchmark interpretation

Benchmark rows must identify the Radixor model ID used. Default rows use the default IDs above. Optional Polish PoliMorf comparisons must be labeled pl-pl-polimorf; they are not interchangeable with the historical default Polish row. Continue with Benchmarking and Reproducibility.