Language Benchmark Pages
This section splits Radixor stemmer benchmark results by language. Each of the 20 registered default models has one language page containing the refreshed corpus, patch-command distribution, exact-root accuracy, runtime performance, and pairwise stemming-quality tables for both dictionary-processing modes.
Reference Pages
| Page |
Purpose |
| Methodology |
Workload design, normalization, speed metrics, and exact-root quality metrics. Pairwise quality definitions are also reproduced on every language page. |
| Corpora |
Dictionary sizes and changed-token timing workloads. |
| Environment and reports |
Hardware, JVM, JMH settings, report files, and badge policy. |
| English dictionary coverage |
Quality/speed operating curve for contracted Radixor tries built from 100% down to 10% of English dictionary rows. |
| Candidate evaluation |
Included and skipped stemmer candidates. |
Languages
Methodology Notes
- Speed benchmarks process only changed dictionary tokens where the surface form differs from the expected root.
- Accuracy benchmarks process the complete dictionary and report
All exact, Changed exact, and Root preserved.
- Radixor speed must be interpreted together with exact-root quality. A slower Radixor row must not be read as a simple performance weakness when Radixor is also the row with accuracy close to 100% and competing stemmers are much lower. Many fast light, minimal, possessive, or aggressive rule-based stemmers are fast because they do much less linguistic work. The measured Radixor cost buys dictionary-trained precision, and that precision is what improves search quality when queries and indexed text are reduced to the same intended roots. The EnglishRadixorDictionaryCoverageBenchmark table shows this contracted-trie operating curve explicitly.
- Results are comparable only within the same language and benchmark family.
- The historical Porter badge is retired; no JMH badge JSON is generated.