Methods and provenance
Use this section to verify what each benchmark measures, which inputs it uses, and where its machine-readable evidence is stored. The three evidence streams answer different questions and must not be collapsed into one score.
| Evidence stream | Question answered | Start here | Audit trail |
|---|---|---|---|
| Dictionary-reference agreement | How closely do outputs reproduce the relations or roots encoded by the evaluated lexical resource? | Comparison methodology and linguistic quality | Complete quality appendix, corpora, tested stemmers, and raw data |
| Held-out-family generalization | How do learned transformations behave on lexical families excluded from model construction? | Generalization protocol | Generalization summary, complete 143×10 appendix, and their linked CSV and source manifests |
| Runtime timing | How long do tested implementations take on one recorded workload and environment? | Environment and reports | JMH CSV/TXT artifacts and checksums linked from the environment and reproducibility pages |
The evidence overview explains how to combine these streams without turning speed, reference agreement, and transfer into a single unsupported ranking. The language explorer links the complete per-language evidence, including benchmark-only models that are not distributed.
Specialized protocols
- Edit-cost protocol defines the frozen cost grid, exact-equivalence classes, selection rule, and limitations.
- English coverage and speed records a separate model-size/quality/runtime experiment. Its timings are not interchangeable with the multilingual runtime suite.
- Candidate evaluation records included and skipped candidate families so that absence is not mistaken for a zero result.
Every published claim remains bounded by the versions, source identities, hardware, adapters, and candidate sets recorded with its campaign.