Skip to content

Methods and provenance

Use this section to verify what each benchmark measures, which inputs it uses, and where its machine-readable evidence is stored. The three evidence streams answer different questions and must not be collapsed into one score.

Evidence stream Question answered Start here Audit trail
Dictionary-reference agreement How closely do outputs reproduce the relations or roots encoded by the evaluated lexical resource? Comparison methodology and linguistic quality Complete quality appendix, corpora, tested stemmers, and raw data
Held-out-family generalization How do learned transformations behave on lexical families excluded from model construction? Generalization protocol Generalization summary, complete 143×10 appendix, and their linked CSV and source manifests
Runtime timing How long do tested implementations take on one recorded workload and environment? Environment and reports JMH CSV/TXT artifacts and checksums linked from the environment and reproducibility pages

The evidence overview explains how to combine these streams without turning speed, reference agreement, and transfer into a single unsupported ranking. The language explorer links the complete per-language evidence, including benchmark-only models that are not distributed.

Specialized protocols

  • Edit-cost protocol defines the frozen cost grid, exact-equivalence classes, selection rule, and limitations.
  • English coverage and speed records a separate model-size/quality/runtime experiment. Its timings are not interchangeable with the multilingual runtime suite.
  • Candidate evaluation records included and skipped candidate families so that absence is not mistaken for a zero result.

Every published claim remains bounded by the versions, source identities, hardware, adapters, and candidate sets recorded with its campaign.