Benchmark
How knayi 2.10.0 converts and detects Burmese text on public Zawgyi and Unicode data, next to knayi 2.8.3, myanmar-tools 1.1.3 and Rabbit 1.0.4.
Accuracy
In the conversion and detection tables, the best value in each row is in bold. The tables of text flagged as Zawgyi have no bold: a detector can lower them just by calling Zawgyi less often, so read them next to the detection table. Every set is measured on its distinct lines.
Conversion: Zawgyi → Unicode
Higher is better. Both reference sets come from Google's i18n work, and CLDR's expected output follows ICU, the converter myanmar-tools ships, so myanmar-tools has a home advantage on them. CLDR pairs that repeat Google's file are counted once, in the Google row. "NFC" compares after Unicode NFC normalization, which treats canonically equivalent spellings as equal (ဦ typed as U+1025 U+102E or as U+1026). The round trip turns Wikipedia lines into Zawgyi with Rabbit and converts them back; Rabbit is left out of that row because it made the input. The Wikipedia text has typing errors of its own, mostly ဝ typed for zero in numbers (၁ဝ for ၁၀). Since 2.10 knayi corrects them, and this row counts each correction as a miss.
| Data | n | knayi 2.10.0 | knayi 2.8.3 | myanmar-tools 1.1.3 | Rabbit 1.0.4 | ||||
|---|---|---|---|---|---|---|---|---|---|
| exact | NFC | exact | NFC | exact | NFC | exact | NFC | ||
| google/language-resources reference pairs | 80 | 100.0% | 100.0% | 81.3% | 88.8% | 97.5% | 97.5% | 97.5% | 97.5% |
| CLDR reference pairs not in Google's file (ICU) | 11 | 72.7% | 72.7% | 36.4% | 45.5% | 100.0% | 100.0% | 45.5% | 45.5% |
| Wikipedia → Rabbit Zawgyi → back | 4,745 | 96.6% | 96.6% | 85.6% | 85.6% | 96.9% | 96.9% | — | — |
Detection: real text recognised
Higher is better. "On evidence" passes fallback unicode, so a word counts only when the detector finds Zawgyi evidence (for myanmar-tools, p above 0.95). 308 of 2390 distinct WaitZar words read the same in both encodings (neither Rabbit nor myanmar-tools changes them) and are left out.
| Data | n | knayi 2.10.0 | knayi 2.8.3 | myanmar-tools 1.1.3 |
|---|---|---|---|---|
| WaitZar hand-typed Zawgyi words, on evidence | 2,082 | 79.6% | 79.6% | 96.5% |
Unicode flagged as Zawgyi
Lower is better, but nothing is bolded: a detector can flag less Unicode just by calling Zawgyi less often, so read these next to the detection table. "default" is a plain fontDetect(text), where a tie falls back to zawgyi. "evidence" uses fallback unicode, so only real Zawgyi evidence counts. For myanmar-tools both use the thresholds of knayi's adapter: Zawgyi above p = 0.95, Unicode below 0.05, and the fallback in between. A few lines in these sets are real Zawgyi, so 0% is not always reachable.
| Data | n | knayi 2.10.0 | knayi 2.8.3 | myanmar-tools 1.1.3 | |||
|---|---|---|---|---|---|---|---|
| default | evidence | default | evidence | default | evidence | ||
| FLORES-200 mya_Mymr | 2,009 | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| Burmese Wikipedia sample | 4,812 | 5.0% | 0.0% | 5.9% | 0.0% | 2.3% | 0.7% |
| Okell corpus | 16,924 | 6.5% | 0.0% | 7.3% | 0.0% | 3.3% | 1.4% |
Other Myanmar-script languages flagged as Zawgyi
Lower is better, but nothing is bolded: a detector can flag less Unicode just by calling Zawgyi less often, so read these next to the detection table. "default" is a plain fontDetect(text), where a tie falls back to zawgyi. "evidence" uses fallback unicode, so only real Zawgyi evidence counts. For myanmar-tools both use the thresholds of knayi's adapter: Zawgyi above p = 0.95, Unicode below 0.05, and the fallback in between.
| Data | n | knayi 2.10.0 | knayi 2.8.3 | myanmar-tools 1.1.3 | |||
|---|---|---|---|---|---|---|---|
| default | evidence | default | evidence | default | evidence | ||
| Shan (GlotCC shn-Mymr) | 9,923 | 6.1% | 0.1% | 6.2% | 0.1% | 0.5% | 0.1% |
| Mon (GlotCC mnw-Mymr) | 2,270 | 10.7% | 0.6% | 12.1% | 0.6% | 11.5% | 5.5% |
| S'gaw Karen (GlotCC ksw-Mymr) | 673 | 90.3% | 73.3% | 95.1% | 81.9% | 88.7% | 82.2% |
| Pa'o (GlotCC blk-Mymr) | 770 | 12.5% | 0.0% | 13.9% | 0.0% | 9.9% | 3.8% |
Web text without labels (mC4 Burmese validation)
Share of lines called Zawgyi on evidence, and agreement with myanmar-tools on the 14,225 lines where myanmar-tools is confident (p < 0.05 or p > 0.95).
| Data | n | knayi 2.10.0 | knayi 2.8.3 | myanmar-tools 1.1.3 |
|---|---|---|---|---|
| Called Zawgyi | 14,304 | 68.6% | 68.6% | 69.8% |
| Agrees with myanmar-tools | 14,225 | 98.7% | 98.7% | — |
Speed
Real text
6,821 distinct lines of FLORES-200 and Wikipedia, and their Zawgyi form made by Rabbit. Mean of 10 runs after 3 warm-ups. Timings depend on the machine, so compare the ratio: below 1 means knayi 2.10.0 is faster.
| Task | knayi 2.10.0 | knayi 2.8.3 | ratio |
|---|---|---|---|
| fontDetect | 33.8 ms | 33.0 ms | 1.03× |
| fontConvert Zawgyi → Unicode, source detected | 112.2 ms | 131.4 ms | 0.85× |
| fontConvert Unicode → Zawgyi | 107.3 ms | 109.9 ms | 0.98× |
| syllBreak | 41.9 ms | 33.6 ms | 1.25× |
| normalize | 84.9 ms | 164.9 ms | 0.52× |
Long input
Inputs that took quadratic time before 2.9.1. One run each.
| Input | knayi 2.10.0 | knayi 2.8.3 |
|---|---|---|
| fontConvert Zawgyi → Unicode, stacked ka + 20k alternating vowel signs | 0 ms | 416 ms |
| fontConvert Zawgyi → Unicode, stacked ka + 40k alternating vowel signs | 1 ms | 1,637 ms |
| fontConvert Zawgyi → Unicode, stacked ka + 80k alternating vowel signs | 1 ms | 6,493 ms |
| fontConvert Zawgyi → Unicode, kinzi + 80k alternating vowel signs | 2 ms | 6,621 ms |
| normalize, 50k × ဝ | 3 ms | 62 ms |
| normalize, 100k × ဝ | 4 ms | 575 ms |
| normalize, 200k × ဝ | 11 ms | 3,208 ms |
Sweep of 1,059 long inputs: every Myanmar code point repeated 30,000 times, every ordered pair of 29 marks repeated 10,000 times, and three base letters each followed by every mark, through 8 call forms. 8,472 runs, slowest 118 ms, 0 over 250 ms.
Limits
- The reference pairs are few, and both sets come from Google's i18n work. CLDR's expected output follows ICU, the converter myanmar-tools ships, and in at least one pair it expects ICU's own ordering of asat before tall aa.
- Exact match counts canonically equivalent spellings as different; the NFC column does not.
- The round trip uses Rabbit to make the Zawgyi input, so Rabbit is not scored on it.
- The Unicode sets contain a few lines of real Zawgyi text, so their labels are slightly noisy.
- The Wikipedia and GlotCC rows come from the current revision of those datasets; the Wikipedia rows are 25 random blocks picked with a fixed seed.
- Timings are from one machine. Compare ratios, not absolute times.
Sources and licenses
The data is downloaded when the benchmark runs and is not copied into this repository. Every download must match a pinned sha256; GitHub files are also pinned to a commit, mC4 to a revision and Okell to a Zenodo record. Only aggregate numbers are published here.
| Data | Used for | Size | License |
|---|---|---|---|
| google/language-resources zawgyi_unicode_test.tsv | Conversion reference pairs | 80 pairs | Apache-2.0 |
| Unicode CLDR my-t-my-s0-zawgyi.txt | Conversion reference pairs (ICU) | 89 pairs, 11 not in Google's file | Unicode License V3 |
| WaitZar words.zawgyi.txt | Detection of hand-typed Zawgyi | 2,390 distinct words | Apache-2.0 |
| FLORES-200 mya_Mymr (dev + devtest) | Unicode flagged as Zawgyi; speed | 2,009 distinct lines | CC BY-SA 4.0 |
| Burmese Wikipedia (wikimedia/wikipedia 20231101.my), 1,000 articles in 25 seeded random blocks | Unicode flagged as Zawgyi; round trip; speed | 1,000 of 109,310 articles (seed 20261002), 4,812 distinct lines (of 7,855) | CC BY-SA 3.0 and GFDL |
| John Okell, A Corpus of Modern Burmese | Unicode flagged as Zawgyi | 16,924 distinct lines (of 17,828) | CC BY 4.0 |
| GlotCC-V1 Shan, Mon, S'gaw Karen, Pa'o (all documents) | Other Myanmar-script languages flagged as Zawgyi | shn 648 documents, mnw 24 documents, ksw 40 documents, blk 34 documents | CC0 1.0 (text from Common Crawl, whose terms of use apply) |
| mC4 c4-my validation (allenai/c4) | Web text without labels | 14,304 distinct lines | ODC-BY (text from Common Crawl, whose terms of use apply) |
Reproduce
npm run bench:page
That runs npm run eval and npm run bench -- --sweep and rebuilds this page. Raw results: benchmark.json. Method: scripts/eval/README.md.