knayi-myscript · GitHub

Benchmark

How knayi 2.10.0 converts and detects Burmese text on public Zawgyi and Unicode data, next to knayi 2.8.3, myanmar-tools 1.1.3 and Rabbit 1.0.4.

Run on 2026-10-02 · Apple M3 Max, Darwin 27.0.0 · Node v26.5.0 · every data set has an open license (see sources) · see limits

96.6%
Zawgyi → Unicode round trip, exact
knayi 2.8.3: 85.6% · myanmar-tools 1.1.3: 96.9% · 4,745 Wikipedia lines
100.0%
Google reference pairs, NFC
Rabbit 1.0.4: 97.5% · myanmar-tools 1.1.3: 97.5% · 80 pairs
5.0%
Wikipedia lines flagged as Zawgyi
plain fontDetect; on evidence 0.0% · myanmar-tools 1.1.3: 2.3% / 0.7%
11 ms
Slowest long input
knayi 2.8.3: 3,208 ms on the same input

Accuracy

In the conversion and detection tables, the best value in each row is in bold. The tables of text flagged as Zawgyi have no bold: a detector can lower them just by calling Zawgyi less often, so read them next to the detection table. Every set is measured on its distinct lines.

Conversion: Zawgyi → Unicode

Higher is better. Both reference sets come from Google's i18n work, and CLDR's expected output follows ICU, the converter myanmar-tools ships, so myanmar-tools has a home advantage on them. CLDR pairs that repeat Google's file are counted once, in the Google row. "NFC" compares after Unicode NFC normalization, which treats canonically equivalent spellings as equal (ဦ typed as U+1025 U+102E or as U+1026). The round trip turns Wikipedia lines into Zawgyi with Rabbit and converts them back; Rabbit is left out of that row because it made the input. The Wikipedia text has typing errors of its own, mostly ဝ typed for zero in numbers (၁ဝ for ၁၀). Since 2.10 knayi corrects them, and this row counts each correction as a miss.

Datanknayi 2.10.0knayi 2.8.3myanmar-tools 1.1.3Rabbit 1.0.4
exactNFCexactNFCexactNFCexactNFC
google/language-resources reference pairs80100.0%100.0%81.3%88.8%97.5%97.5%97.5%97.5%
CLDR reference pairs not in Google's file (ICU)1172.7%72.7%36.4%45.5%100.0%100.0%45.5%45.5%
Wikipedia → Rabbit Zawgyi → back4,74596.6%96.6%85.6%85.6%96.9%96.9%——

Detection: real text recognised

Higher is better. "On evidence" passes fallback unicode, so a word counts only when the detector finds Zawgyi evidence (for myanmar-tools, p above 0.95). 308 of 2390 distinct WaitZar words read the same in both encodings (neither Rabbit nor myanmar-tools changes them) and are left out.

Datanknayi 2.10.0knayi 2.8.3myanmar-tools 1.1.3
WaitZar hand-typed Zawgyi words, on evidence2,08279.6%79.6%96.5%

Unicode flagged as Zawgyi

Lower is better, but nothing is bolded: a detector can flag less Unicode just by calling Zawgyi less often, so read these next to the detection table. "default" is a plain fontDetect(text), where a tie falls back to zawgyi. "evidence" uses fallback unicode, so only real Zawgyi evidence counts. For myanmar-tools both use the thresholds of knayi's adapter: Zawgyi above p = 0.95, Unicode below 0.05, and the fallback in between. A few lines in these sets are real Zawgyi, so 0% is not always reachable.

Datanknayi 2.10.0knayi 2.8.3myanmar-tools 1.1.3
defaultevidencedefaultevidencedefaultevidence
FLORES-200 mya_Mymr2,0090.0%0.0%0.0%0.0%0.0%0.0%
Burmese Wikipedia sample4,8125.0%0.0%5.9%0.0%2.3%0.7%
Okell corpus16,9246.5%0.0%7.3%0.0%3.3%1.4%

Other Myanmar-script languages flagged as Zawgyi

Lower is better, but nothing is bolded: a detector can flag less Unicode just by calling Zawgyi less often, so read these next to the detection table. "default" is a plain fontDetect(text), where a tie falls back to zawgyi. "evidence" uses fallback unicode, so only real Zawgyi evidence counts. For myanmar-tools both use the thresholds of knayi's adapter: Zawgyi above p = 0.95, Unicode below 0.05, and the fallback in between.

Datanknayi 2.10.0knayi 2.8.3myanmar-tools 1.1.3
defaultevidencedefaultevidencedefaultevidence
Shan (GlotCC shn-Mymr)9,9236.1%0.1%6.2%0.1%0.5%0.1%
Mon (GlotCC mnw-Mymr)2,27010.7%0.6%12.1%0.6%11.5%5.5%
S'gaw Karen (GlotCC ksw-Mymr)67390.3%73.3%95.1%81.9%88.7%82.2%
Pa'o (GlotCC blk-Mymr)77012.5%0.0%13.9%0.0%9.9%3.8%

Web text without labels (mC4 Burmese validation)

Share of lines called Zawgyi on evidence, and agreement with myanmar-tools on the 14,225 lines where myanmar-tools is confident (p < 0.05 or p > 0.95).

Datanknayi 2.10.0knayi 2.8.3myanmar-tools 1.1.3
Called Zawgyi14,30468.6%68.6%69.8%
Agrees with myanmar-tools14,22598.7%98.7%—

Speed

Real text

6,821 distinct lines of FLORES-200 and Wikipedia, and their Zawgyi form made by Rabbit. Mean of 10 runs after 3 warm-ups. Timings depend on the machine, so compare the ratio: below 1 means knayi 2.10.0 is faster.

Taskknayi 2.10.0knayi 2.8.3ratio
fontDetect33.8 ms33.0 ms1.03×
fontConvert Zawgyi → Unicode, source detected112.2 ms131.4 ms0.85×
fontConvert Unicode → Zawgyi107.3 ms109.9 ms0.98×
syllBreak41.9 ms33.6 ms1.25×
normalize84.9 ms164.9 ms0.52×

Long input

Inputs that took quadratic time before 2.9.1. One run each.

Inputknayi 2.10.0knayi 2.8.3
fontConvert Zawgyi → Unicode, stacked ka + 20k alternating vowel signs0 ms416 ms
fontConvert Zawgyi → Unicode, stacked ka + 40k alternating vowel signs1 ms1,637 ms
fontConvert Zawgyi → Unicode, stacked ka + 80k alternating vowel signs1 ms6,493 ms
fontConvert Zawgyi → Unicode, kinzi + 80k alternating vowel signs2 ms6,621 ms
normalize, 50k × ဝ3 ms62 ms
normalize, 100k × ဝ4 ms575 ms
normalize, 200k × ဝ11 ms3,208 ms

Sweep of 1,059 long inputs: every Myanmar code point repeated 30,000 times, every ordered pair of 29 marks repeated 10,000 times, and three base letters each followed by every mark, through 8 call forms. 8,472 runs, slowest 118 ms, 0 over 250 ms.

Limits

Sources and licenses

The data is downloaded when the benchmark runs and is not copied into this repository. Every download must match a pinned sha256; GitHub files are also pinned to a commit, mC4 to a revision and Okell to a Zenodo record. Only aggregate numbers are published here.

DataUsed forSizeLicense
google/language-resources zawgyi_unicode_test.tsvConversion reference pairs80 pairsApache-2.0
Unicode CLDR my-t-my-s0-zawgyi.txtConversion reference pairs (ICU)89 pairs, 11 not in Google's fileUnicode License V3
WaitZar words.zawgyi.txtDetection of hand-typed Zawgyi2,390 distinct wordsApache-2.0
FLORES-200 mya_Mymr (dev + devtest)Unicode flagged as Zawgyi; speed2,009 distinct linesCC BY-SA 4.0
Burmese Wikipedia (wikimedia/wikipedia 20231101.my), 1,000 articles in 25 seeded random blocksUnicode flagged as Zawgyi; round trip; speed1,000 of 109,310 articles (seed 20261002), 4,812 distinct lines (of 7,855)CC BY-SA 3.0 and GFDL
John Okell, A Corpus of Modern BurmeseUnicode flagged as Zawgyi16,924 distinct lines (of 17,828)CC BY 4.0
GlotCC-V1 Shan, Mon, S'gaw Karen, Pa'o (all documents)Other Myanmar-script languages flagged as Zawgyishn 648 documents, mnw 24 documents, ksw 40 documents, blk 34 documentsCC0 1.0 (text from Common Crawl, whose terms of use apply)
mC4 c4-my validation (allenai/c4)Web text without labels14,304 distinct linesODC-BY (text from Common Crawl, whose terms of use apply)

Reproduce

npm run bench:page

That runs npm run eval and npm run bench -- --sweep and rebuilds this page. Raw results: benchmark.json. Method: scripts/eval/README.md.