Skip to content

Repository files navigation

knayi-myscript

JavaScript library for Myanmar (Burmese) text stored as Unicode or Zawgyi. Version 2.10.0. MIT license.

It detects the encoding, converts between them, inserts syllable breaks, collapses repeated spelling marks, normalizes some Unicode typing errors, and truncates on those breaks. It does not segment dictionary words, translate, or tokenize for a language model.

Try every function in the browser at https://greenlikeorange.github.io/knayi-myscript/.

Install it from npm. npm, Yarn, pnpm, and Bun all read that registry.

npm install knayi-myscript
yarn add knayi-myscript
pnpm add knayi-myscript
bun add knayi-myscript

Browser script, global name knayi:

<script src="https://unpkg.com/knayi-myscript@2.10.0/dist/knayi-myscript.min.js"></script>

Runtime

Node.js 16 or newer, checked on Node 16, 18, 20, and 26. Building and testing the package needs Node 22 or newer. Node 24 is the version in .nvmrc.

const knayi = require('knayi-myscript')
import knayi from 'knayi-myscript'
import knayi from 'knayi-myscript'

TypeScript types are index.d.ts. Named imports such as import { fontConvert } from 'knayi-myscript' work in Node and in bundlers, next to the default import. The default import compiles with or without esModuleInterop. The option types (DetectorOptions, GlobalOptions, TruncateOptions) are exported.

In Node, require and import both load main.js and share setGlobalOptions. A bundler that follows the module field loads dist/knayi-myscript.es.js instead. That file is a second copy. If one part of an app uses main.js and another uses dist/knayi-myscript.es.js, silent mode and detector settings do not cross between them.

The script build sets the global knayi, both in a <script> tag and when a bundler loads it with import 'knayi-myscript/dist/knayi-myscript.min.js'.

The dist/ builds are ES2015. They run in Chrome 49, Edge 14, Firefox 34, Safari 10 (iOS 10), Samsung Internet 5 and Opera 36, or newer. Internet Explorer needs knayi 2.8.3.

These paths load without an exports map:

  • knayi-myscript
  • knayi-myscript/library/converter
  • knayi-myscript/dist/knayi-myscript.min.js
  • knayi-myscript/dist/knayi-myscript.es.js

Font names

unicode, uni, zawgyi, zaw, and win. uni is Unicode. zaw is Zawgyi. win is the Win Innwa family of legacy fonts, which fontConvert converts to Unicode. Any other string is an unknown font.

Missing content

null, undefined, '', 0, false, and NaN are missing content, as in 2.8.3.

Function Missing content
fontDetect The fallback, or 'en' when the fallback is omitted. Warns unless silent.
fontConvert, syllBreak, spellingFix, normalize ''. Warns unless silent.
truncate ''. Warns unless silent. An empty string '' returns the omission instead.

Text with no Myanmar letters (U+1000–U+109F) is returned unchanged by detect, convert, break, spelling fix, and normalize. fontDetect returns the fallback or 'en'. truncate still appends the omission.

Other values, such as numbers and objects, are returned unchanged the same way, and no function throws on them. truncate turns them into strings first, like lodash.truncate. String objects work like the strings they hold.

setGlobalOptions({ silent_mode: true }) hides those warnings. The option applies to the copy of the library that received the call.

fontDetect(content, fallbackFontType?, options?)

Returns 'unicode', 'zawgyi', or the fallback / 'en'.

When the rule scores tie, including a single consonant such as က, the result is the fallback, or 'zawgyi' if the fallback is omitted.

knayi.fontDetect('မဂၤလာပါ') // 'zawgyi'
knayi.fontDetect('မင်္ဂလာပါ') // 'unicode'
knayi.fontDetect('ကျ') // 'unicode'
knayi.fontDetect('က') // 'zawgyi'
knayi.fontDetect('က', 'unicode') // 'unicode'
knayi.fontDetect(null) // 'en'

options.adapter chooses the detector for that call. 'rules' is the built-in scorer and the default. 'myanmartools' uses the myanmar-tools package. Install it only for that adapter, and use 1.1.x: myanmar-tools 1.2.0 on npm was published without its built files and cannot be loaded.

npm install myanmar-tools@1.1.3
knayi.fontDetect('မဂၤလာပါ', null, { adapter: 'myanmartools' })
knayi.fontDetect('မင်္ဂလာပါ', null, {
  use_myanmartools: true,
  myanmartools_zg_threshold: [0.05, 0.95]
})

use_myanmartools: true selects the same adapter. A probability below the first threshold returns 'unicode'. A probability above the second returns 'zawgyi'. A probability between them returns the fallback. The default pair is [0.05, 0.95]. If the package is not installed or cannot be loaded, the call uses the rule scorer and warns once. The warning says which of the two happened.

setGlobalOptions({ detector: { use_myanmartools: true } }) changes the default. An explicit adapter on a later call wins. A later call that only sets use_myanmartools keeps a previously stored threshold.

The rule scorer does not count a consonant, U+1039, consonant sequence such as က္က as Unicode. In Zawgyi, U+1039 is the visible asat, so ပ္က is a common Zawgyi sequence. A lone stack is a tie and returns the fallback. In longer Unicode text such as ရန်ကုန်တက္ကသိုလ်, the other signs decide.

fontConvert(content, targetFontType, originalFontType?)

Returns a string. targetFontType is required. When originalFontType is omitted, fontDetect chooses it.

The text is trimmed first. Zero-width spaces (U+200B) and non-joiners (U+200C) are kept, because they mark word breaks. When the two fonts are the same, the trimmed text is returned.

knayi.fontConvert('မဂၤလာပါ', 'unicode', 'zawgyi') // 'မင်္ဂလာပါ'
knayi.fontConvert('မဂၤလာပါ', 'unicode') // 'မင်္ဂလာပါ'
knayi.fontConvert('မြန်မာ', 'zawgyi', 'unicode') // 'ျမန္မာ'
knayi.fontConvert('ကျ', 'unicode') // 'ကျ'
knayi.fontConvert(' ကာာ ', 'unicode', 'unicode') // 'ကာာ'
knayi.fontConvert('မဂၤလာပါ', 'uni', 'zaw') // 'မင်္ဂလာပါ'
knayi.fontConvert(null, 'unicode') // ''
knayi.fontConvert('က') // 'က'  (no target font; warns)

fontConvert.debugging(content, targetFontType, originalFontType) returns { to, from, matched_patterns, steps }. steps is an array of strings. The last step equals fontConvert for the same arguments. From Unicode, matched_patterns holds the source of each rule pattern that matched. From Zawgyi or Win, it names each stage that changed the text: sequences, glyphs, syllables, zero as wa, look-alikes, typos, NFC.

Zawgyi to Unicode

Zawgyi stores text in the order the glyphs are drawn: ေ and medial ra before the consonant, kinzi and stacked consonants after it, and the marks in any order. knayi reads each Zawgyi glyph as Unicode characters and writes every syllable in Unicode storage order (UTN #11). Win fonts use the same rules.

knayi.fontConvert('ေယာက္်ား', 'unicode', 'zawgyi') // 'ယောက်ျား'
knayi.fontConvert('ေစ်း', 'unicode', 'zawgyi') // 'ဈေး'
knayi.fontConvert('ႏို္င္ငံ', 'unicode', 'zawgyi') // 'နိုင်ငံ'
knayi.fontConvert('ၿမိဳ ့', 'unicode', 'zawgyi') // 'မြို့'
  • Marks: a mark typed twice counts once.
  • Asat on a consonant: stored right after the consonant, before the medials and vowels: ယောက်ျား, ကျွန်ုပ်, ခ်ျ.
  • Asat stored last:
    • after ာ, as in ကျော်, even when typed before the ာ of a word with no medial (ကော်ဖီ);
    • with a dot below (ကြောင့်);
    • after medial ha (ရှ်).
  • Asat dropped: typed with ိ or ီ, or on a stacked consonant, an asat is a slip (နိုင်ငံ, ကုလသမဂ္ဂ).
  • Letters Zawgyi draws alike:
    • စ with medial ya is ဈ (ဈေး);
    • ဥ with a stacked consonant, asat or ာ is ဉ (ပဉ္စ, ဉာဏ်);
    • ၄ before င်း is ၎ (၎င်း);
    • ၇ with a vowel sign or medial is ရ (ရေး).
  • Zero: ၀ is also ဝ. A zero stays a digit next to a digit or an arithmetic sign, or across a decimal point from a digit (၁၀၀, ၅.၀).
  • Typing fixes, as in normalize: ဝ or ရ typed in a number is a digit (၂ဝ၁၉ is ၂၀၁၉). ၇ starting a closed syllable is ရ (ဆိုရင်). ိ with ီ is ီ (ဦး), and ု with ူ is ူ.
  • Spaces: a space typed before a mark only moved the mark, so it is dropped: ၿမိဳ ့ is မြို့ and တစ္ခ ု is တစ်ခု. A line break stays.
  • Zero-width characters: a zero-width space or non-joiner typed inside a syllable moves to the end of the syllable.
  • NFC: the result is NFC.

Converting from Unicode collapses a mark typed twice in a row, as spellingFix does, then applies knayi's pattern rules.

Win fonts

Win Innwa, Win Researcher, Win Kalaw and the other Win fonts by WinMyanmar Systems (1992–2005) draw Burmese glyphs on the keys that type them. Win text is ASCII and Latin-1: jrefrm shows as မြန်မာ in a Win font. Name the source font, because fontDetect never returns win.

knayi.fontConvert('jrefrm', 'unicode', 'win') // 'မြန်မာ'
knayi.fontConvert('ajumifh', 'unicode', 'win') // 'ကြောင့်'
knayi.fontConvert('ZvGefaps;', 'unicode', 'win') // 'ဇလွန်ဈေး'
knayi.fontConvert('jrefrm', 'unicode') // 'jrefrm'  (no source font: plain ASCII)

knayi converts Win to Unicode only. Any other target returns the text unchanged, with an error unless silent.

  • Win text is stored in drawing order, like Zawgyi, and knayi converts it with the same rules (see Zawgyi to Unicode). So ajumifh (asat before the dot below) becomes ကြောင့် with the dot below first, a,musfm; is ယောက်ျား and usGefkyf is ကျွန်ုပ်.
  • 0 is both ဝ and ၀ in Win, and 7 can be ရ, as in Zawgyi. ps, Mo, aMomf and OD become ဈ, ဩ, ဪ and ဦ.
  • Text read as ISO-8859-1 instead of Windows-1252 converts the same way.
  • Fractions become text such as ၁/၂. Dingbats become the Unicode symbols they show. The vendor logo at byte 0xB0 is dropped.
  • English typed in another font run is ASCII too. Once the font names are gone, convert only the Win text.
  • Wwin_Burmese and other ASCII fonts use different mappings and are not supported.

syllBreak(content, fontType?, breakPoint?)

Returns one string. The default break character is U+200B. This is the current public break, not a split into မ|င်္ဂ|လာ|ပါ.

knayi.syllBreak('မင်္ဂလာပါ', null, '$$') // 'မင်္ဂလာ$$ပါ'
knayi.syllBreak('မင်္ဂလာပါ') // 'မင်္ဂလာ' + '\u200b' + 'ပါ'
knayi.syllBreak('မြန်မာ', 'unicode', '|') // 'မြန်|မာ'
knayi.syllBreak('ထို့ကြောင့်', 'unicode', '|') // 'ထို့|ကြောင့်'
knayi.syllBreak('က္က', 'unicode', '|') // 'က္က'
knayi.syllBreak('က္က', 'zawgyi', '|') // 'က္|က'
knayi.syllBreak('က္က', 'uni', '|') // 'က္က'
knayi.syllBreak('ကက', 'unicode', '|') // 'ကက'
knayi.syllBreak('ၾကပါ', 'zawgyi', '|') // 'ၾက|ပါ'

When fontType is omitted, detection runs first. Unknown font names throw.

Zawgyi types ေ and the medial ra before the consonant. A consonant typed after them ends its syllable, as ကြ does in Unicode.

spellingFix(content, fontType?)

Collapses a mark repeated two or more times into one mark. It does not reorder marks.

knayi.spellingFix('မင်္ဂလာာပါါ', 'unicode') // 'မင်္ဂလာပါ'
knayi.spellingFix('ကိီ', 'unicode') // 'ကိီ'
knayi.spellingFix('\u1033\u1033', 'zawgyi') // '\u1033'
knayi.spellingFix('\u1033\u1033', 'zaw') // '\u1033'

normalize(content)

Unicode only, written for Burmese. Puts every syllable in Unicode storage order (UTN #11) with the rules of Zawgyi to Unicode, makes a few typing fixes, and returns NFC.

  • What stays the same: text that is already right, text normalized a second time, and the output of fontConvert all come back unchanged.
  • What it keeps: surrounding spaces, zero-width spaces and joiners.

It is not the same operation as spellingFix.

knayi.normalize('မိြုင်မိြုင်\nဆိုင်ဆုိင်') // 'မြိုင်မြိုင်\nဆိုင်ဆိုင်'
knayi.normalize(' မိြုင် ') // ' မြိုင် '
knayi.normalize('ယောကျ်ား') // 'ယောက်ျား'
knayi.normalize('လည်းေကာင်း') // 'လည်းကောင်း'
knayi.normalize('၂ဝ၁၉') // '၂၀၁၉'
knayi.normalize('ကိီ') // 'ကီ'
knayi.normalize('ဝ') // 'ဝ'
  • Order: marks typed in any order are sorted, and a mark typed twice counts once. Asat goes where UTN #11 puts it (ကျွန်ုပ်, ခ်ျ, ရှ်), and the dot below comes before asat, as NFC requires.
  • Zawgyi typing habits:
    • ေ or medial ra typed before its consonant moves after it: လည်းေကာင်း is လည်းကောင်း.
    • A space typed before a mark is dropped (သုံ း is သုံး). A line break stays.
  • Look-alikes: only clear cases change.
    • ဝ and ရ inside a number are digits: ၄ဝဝ is ၄၀၀.
    • ၀ and ၇ that carry a vowel sign or start a closed syllable are letters, as is ၀ inside a word: ဘ၀ is ဘဝ, ဆို၇င် is ဆိုရင်.
    • Words such as လုံးဝ, ဘဝ and ထာဝရ, and numbers such as ၁၉၇၇, stay as they are.
  • Spelling:
    • စ with medial ya is ဈ.
    • ဥ with asat, aa or a stacked consonant is ဉ (ညဉ့်, ဉာဏ်), except right after a vowel sign, where Pa'o writes ဥ်း.
    • ၄င်း is ၎င်း, ိ with ီ is ီ, ု with ူ is ူ, and ဩော် is ဪ.
  • Other languages: Mon, Karen, Pa'o and Shan letters stay as they are, and so do spellings that differ from Burmese (တုဲ, ခရံာ်).

truncate(content, options?)

Cuts on the current syllable breaks, then on spaces inside a syllable that does not fit. Defaults are length: 30 and omission: '...'. The omission is appended even when the text is shorter than length. options.fontType accepts the same font names. When omitted, detection runs.

knayi.truncate('အာယုဝဍ်ဎနဆေးညွှန်းစာကို ဇလွန်ဈေးဘေးဗာဒံပင်ထက် အဓိဋ္ဌာန်လျက် ဂဃနဏဖတ်ခဲ့သည်။', { length: 30, omission: '...' })
// 'အာယုဝဍ်ဎနဆေးညွှန်းစာကို ဈေး...'
knayi.truncate('က') // 'က...'
knayi.truncate('') // '...'
knayi.truncate(null) // ''

Build

npm test builds the browser and ESM files, runs the tests, and type-checks typecheck/. npm run test:bun runs the Bun checks. npm run test:pack packs the tarball, installs it with Bun, and converts the Zawgyi greeting through require and import. npm run build writes:

  • dist/knayi-myscript.mjs
  • dist/knayi-myscript.es.js (same bytes as the .mjs file)
  • dist/knayi-myscript.js
  • dist/knayi-myscript.min.js

npm run eval measures conversion and detection on public Zawgyi and Unicode data, next to a published knayi release, myanmar-tools, and Rabbit. npm run bench measures speed on real text and long input. Both download their data on first use. See scripts/eval/README.md. The latest results are published at https://greenlikeorange.github.io/knayi-myscript/benchmark.html.

Releases

Packages

Used by

Contributors

Languages