About this romanisation
« Back to the romanisation page
1. What the tool does

Japanese text you enter is analysed into short units by MeCab and UniDic, and romanised with word division on the basis of the kana form of each unit. The grammatical information that word division requires — telling particles and auxiliaries apart, identifying prefixes and suffixes, recognising proper nouns — is taken from the result of the morphological analysis.

The work is divided as follows.

2. The two spelling standards

The two agree for the most part. They differ in the following respects.

3. How the two standards differ
ALA-LC 2022 Cabinet Notification 2025
Geminate before the chi row add t
密着 mitchaku
double the c (note 2)
日直 nicchoku, 抹茶 maccha
Long vowels macron only (§5.1)
東北 Tōhoku
macron, or repeated vowel letters (note 3)
Tōhoku / Touhoku
Adjacent vowels that are not long no mark (§2.5.5)
小唄 kouta, 本居 motoori
an apostrophe may mark the break (note 4)
ko'uta, moto'ori
Sounds used only in loanwords
(ファ, ティ, ヴ and the like)
given in chart B outside its scope (preamble 4)
→ this tool falls back on chart B
Word division detailed rules (§6) almost no provision
→ this tool always uses ALA-LC §6

Because the Cabinet Notification places sounds used only in loanwords outside its scope (preamble 4), those are converted using ALA-LC chart B even when the Notification is selected. And because it says almost nothing about word division — only that a hyphen may be used (note 5) — word division always follows ALA-LC §6.

4. Where the two agree
5. How the conversion works
6. How word division is decided

The short-unit boundaries given by UniDic are the starting point. Joining, separation and hyphenation are then decided by the rules of ALA-LC §6.

Where the rules turn on on'yomi versus kun'yomi (§6.2.3 against §6.3.1, §6.4.1.1.1 against §6.4.1.1.2), the decision is made from the word origin recorded by UniDic: 漢 Sino-Japanese, 和 native, 外 foreign.

7. Limitations

For these reasons, always check the output. Selecting the per-unit breakdown on the input page shows which rule caused each word to be joined or separated.

« Back to the romanisation page