1. What the tool does
Japanese text you enter is analysed into short units by MeCab and UniDic, and romanised
with word division on the basis of the kana form of each unit. The
grammatical information that word division requires — telling particles and auxiliaries
apart, identifying prefixes and suffixes, recognising proper nouns — is taken from the
result of the morphological analysis.
The work is divided as follows.
- Readings follow the result of the analysis. The romaniser does not,
as a rule, override them.
- Spelling — the correspondence between kana and roman letters — can be
chosen from two standards on the input page.
- Word division always follows ALA-LC §6, whichever spelling you
choose.
2. The two spelling standards
The two agree for the most part. They differ in the following respects.
3. How the two standards differ
|
ALA-LC
2022 |
Cabinet Notification 2025 |
| Geminate before the chi row |
add t 密着 mitchaku |
double the c (note 2) 日直 nicchoku, 抹茶 maccha |
| Long vowels |
macron only (§5.1) 東北 Tōhoku |
macron, or repeated vowel letters (note 3) Tōhoku / Touhoku |
| Adjacent vowels that are not long |
no mark (§2.5.5) 小唄 kouta, 本居 motoori |
an apostrophe may mark the break (note 4) ko'uta, moto'ori |
Sounds used only in loanwords (ファ, ティ, ヴ and the like) |
given in chart B |
outside its scope (preamble 4) → this tool falls back on chart B |
| Word division |
detailed rules (§6) |
almost no provision → this tool always uses ALA-LC §6 |
Because the Cabinet Notification places sounds used only in loanwords outside its scope
(preamble 4), those are converted using ALA-LC chart B even when the Notification is
selected. And because it says almost nothing about word division — only that a hyphen
may be used (note 5) — word division always follows ALA-LC §6.
4. Where the two agree
- Syllabic n is always written n. 新橋 Shinbashi, 天ぷら tenpura, 欄間 ranma — never m.
- The particles は, へ and を are written wa, e and o.
- シ shi, チ chi, ツ tsu, フ fu, ジ ji. ヲ is o. ヂ and ヅ are ji and zu.
- A long i is written ii. 新潟 Niigata, 兄さん niisan.
- ei is written ei. 平成 Heisei, 時計台 tokeidai, 経済的 keizaiteki, 英語 Eigo.
- An apostrophe follows n before a vowel or y. 翻訳 hon'yaku, 単位 tan'i.
- Proper nouns and the first word of a sentence are capitalised.
5. How the conversion works
-
The spelling follows the kana form given by UniDic. The
pronunciation form is used only to decide whether a sequence of
vowels has merged into a long vowel. This is why 労働 comes out as rōdō while 子牛
comes out as koushi (ko'ushi under the Cabinet Notification). Transcribing the
pronunciation form directly would make 子牛 kōshi as well.
-
Even though the pronunciation form of 政治 is セージ, it is romanised seiji. Both
standards write ei as ei — compare ALA-LC's keizaiteki and Eigo with the Cabinet
Notification's Heisei and tokeidai.
-
Where the kana form and the pronunciation form diverge by more than the presence of a
long vowel, the pronunciation form is taken as the basis, since it represents the
modern reading (てふてふ chōchō, あふ坂 Ōsaka). ALA-LC §2.3.2 prescribes modern
readings for historical kana usage.
-
A geminate is written as a doubled consonant. Where a short unit ends in a geminate
(帰っ + て), the consonant of the following unit is supplied when the two are written
together: kaette.
6. How word division is decided
The short-unit boundaries given by UniDic are the starting point. Joining, separation and
hyphenation are then decided by the rules of ALA-LC §6.
- Particles are written separately (§6.8.4): 幸福への道 kōfuku e no michi.
- Auxiliaries, prefixes and suffixes are joined to the adjacent word (§6.5, §6.6,
§6.8.1.1): 支配する shihaisuru, 経済的 keizaiteki, 新幹線 shinkansen.
- An auxiliary verb following the te-form, and auxiliaries of potential or respect, are
written separately (§6.8.1.2, §6.8.1.3): 生きていた ikite ita, 我慢出来ない gaman
dekinai.
- Jurisdictions, stations and harbours, titles following personal names, and single
characters suffixed to proper nouns are hyphenated (§6.9.2.3.2, §6.9.5, §6.10.1,
§6.10.6): 東京都 Tōkyō-to, 井上さん Inoue-san, 日本的 Nihon-teki.
- A single-character element before or after a word is joined (§6.2.3, §6.6.1,
§6.4.1.1.1): 日本語 Nihongo, 核戦争 kakusensō, 都道府県 todōfuken. But 編, 抄, 考 and 展
are written separately (§6.6.2.1): 君が代考 Kimigayo kō.
- A modifier written in kana or read with kun'yomi is separated (§6.3.1, §6.4.1.1.2):
米騒動 kome sōdō, 水資源 mizu shigen, 春夏秋冬 haru natsu aki fuyu.
Where the rules turn on on'yomi versus kun'yomi (§6.2.3 against §6.3.1, §6.4.1.1.1 against
§6.4.1.1.2), the decision is made from the word origin recorded by UniDic:
漢 Sino-Japanese, 和 native, 外 foreign.
7. Limitations
-
Choice of reading. Where one orthography admits several readings, the
tool follows whichever the analyser chose. 描ける (kakeru or egakeru), 日本 (Nihon or
Nippon) and 兄妹 (ani imōto or kyōdai) all depend on context. A different reading can
change the word division as well.
-
Where a word begins and ends. A word in the ALA-LC sense corresponds
roughly to a dictionary headword, which is larger than a UniDic short unit. Compounds
conventionally written as one word are sometimes split.
-
Proper nouns. Readings of personal and place names depend on the
analysis dictionary. Conventional romanised forms such as Tokyo or judo, widely used
internationally, are not reproduced.
-
Numerals. Whether kanji numerals should become Arabic numerals depends
on context; §7.2.1 gives 二十四の瞳 Nijūshi no hitomi as a case where they should not.
This can be switched on the input page.
For these reasons, always check the output. Selecting the per-unit
breakdown on the input page shows which rule caused each word to be joined or separated.