README

Lang::JA::Kana - Japanese Kana Conversion Utilities

Languages: English â€ĸ æ—ĨæœŦčĒž

Documentation:

Overview

Lang::JA::Kana is a Raku module for converting between different Japanese kana scripts (Hiragana and Katakana) and their various forms. It provides support for modern kana, historical variants, half-width characters, and specialized Unicode symbols.

Features

  • Bidirectional Script Conversion: Seamless conversion between Hiragana and Katakana

  • Half-width Support: Handling of half-width katakana (īžŠīžīŊļīŊ¸) conversion

  • Historical Kana: Support for Hentaigana (変äŊ“äģŽå) and obsolete characters

  • Modern Extensions: Foreign sound adaptations (ãƒ•ã‚Ą, ãƒ†ã‚Ŗ, ã‚Ļã‚Ŗ, etc.)

  • Specialized Symbols: Circled and squared katakana processing

  • Sound Mark Analysis: Diacritical mark separation and analysis

  • Cross-script Integration: Built-in integration with Romaji, Cyrillic, and Hangul converters

Installation

use Lang::JA::Kana;

Basic Usage

Hiragana ↔ Katakana Conversion

use Lang::JA::Kana;

# Basic conversions
say to-katakana("こんãĢãĄã¯");     # → ã‚ŗãƒŗãƒ‹ãƒãƒ
say to-hiragana("ã‚ŗãƒŗãƒ‹ãƒãƒ");     # → こんãĢãĄã¯

# With modern extensions
say to-katakana("ãĩぁãŋりãƒŧ");     # → ãƒ•ã‚ĄãƒŸãƒĒãƒŧ
say to-hiragana("ãƒ•ã‚ĄãƒŸãƒĒãƒŧ");     # → ãĩぁãŋりãƒŧ

# Combination sounds (æ‹—éŸŗ)
say to-katakana("きゃりãƒŧãąãŋã‚…ãąãŋゅ");  # → ã‚­ãƒŖãƒĒãƒŧパミãƒĨパミãƒĨ
say to-hiragana("ã‚­ãƒŖãƒĒãƒŧパミãƒĨパミãƒĨ");  # → きゃりãƒŧãąãŋã‚…ãąãŋゅ

# Mixed text (non-kana characters pass through unchanged)
say to-katakana("Hello こんãĢãĄã¯ World");  # → Hello ã‚ŗãƒŗãƒ‹ãƒãƒ World
say to-hiragana("Hello ã‚ŗãƒŗãƒ‹ãƒãƒ World");  # → Hello こんãĢãĄã¯ World

Half-width Katakana Conversion

# Half-width to full-width conversion
say to-fullwidth-katakana("īŊąīŊ˛īŊŗīŊ´īŊĩ");     # → ã‚ĸイã‚Ļエã‚Ē
say to-fullwidth-katakana("īŊļīžžīŊˇīžžīŊ¸īžž");    # → ã‚Ŧゎグ (voiced combinations)
say to-fullwidth-katakana("īžŠīžŸīž‹īžŸīžŒīžŸ");    # → パピプ (semi-voiced combinations)

# Full-width to half-width conversion
say to-halfwidth-katakana("ã‚ĸイã‚Ļエã‚Ē");   # → īŊąīŊ˛īŊŗīŊ´īŊĩ
say to-halfwidth-katakana("ã‚Ŧゎグ");      # → īŊļīžžīŊˇīžžīŊ¸īžž
say to-halfwidth-katakana("パピプ");      # → īžŠīžŸīž‹īžŸīžŒīžŸ

# Integration with other conversions
say to-hiragana("īŊļīž€īŊļīž…");               # → かたかãĒ (auto-converts half-width)

Character Support

Standard Kana

All 50-sound (äē”åéŸŗ) characters:

# Basic vowels (æ¯éŸŗ)
to-katakana("あいうえお");  # → ã‚ĸイã‚Ļエã‚Ē

# K-series (ã‚Ģ行)
to-katakana("かきくけこ");  # → ã‚Ģã‚­ã‚¯ã‚ąã‚ŗ
to-katakana("がぎぐげご");  # → ã‚Ŧã‚Žã‚°ã‚˛ã‚´

# S-series (ã‚ĩ行)
to-katakana("さしすせそ");  # → ã‚ĩã‚ˇã‚šã‚ģã‚Ŋ
to-katakana("ざじずぜぞ");  # → ã‚ļジã‚ēã‚ŧゞ

# T-series (ã‚ŋ行)
to-katakana("ãŸãĄã¤ãĻと");  # → ã‚ŋチツテト
to-katakana("だãĸãĨでお");  # → ダヂヅデド

# N-series (ãƒŠčĄŒ)
to-katakana("ãĒãĢãŦねぎ");  # → ナニヌネノ

# H-series (ãƒčĄŒ)
to-katakana("ã¯ã˛ãĩへãģ");  # → ハヒフヘホ
to-katakana("ã°ãŗãļずãŧ");  # → バビブベボ (voiced)
to-katakana("ãąã´ãˇãēãŊ");  # → パピプペポ (semi-voiced)

# M-series (ãƒžčĄŒ)
to-katakana("ぞãŋむめも");  # → ãƒžãƒŸãƒ ãƒĄãƒĸ

# Y-series (ãƒ¤čĄŒ)
to-katakana("やゆよ");     # → ヤãƒĻヨ

# R-series (ãƒŠčĄŒ)
to-katakana("らりるれろ");  # → ナãƒĒãƒĢãƒŦロ

# W-series (ãƒ¯čĄŒ) and N
to-katakana("わゐゑをん");  # → ãƒ¯ãƒ°ãƒąãƒ˛ãƒŗ

Small Kana (小文字)

# Small vowels
to-katakana("ぁぃぅぇぉ");  # → ã‚Ąã‚Ŗã‚Ĩェり

# Small Y-sounds
to-katakana("ゃゅょ");     # → ãƒŖãƒĨョ

# Small tsu (äŋƒéŸŗ)
to-katakana("ãŖ");         # → ッ

# Small wa
to-katakana("ゎ");         # → ノ

Combination Sounds (æ‹—éŸŗ)

# Y-combinations
to-katakana("きゃきゅきょ");  # → ã‚­ãƒŖã‚­ãƒĨキョ
to-katakana("しゃしゅしょ");  # → ã‚ˇãƒŖã‚ˇãƒĨã‚ˇãƒ§
to-katakana("ãĄã‚ƒãĄã‚…ãĄã‚‡");  # → ãƒãƒŖãƒãƒĨチョ
to-katakana("ãĢゃãĢゅãĢょ");  # → ãƒ‹ãƒŖãƒ‹ãƒĨニョ
to-katakana("ã˛ã‚ƒã˛ã‚…ã˛ã‚‡");  # → ãƒ’ãƒŖãƒ’ãƒĨヒョ
to-katakana("ãŋゃãŋゅãŋょ");  # → ãƒŸãƒŖãƒŸãƒĨミョ
to-katakana("りゃりゅりょ");  # → ãƒĒãƒŖãƒĒãƒĨãƒĒョ

# Voiced Y-combinations
to-katakana("ぎゃぎゅぎょ");  # → ã‚ŽãƒŖã‚ŽãƒĨゎョ
to-katakana("じゃじゅじょ");  # → ã‚¸ãƒŖã‚¸ãƒĨジョ
to-katakana("ãŗã‚ƒãŗã‚…ãŗã‚‡");  # → ãƒ“ãƒŖãƒ“ãƒĨビョ
to-katakana("ぴゃぴゅぴょ");  # → ãƒ”ãƒŖãƒ”ãƒĨピョ

Modern Extensions for Foreign Sounds

# F-sounds
to-katakana("ãĩぁãĩぃãĩぇãĩぉ");  # → ãƒ•ã‚Ąãƒ•ã‚Ŗãƒ•ã‚§ãƒ•ã‚Š

# T/D-sounds
to-katakana("ãĻぃでぃ");         # → ãƒ†ã‚Ŗãƒ‡ã‚Ŗ
to-katakana("とぅおぅ");         # → トã‚Ĩドã‚Ĩ

# W-sounds
to-katakana("うぃうぇうぉ");     # → ã‚Ļã‚Ŗã‚Ļェã‚Ļり

# V-sounds
to-katakana("ゔぁゔぃゔぇゔぉ"); # → ãƒ´ã‚Ąãƒ´ã‚Ŗãƒ´ã‚§ãƒ´ã‚Š

# Kw/Gw-sounds
to-katakana("くぁくぃくぇくぉ"); # → ã‚¯ã‚Ąã‚¯ã‚Ŗã‚¯ã‚§ã‚¯ã‚Š
to-katakana("ぐぁぐぃぐぇぐぉ"); # → ã‚°ã‚Ąã‚°ã‚Ŗã‚°ã‚§ã‚°ã‚Š

# Ts-sounds
to-katakana("つぁつぃつぇつぉ"); # → ãƒ„ã‚Ąãƒ„ã‚Ŗãƒ„ã‚§ãƒ„ã‚Š

# Other combinations
to-katakana("ãĄã‡ã˜ã‡ã—ã‡ã„ã‡"); # → ãƒã‚§ã‚¸ã‚§ã‚ˇã‚§ã‚¤ã‚§

Historical and Obsolete Kana

# Historical Wi/We/Wo
to-katakana("ゐゑを");  # → ãƒ°ãƒąãƒ˛

# VU sound
to-katakana("ゔ");      # → ヴ

# Digraph Yori
to-katakana("ゟ");      # → ãƒŋ

Hentaigana (変äŊ“äģŽå) Support

Hentaigana are historical variant forms of kana characters used before standardization. The module provides support for these Unicode characters.

Hentaigana to Hiragana Conversion

# Basic conversion
say hentaigana-to-hiragana("𛀁𛀂𛀃");  # → あいう

# Multiple readings (some Hentaigana have ambiguous readings)
say hentaigana-to-hiragana("𛀒");      # → しãƒģせ (can be "shi" or "se")
say hentaigana-to-hiragana("𛁆");      # → ゐãƒģい (can be "wi" or "i")

# Complex examples
say hentaigana-to-hiragana("𛀆𛀈𛀊");  # → かきく

Hiragana to Hentaigana Conversion

# Single variant
say hiragana-to-hentaigana("る");      # → 𛁂

# Multiple variants (shows all possibilities)
say hiragana-to-hentaigana("あ");      # → 𛀁ãƒģ𛄀ãƒģ𛄁
say hiragana-to-hentaigana("し");      # → 𛀒ãƒģ𛀖ãƒģ𛂡

# With voiced marks
say hiragana-to-hentaigana("が");      # → (variants)゛

Hentaigana Character Origins

Many Hentaigana derive from specific Chinese characters (kanji):

  • 𛀁 (A): From 厉 (an)

  • 𛀂 (I): From äģĨ (i)

  • 𛀃 (U): From 厇 (u)

  • 𛀆 (KA): From 加 (ka)

  • 𛀈 (KI): From åšž (ki)

  • 𛀐 (SA): From åˇĻ (sa)

  • 𛀒 (SHI/SE): From 之 (shi/se) - ambiguous reading

  • 𛀚 (TA): From å¤Ē (ta)

  • 𛀜 (CHI): From įŸĨ (chi)

Sound Mark Processing

The module provides utilities for analyzing and manipulating diacritical marks (æŋį‚šãƒģ半æŋį‚š).

Sound Mark Splitting

# Split voiced characters into base + mark
my @parts = split-sound-marks("が");
say @parts[0];  # → か (base character)
say @parts[1];  # → ゛ (voiced mark)

# Split semi-voiced characters
@parts = split-sound-marks("ãą");
say @parts[0];  # → は (base character)
say @parts[1];  # → ゜ (semi-voiced mark)

# Regular characters return as-is
@parts = split-sound-marks("あ");
say @parts[0];  # → あ (no splitting)

Practical Applications

# Analyze character composition
sub analyze-kana($char) {
    my @parts = split-sound-marks($char);
    if @parts.elems == 2 {
        say "$char = {@parts[0]} + {@parts[1]}";
    } else {
        say "$char = base character";
    }
}

analyze-kana("が");  # → が = か + ゛
analyze-kana("ãą");  # → ãą = は + ゜
analyze-kana("あ");  # → あ = base character

Specialized Unicode Symbols

Circled Katakana

# Convert circled katakana to components
say decircle-katakana("㋐㋑㋒");  # → ã‚ĸイã‚Ļ
say decircle-katakana("㋕㋖㋗");  # → ã‚Ģキク

# Convert components to circled katakana
say encircle-katakana("ã‚ĸイã‚Ļ");  # → ㋐㋑㋒
say encircle-katakana("ã‚Ģキク");  # → ㋕㋖㋗

Squared Katakana (Units and Abbreviations)

# Convert squared katakana to full forms
say desquare-katakana("㌔");     # → キロ (kilo)
say desquare-katakana("㌧");     # → ãƒˆãƒŗ (ton)
say desquare-katakana("㍍");     # → ãƒĄãƒŧトãƒĢ (meter)
say desquare-katakana("㍑");     # → ãƒĒットãƒĢ (liter)

# Convert full forms to squared katakana
say ensquare-katakana("キロ");     # → ㌔
say ensquare-katakana("ãƒĄãƒŧトãƒĢ");  # → ㍍

# Complex examples
say desquare-katakana("㌔㍍");     # → ã‚­ãƒ­ãƒĄãƒŧトãƒĢ
say desquare-katakana("㍉㍍");     # → ミãƒĒãƒĄãƒŧトãƒĢ

Common Squared Katakana Units:

  • ㌔ (キロ) - kilo

  • ㌧ (ãƒˆãƒŗ) - ton

  • ㍍ (ãƒĄãƒŧトãƒĢ) - meter

  • ㍑ (ãƒĒットãƒĢ) - liter

  • ㍉ (ミãƒĒ) - milli

  • ãŒĸ (ã‚ģãƒŗãƒ) - centi

  • ãŒĻ (ドãƒĢ) - dollar

  • ãŒĢ (パãƒŧã‚ģãƒŗãƒˆ) - percent

  • ㍗ (ワット) - watt

Cross-script Conversion Integration

The module seamlessly integrates with other script converters, automatically handling half-width conversion.

Romaji Conversion

# Automatic half-width handling
say kana-to-romaji("īŊēīžīž†īžīžŠ");                    # → konnichiha
say kana-to-romaji("こんãĢãĄã¯");                 # → konnichiha

# Multiple romanization systems
say kana-to-romaji("しんãļん", :system<hepburn>);  # → shinbun
say kana-to-romaji("しんãļん", :system<kunrei>);   # → sinbun
say kana-to-romaji("しんãļん", :system<nihon>);    # → sinbun

# Sokuon (ãŖ) handling
say kana-to-romaji("ãŒãŖã“ã†");                   # → gakkou
say kana-to-romaji("ãĄã‚‡ãŖã¨");                   # → chotto

Cyrillic Conversion

# Polivanov system (default)
say kana-to-kuriru-moji("こんãĢãĄã¯");              # → ĐēĐžĐŊĐŊĐ¸Ņ‡Đ¸Ņ…Đ°
say kana-to-kuriru-moji("ã˛ã‚‰ãŒãĒ");               # → Ņ…Đ¸Ņ€Đ°ĐŗĐ°ĐŊа

# Phonetic system
say kana-to-kuriru-moji("しんãļん", :system<phonetic>);  # → ŅˆĐ¸ĐŊĐąŅƒĐŊ
say kana-to-kuriru-moji("ãĄã‚ƒãĄã‚…ãĄã‚‡", :system<phonetic>); # → Ņ‡Đ°Ņ‡ŅƒŅ‡Đž

# Slavic language variants
say kana-to-kuriru-moji("さくら", :system<ukrainian>);    # → ŅĐ°ĐēŅƒŅ€Đ°
say kana-to-kuriru-moji("さくら", :system<serbian>);      # → ŅĐ°ĐēŅƒŅ€Đ°

Hangul Conversion

# Standard system (default)
say kana-to-hangul("こんãĢãĄã¯");                 # → ęŗ¤ë‹ˆėš˜í•˜
say kana-to-hangul("ã˛ã‚‰ãŒãĒ");                  # → 히ëŧ가나

# Academic system (with consonant doubling)
say kana-to-hangul("ãŒãŖã“ã†", :system<academic>);  # → 깍ėŊ”ėš°
say kana-to-hangul("ã°ãŖã°", :system<academic>);    # → ëšąëš 

# Phonetic system (preserves Japanese pronunciation)
say kana-to-hangul("ãĄã‚ƒãĄã‚…ãĄã‚‡", :system<phonetic>); # → ėš˜ė•ŧėš˜ėœ ėš˜ėš”

# Popular system (K-pop/media usage)
say kana-to-hangul("ãĄã‚…ã†", :system<popular>);     # → ėļ”ėš°

Advanced Features

Mixed Text Processing

All functions handle mixed text gracefully, processing only kana characters:

say to-katakana("Hello こんãĢãĄã¯ 123");           # → Hello ã‚ŗãƒŗãƒ‹ãƒãƒ 123
say to-hiragana("Hello ã‚ŗãƒŗãƒ‹ãƒãƒ 123");           # → Hello こんãĢãĄã¯ 123
say to-fullwidth-katakana("Hello īŊąīŊ˛īŊŗ 123");      # → Hello ã‚ĸイã‚Ļ 123

Chained Conversions

# Complex conversion chains
my $text = "īŊēīžīž†īžīžŠ";                              # Half-width katakana
$text = to-fullwidth-katakana($text);              # → ã‚ŗãƒŗãƒ‹ãƒãƒ
$text = to-hiragana($text);                        # → こんãĢãĄã¯
say kana-to-romaji($text);                         # → konnichiha

# Historical processing
$text = "𛀆𛀈𛀊";                                  # Hentaigana
$text = hentaigana-to-hiragana($text);             # → かきく
$text = to-katakana($text);                        # → ã‚Ģキク
say encircle-katakana($text);                      # → ㋕㋖㋗

Empty String and Edge Case Handling

say to-katakana("");                 # → "" (empty string)
say to-hiragana("");                 # → "" (empty string)
say to-fullwidth-katakana("");       # → "" (empty string)
say hentaigana-to-hiragana("");      # → "" (empty string)

Character Coverage

Unicode Ranges Supported

  • Hiragana: U+3040-U+309F (ã˛ã‚‰ãŒãĒ)

  • Katakana: U+30A0-U+30FF (ã‚Ģã‚ŋã‚Ģナ)

  • Half-width Katakana: U+FF61-U+FF9F (īžŠīžīŊļīŊ¸)

  • Hentaigana: U+1B001-U+1B11E (𛀁-𛄟)

  • Circled Katakana: U+32D0-U+32FE (㋐-㋞)

  • Squared Katakana: U+3300-U+3357 (㌀-㍗)

Character Count

  • Basic Hiragana/Katakana: 46 + 25 (voiced/semi-voiced) = 71 characters

  • Small Kana: 10 characters

  • Y-combinations: 33 combinations

  • Modern Extensions: 25+ foreign sound adaptations

  • Half-width Forms: 63 characters

  • Hentaigana: 300+ historical variants

  • Circled Katakana: 47 symbols

  • Squared Katakana: 88 unit abbreviations

Performance Considerations

Optimization Features

  • Longest-First Matching: Multi-character combinations processed before single characters

  • Efficient Hash Lookups: O(1) character mapping using Raku hashes

  • Minimal Regex Usage: Direct string substitution where possible

  • Lazy Evaluation: Conversion tables computed only when needed

Best Practices

# Efficient: Single conversion call
my $result = to-katakana($large-text);

# Less efficient: Multiple small conversions
for @small-texts -> $text {
    $result ~= to-katakana($text);  # Consider batching
}

# Efficient: Reuse conversion results
my $katakana = to-katakana($text);
my $romaji = kana-to-romaji($katakana);  # Uses already-converted katakana

Error Handling and Edge Cases

Robust Input Processing

# Invalid or unknown characters are preserved
say to-katakana("こんãĢãĄã¯đŸŽŒ");     # → ã‚ŗãƒŗãƒ‹ãƒãƒđŸŽŒ
say to-hiragana("ã‚Ģã‚ŋã‚ĢãƒŠđŸ—ž");       # → かたかãĒ🗾

# Mixed scripts handled appropriately
say to-katakana("ã˛ã‚‰ãŒãĒã‚Ģã‚ŋã‚Ģナ");  # → ヒナã‚Ŧナã‚Ģã‚ŋã‚Ģナ
say to-hiragana("ã‚Ģã‚ŋã‚ĢãƒŠã˛ã‚‰ãŒãĒ");  # → かたかãĒã˛ã‚‰ãŒãĒ

# Partial conversions work correctly
say to-fullwidth-katakana("Normal īŊąīŊ˛īŊŗ text");  # → Normal ã‚ĸイã‚Ļ text

Ambiguous Character Handling

# Hentaigana with multiple readings
say hentaigana-to-hiragana("𛀒");  # → しãƒģせ (shows all possibilities)

# Historical characters preserved if no modern equivalent
say to-katakana("å¤ã„đ›€æ–‡å­—");      # → å¤ã‚¤đ›€æ–‡å­— (𛀁 processed separately)

Integration Examples

Text Processing Pipeline

sub normalize-japanese-text($text) {
    # Step 1: Convert half-width to full-width
    my $normalized = to-fullwidth-katakana($text);
    
    # Step 2: Standardize to hiragana for processing
    $normalized = to-hiragana($normalized);
    
    # Step 3: Convert historical kana
    $normalized = hentaigana-to-hiragana($normalized);
    
    # Step 4: Expand abbreviated forms
    $normalized = desquare-katakana($normalized);
    $normalized = decircle-katakana($normalized);
    
    return $normalized;
}

# Example usage
my $text = "𛀁īŊ˛īŊŗãŒ”ã‹–";
say normalize-japanese-text($text);  # → あいうキロキ

Multilingual Conversion

sub convert-to-all-scripts($japanese-text) {
    # Normalize input
    my $normalized = to-fullwidth-katakana($japanese-text);
    
    return {
        hiragana => to-hiragana($normalized),
        katakana => to-katakana($normalized),
        romaji => kana-to-romaji($normalized),
        cyrillic => kana-to-kuriru-moji($normalized),
        hangul => kana-to-hangul($normalized)
    };
}

# Example usage
my %scripts = convert-to-all-scripts("īŊēīžīž†īžīžŠ");
say %scripts<hiragana>;  # → こんãĢãĄã¯
say %scripts<romaji>;    # → konnichiha
say %scripts<cyrillic>;  # → ĐēĐžĐŊĐŊĐ¸Ņ‡Đ¸Ņ…Đ°
say %scripts<hangul>;    # → ęŗ¤ë‹ˆėš˜í•˜

Limitations

  1. Kanji Processing: Does not convert Kanji characters (æŧĸ字)

  2. Context Sensitivity: Pure character-level conversion without semantic analysis

  3. Historical Accuracy: Hentaigana mappings based on Unicode standards, not historical manuscripts

  4. Regional Variants: Based on standard Japanese, not dialectal pronunciations

Use Cases

Educational Applications

  • Japanese language learning materials

  • Script conversion exercises

  • Historical text modernization

  • Unicode character reference

Text Processing

  • Document normalization

  • Search and indexing systems

  • Legacy text conversion

  • Character encoding migration

Digital Humanities

  • Historical manuscript digitization

  • Classical Japanese text processing

  • Unicode compliance testing

  • Script evolution research

Entertainment Industry

  • Game localization

  • Anime subtitle processing

  • Manga text conversion

  • Social media content adaptation

Contributing

Contributions are welcome. Please visit the project repository at:https://github.com/slavenskoj/raku-lang-ja-kana

We apologize for any errors and welcome suggestions for improvements.

References

Unicode Standards

  • Unicode Standard Annex #15: Unicode Normalization Forms

  • Unicode block specifications for Japanese scripts

  • Unicode Consortium Hentaigana guidelines

Academic Sources

  • Japanese Ministry of Education kana standardization

  • Historical kana usage studies

  • Unicode Consortium technical reports

License

This library is free software; you can redistribute it and/or modify it under the Artistic License 2.0.

Author

Danslav Slavenskoj

For specific script conversions (Romaji, Cyrillic, Hangul), see the specialized README files:

Lang::JA::Kana v1.2.1

Japanese Hiragana and Katakana conversion utilities

Authors

  • Danslav Slavenskoj

License

Artistic-2.0

Dependencies

Test Dependencies

Provides

  • Lang::JA::Kana
  • Lang::JA::Kana::Hangul
  • Lang::JA::Kana::Kuriru-moji
  • Lang::JA::Kana::Romaji

The Camelia image is copyright 2009 by Larry Wall. "Raku" is trademark of the Yet Another Society. All rights reserved.