논문 상세보기

Automatic Mapping of a Large-Scale Dataset of Korean and Vietnamese Readings of Chinese Characters: An Analysis of Phonological Correspondences KCI 등재

Hyeon-yeol IM
  • 언어ENG
  • URLhttps://db.koreascholar.com/Article/Detail/452444
구독 기관 인증 시 무료 이용이 가능합니다. 8,000원
漢字硏究 (한자연구)
경성대학교 한국한자연구소 (Center For The Study of Chinese Charaters in Korea, Kyungsung University)
초록

This study develops an automatic procedure for mapping and analyzing Korean and Vietnamese readings of Chinese characters. The dataset was constructed from the graded Chinese-character list (급수별 배정한자) provided on the official website of the Society for Korean Language & Literary Research (한국어문교육연구회). Characters with multiple Sino-Korean readings were divided into separate records according to their meaning–reading distinctions. When multiple Sino-Vietnamese readings remained unresolved because sufficient lexical context was unavailable, the first-listed reading was used for analysis. The final dataset contained 6,236 records. Unicode code points were used to match the same characters across the Korean and Vietnamese data. Sino-Korean readings were decomposed into onsets, nuclei, and codas by using the structure of precomposed Hangul syllables. Sino-Vietnamese readings were decomposed into onsets, nuclei, codas, and tones through NFD normalization, tone extraction, NFC recomposition, and rule-based parsing. The correspondence analysis was conducted on 6,219 records with valid readings in both languages. The analysis was designed to establish an initial quantitative baseline for the major correspondence patterns across the full dataset. The results showed that coda correspondences were generally more regular than onset and nucleus correspondences. Sino-Korean ㄴ, ㅁ, ㄹ, and ㅂ corresponded mainly to Sino-Vietnamese -n, -m, -t, and -p, respectively. Sino-Korean ㅇ corresponded mainly to -ng and -nh, while ㄱ corresponded mainly to -c and -ch. Onset and nucleus correspondences showed greater variation. The results provide a quantitative basis for comparing the two reading systems and may support further research on historical phonology and Sino-Korean vocabulary education for Vietnamese learners.

키워드
Chinese Character ReadingsSino-Korean ReadingsSino-Vietnamese ReadingsAutomatic MappingSyllable DecompositionPhonological Correspondences
목차
目 錄
1.Introduction
2.Data Construction and Automatic Mapping
    2.1 Dataset Scope and Expansion of Characters with Multiple Readings
    2.2 Unicode-Based Mapping
    2.3 Data Consistency Checks and Exception Handling
3.Automatic Syllable Decomposition and Descriptive Statistics
    3.1 Decomposition of Sino-Korean Readings
    3.2 Decomposition of Sino-Vietnamese Readings
    3.3 Descriptive Statistics of the Full Dataset
4.Phonological Correspondences between Sino-Korean and Sino-Vietnamese Readings
    4.1 Onset Correspondences
    4.2 Nucleus Correspondences and Tone Patterns
    4.3 Coda Correspondences
    4.4 Implications for Teaching Sino-Korean Vocabulary to Vietnamese Learners
5.Conclusion


저자
  • Hyeon-yeol IM(associate professor, the Department of Korean Language and Literature at Chung-Ang University, Republic of Korea)
같은 권호 다른 논문