Generative AI in Japanese grammar education: Assessing output quality in honorific, benefactive, and modality expressions
This study asks whether generative AI can be adopted in Japanese grammar education and, if so, in which grammatical areas it can be trusted. Rather than analyzing how a model processes language internally, it treats AI as an educational tool and evaluates the quality of its outputs on classroom-type tasks. A controlled probe set covers honorific (keigo), benefactive (yarimorai), and modality expressions, with case particles and transitivity pairs as control categories, and encodes points of Korean L1 transfer. Outputs from three commercial models (Gemini, ChatGPT, and Claude) are scored through two-pass human coding: a first pass from a Japanese learner’s perspective and a second from a linguistically informed native-speaker perspective. On discrete items the three models are at or near ceiling in accuracy; in context-dependent open generation they diverge in aspect, construction choice, and register—differences the second pass detects where the first does not. Notably, none of the models reproduced the L1-transfer errors typical of Korean learners, suggesting increasingly native-like performance; the residual risks lie instead in avoiding target constructions and in over-formalization. The study offers a category-by-task suitability map and prompt-design and teacher-AI guidelines, arguing for selective rather than wholesale adoption.
本稿は、生成AIを日本語文法教育に導入できるか、できるとすればどの文法領域 まで信頼できるかを問う。モデルの内部処理ではなく、AIを教育の道具とみな し、教室型の課題における産出の質を評価する。統制刺激は敬語・授受表現・モ ダリティを中核に、格助詞と自他動詞対を対照群として含み、韓国語のL1転移 の地点を符号化する。三つの商用モデル(Gemini・ChatGPT・Claude)の産 出を二段階の人手コーディング(日本語の学習者の一次、言語学的知識を持つ母 語話者視点の二次)で評価した。個別項目では三モデルともほぼ天井水準の正確 さを示したが、文脈依存の開放型生成では相・構文選択・文体に差が現れ、その 差は一次では見落とされ二次でのみ検出された。特に、三モデルとも韓国語学 習者に典型的なL1転移の誤りを再現せず、母語話者に近づきつつあることが示 された。残るリスクはむしろ目標構文の回避と過剰な格式化にある。本稿は領 域×課題の適合度マップと、プロンプト設計および教師とAIの役割分担の指針 を示し、導入は一括ではなく選別的であるべきだと論じる。