LLM 멀티에이전트 정책 숙의에서 역할 분화와 증거 활용이 숙의 품질에 미치는 영향
This study analyzes the effects of communicative role differentiation and web-search-based evidence utilization on deliberation quality in LLM (Large Language Model) multi-agent policy deliberation. Drawing on Habermas's theory of communicative action as a theoretical framework, the study evaluates deliberation outcomes not merely by whether or how quickly final responses converge, but by the procedural quality of mutual understanding, validity examination, evidence integration, and productive argumentation. The research design employs a 2×2 full factorial design crossing role differentiation (present/absent) and web search (present/absent), analyzing a total of 48 LLM multi-agent policy deliberation cases. The analytical indicators include Variance and Jensen-Shannon Divergence (JSD) for measuring belief convergence; rapid convergence (CON-R), non-convergence (DIV), and oscillation (OSC) for classifying deliberation patterns; Communicative Differentiation Index (CDI) for measuring communicative act differentiation; Evidence Integration Ratio (EIR) for measuring evidence integration; and deliberative (DEL) and groupthink (GRP) patterns. The analysis found that H1, which posited that role differentiation directly increases final belief convergence, was not supported by either Variance or JSD. However, H2, which posited that role differentiation increases communicative act differentiation, was supported by CDI, and H3, which posited that web search increases external evidence integration, was supported by EIR. H4, which posited that productive deliberation patterns (DEL) are strengthened in the C4 condition combining role differentiation and web search, was also supported. H5, which posited that role differentiation reduces groupthink, was inconclusive due to the absence of GRP cases throughout the entire experiment. Pattern analysis further revealed condition-dependent differences in deliberation pattern distributions: DEL proportions increased under role differentiation conditions, while CON-R and OSC were observed under homogeneous conditions. Although this study is limited by a small sample size and a single LLM environment, it contributes by shifting the evaluative focus of LLM multi-agent policy deliberation from final answers or consensus achievement to the quality of the communicative process. In particular, the study proposes an analytical framework for multidimensional evaluation of deliberation quality through diverse process indicators including CDI, EIR, JSD, and DEL. Future research should incorporate expert validation, sensitivity analysis, and replication across diverse models and policy agendas.