This exploratory study investigates rater severity, internal consistency, and rater-related interaction effects in a graph-based writing test for English placement purposes. A total of 101 ESL test-takers completed two graph-based writing tasks, and three trained raters evaluated the performances using a five-point analytic scoring rubric. Multi-faceted Rasch Measurement was employed, and the results indicated that, although the raters demonstrated acceptable levels of internal consistency, they differed significantly in overall severity, suggesting that they were not fully interchangeable. Interaction analyses further revealed no unexpected severity or leniency toward specific test-takers. However, differential rater functioning emerged regarding organization, with two raters exhibiting opposing severity tendencies. These findings suggest that rater training may enhance internal consistency but does not fully eliminate differences in rating severity. They further highlight the need for targeted, criterion-focused calibration for organization in graph-based writing tasks, and scoring rubric refinement to support valid score interpretations and fair placement decisions.
The purpose of this study is to develop and implement a customized AI-based speaking diagnosis, learning, and assessment system, SpeakMaster, in order to overcome the lack of systematic evaluation and practice opportunities in school English speaking class. This system integrates automated speaking scoring to provide students with feedback on their speaking abilities across pronunciation, conversation, and presentation. This study adopts a design-based research methodology, demonstrating the development and implementation process. 1,451 students and eight teachers in elementary, middle, and high schools participated in the experiment. Data were collected through learning logs, teacher journals, interviews, and post-surveys. The findings indicate that the system design is appropriate for English class, promoting students’ flow in engaging speaking practice. Students showed motivation and satisfaction while teachers found the system valuable for monitoring student progress and facilitating speaking assessments. Despite the challenges of improving chatbot performance and enhancing scoring reliability, the results suggest that SpeakMaster shows potential to enhance English speaking education.