AN ANALYSIS OF TEACHER-MADE ENGLISH SUMMATIVE TEST ITEMS AT SENIOR HIGH SCHOOL
Main Article Content
Abstract
The implementation of Kurikulum Merdeka has changed English teaching and assessment in Indonesian senior high schools by emphasizing communicative and functional language use. In this context, teacher-made summative tests are expected to measure students’ actual achievement accurately and fairly. This study evaluated the psychometric quality of a teacher-made English summative test under Kurikulum Merdeka. Using a quantitative descriptive evaluative design, the research analyzed 35 multiple-choice items administered to 108 Grade XI students from three classes at SMA Negeri 2 Slawi. Item validity was examined using Point-Biserial correlation, reliability using KR-20, and difficulty using item difficulty index. The pooled analysis showed that all items were valid, with 25.7% highly valid, 62.9% moderately valid, and 11.4% low valid. However, cross-class verification revealed instability, with 1 invalid item in XI-1, 3 in XI-2, and 10 in XI-5, indicating sample dependency bias. The test also showed very high reliability (KR-20 = 0.922). In terms of difficulty, 57.1% of items were easy, 42.9% moderate, and none difficult. These results suggest that although the test was reliable and generally valid, its item functioning was not fully stable across classes and its difficulty level was unbalanced