Research Article
Abstract
References
Information
The purpose of this study was to examine the applicability of AI-based automated scoring for text- and image-based responses in geography assessment from the perspective of reliability. To this end, a restricted-response text item, an extended-response text item, and an image-based item were developed for a high school World Geography course. Responses from 85 high school students were automatically scored using Gemini and GPT models. The AI-generated scores were then compared with scores assigned by five in-service geography teachers to analyze the intra-rater and inter-rater reliability of AI-based automated scoring. The results showed that the reliability of AI-based automated scoring varied by item type and model. First, for the restricted-response text item, both AI models demonstrated high intra-rater and inter-rater reliability. Second, for the extended-response text item, intra-rater reliability was good, whereas inter-rater reliability was lower than the reliability among the teachers. Both AI models also assigned lower scores than the teachers, with GPT exhibiting a particularly strict scoring tendency. Third, for the image-based item, GPT demonstrated high intra-rater reliability and inter-rater reliability approaching the level observed among the teachers. In contrast, Gemini showed moderate intra-rater and inter-rater reliability because of instability in handling blank responses. However, when blank responses were recoded as zero, both reliability measures improved to at least a satisfactory level. By demonstrating that AI-based automated scoring achieved meaningful levels of intra-rater and inter-rater reliability for text- and image-based responses in geography assessment, this study provides empirical evidence for expanding the scope of its application in geography education. Nevertheless, its use requires careful consideration of multiple factors, including prompt design tailored to item types and response characteristics, procedures for handling exceptional responses, and the incorporation of teacher review.
본 연구의 목적은 지리과 텍스트 및 이미지 답안에 대한 AI 자동 채점의 적용 가능성을 신뢰도 측면에서 검증하는 것이다. 이를 위해 고등학교 「세계지리」 과목을 대상으로 제한형, 확장형 텍스트 문항 및 이미지 기반 문항을 개발하고, 해당 문항에 대한 고등학생 85명의 답안을 Gemini와 GPT 모델을 활용하여 자동 채점하였다. 이후 AI 채점 결과를 현직 지리 교사 5명의 채점 결과와 비교하여 AI 자동 채점의 평가자 내 신뢰도와 평가자 간 신뢰도를 분석하였다. 분석 결과, AI 자동 채점의 신뢰도는 문항 유형과 모델에 따라 차이가 있었다. 첫째, 제한형 텍스트 문항에서는 두 AI 모델 모두 평가자 내 신뢰도와 평가자 간 신뢰도가 높게 나타났다. 둘째, 확장형 텍스트 문항에서는 평가자 내 신뢰도는 양호하였으나, 평가자 간 신뢰도는 교사 간 신뢰도보다 낮게 나타났다. 또한 두 AI 모델 모두 교사보다 낮은 점수를 부여하였으며, 특히 GPT의 엄격한 채점 경향이 두드러졌다. 셋째, 이미지 기반 문항에서는 GPT의 경우 높은 평가자 내 신뢰도와 교사 간 신뢰도에 근접한 평가자 간 신뢰도가 확인되었다. 반면 Gemini는 공백 답안 처리의 불안정성으로 평가자 내 및 평가자 간 신뢰도가 중간 수준이었으나, 공백 답안을 0점으로 보정한 경우 일정 수준 이상으로 개선되었다. 본 연구는 지리과의 텍스트 및 이미지 기반 답안에 대한 AI 자동 채점이 의미 있는 수준의 평가자 내 및 평가자 간 신뢰도를 확보하였다는 점에서 지리교육 평가에 AI 자동 채점의 적용 범위를 확장할 수 있는 실증적 근거를 제시한다. 다만 문항 유형과 답안 특성에 적합한 프롬프트 설계, 예외 답안 처리, 교사 검토 절차 도입 등 다양한 측면을 주의 깊게 고려하여 활용할 필요가 있다.
- 경기도교육청, 2025, 「2025 중등 논술형 평가 길라잡이」, 경기: 경기도교육청.
- 교육부, 2024.01.24., 2024년 주요정책 추진계획: 교육개혁으로 사회 난제 해결.
- 교육부, 2025.01.10., 2025년 교육부 주요업무 추진계획: 기회의 사다리가 되는 교육 실현.
- 김주현, 2026, “사회과 서·논술형 문항에 대한 인공지능 자동채점의 신뢰도 분석”, 시민교육연구, 58(1), 133-166.
- 김현미, 2018, “뉴질랜드 지리 평가 탐색: NCEA Level 3을 중심으로”, 한국지리환경교육학회지, 26(1), 135-160. 10.17279/jkagee.2018.26.1.135
- 박종임·최숙기·박강윤·김길재, 2023, 채점자질을 활용한 인공지능 기반 글쓰기 자동채점 방안 탐색, 청람어문교육, 96, 135-172.
- 박혜영·김성숙·김경희·이명진·김광규·김지영, 2019, 「수업-평가 연계 강화를 통한 서·논술형 평가 내실화 방안(연구보고 RRE 2019-6)」, 충북: 한국교육과정평가원.
- 성정원·신병철, 2023, “ChatGPT를 활용한 서·논술형 평가 자동 채점 가능성 탐색: 세계지리 서·논술형 평가를 중심으로”, 한국지리학회지, 12(3), 415-432. 10.25202/JAKG.12.3.3
- 오세준·박정환·백한결·신병철, 2026, “수학 서·논술형 손글씨 이미지 답안에 대한 생성형 AI 자동채점의 특성”, 교과교육학연구, 30(1), 32-44.
- 이간용, 2015, “아프리카 국가들의 지리 평가 특성 비교 분석: 서아프리카 5국, 케냐, 남아프리카공화국의 대학 입학시험을 중심으로”, 한국지리환경교육학회지, 23(3), 127-144.
- 이간용, 2021, “남아시아 4국의 지리평가 특성 비교 분석: 인도의 HSC, 파키스탄의 HSCC, 방글라데시의 HSC, 스리랑카의 GCE A-Level 지리시험을 중심으로”, 한국지리환경교육학회지, 29(3), 1-21.
- 이상균·마갈리 아흐두앙, 2017, “프랑스 지리교육에서 크로키의 등장배경과 제도적 위상, 그리고 활용사례”, 한국지리학회지, 6(3), 369-380. 10.25202/JAKG.6.3.5
- 이인호·이상일·김승현·이정우·서민철·조윤동·이광상·김현경·동효관·배주경·김성혜·권경필·이규호·정기문, 2015, 「국가수준 학업성취도 평가의 서답형 문항 심층 분석(연구보고 RRE 2015-12-2)」, 충북: 한국교육과정평가원.
- 장의선, 2012, “지리과 서답형 문항의 주요 유형에 관한 연구: NAEP의 지리과 4학년 문항을 사례로”, 대한지리학회지, 47(6), 934-954.
- 장의선, 2013, “지리과 서답형 문항의 채점 및 평가 방안에 관한 연구: NAEP 4학년 서답형 답안의 채점 사례를 중심으로”, 사회과교육연구, 20(3), 69-87.
- 최지예·허영수, 2024, “ChatGPT를 활용한 한국어 고급 학습자의 작문 채점의 신뢰도 분석 연구”, 언어사실과 관점, 63, 189-221.
- 함은혜·박소영·이병윤·이성혜·이유경·홍유정, 2024, “GPT-4를 활용한 과학탐구역량 자동 채점의 특성 분석”, 교육정보미디어연구, 30(3), 713-742. 10.15833/KAFEIAM.30.3.713
- 홍세희·노언경·정송·조기현·이현정·이영리, 2020, 교육평가의 기초와 이해, 박영스토리.
- AERA(American Educational Research Association), APA(American Psychological Association), NCME(National Council on Measurement in Education), 2014,
Standards for Educational and Psychological Testing , AERA, Washington, DC. - Alikaniotis, D., Yannakoudakis, H., and Rei, M., 2016, Automatic text scoring using neural networks,
in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics , 1, 715-725. 10.18653/v1/P16-1068 - Baral, S., Botelho, A. F., Santhanam, A., Gurung, A., Cheng, L., and Heffernan, N. T., 2023, Auto-scoring student responses with images in mathematics,
in Proceedings of the 16th International Conference on Educational Data Mining , 362-369. - Bednarz, S. W., Heffron, S., and Huynh, N. T. (eds.), 2013,
A Road Map for 21st Century Geography Education: Geography Education Research (A Report from the Geography Education Research Committee of the Road Map for 21st Century Geography Education Project) , Association of American Geographers, Washington, DC. - Bijsterbosch, E., van der Schee, J., and Kuiper, W., 2017, Meaningful learning and summative assessment in geography education: An analysis in secondary education in the Netherlands,
International Research in Geographical and Environmental Education , 26(1), 17-35. 10.1080/10382046.2016.1217076 - Bourke, T. and Mills, R., 2022, Binaries and silences in geography education assessment research, in Bourke, T., Mills, R., and Lane, R. (eds.),
Assessment in Geographical Education: An International Perspective , Springer, Cham, 3-27. 10.1007/978-3-030-95139-9_1 - Budke, A., Schiefele, U., and Uhlenwinkel, A., 2010, ‘I think it’s stupid’ is no argument: Investigating how students argue in writing,
Teaching Geography , 35(2), 66-69. - Burstein, J., 2003, The e-rater scoring engine: automated essay scoring with natural language processing, in Shermis, M. D. and Burstein, J. (eds.),
Automated Essay Scoring: A Cross-disciplinary Perspective , Lawrence Erlbaum Associates, Mahwah, NJ, 113-121. - Chen, Y., Li, Y., Ren, Y., Liu, Y., and Ma, Y., 2025, Educational evaluation with MLLMs: Framework, dataset, and comprehensive assessment,
Electronics , 14(18), 3713. 10.3390/electronics14183713 - Cohen, J., 1988,
Statistical Power Analysis for the Behavioral Sciences , 2nd ed., Lawrence Erlbaum Associates, Hillsdale, NJ. - de Vet, H. C., Terwee, C. B., Knol, D. L., and Bouter, L. M., 2006, When to use agreement versus reliability measures,
Journal of Clinical Epidemiology , 59(10), 1033-1039. 10.1016/j.jclinepi.2005.10.015 - Edelson, D. C., Shavelson, R. J., and Wertheim, J. A. (eds.), 2013,
A Road Map for 21st Century Geography Education: Assessment (A Report from the Assessment Committee of the Road Map for 21st Century Geography Education Project) , National Geographic Society, Washington, DC. 10.1080/19338341.2012.758045 - Foltz, P. W., Yan, D., and Rupp, A. A., 2020, The past, present, and future of automated scoring, in Yan, D., Rupp, A. A., and Foltz, P. W. (eds.),
Handbook of Automated Scoring: Theory into Practice , CRC Press, Boca Raton, FL, 1-9. 10.1201/9781351264808-1 - Greenhouse, S. W. and Geisser, S., 1959, On methods in the analysis of profile data,
Psychometrika , 24(2), 95-112. 10.1007/BF02289823 - Gwet, K. L., 2014,
Handbook of Inter-Rater Reliability: The Definitive Guide to Measuring the Extent of Agreement among Raters , 4th ed., Advanced Analytics LLC., Gaithersburg, MD. - Harwood, D. and Rawlings, K., 2001, Assessing young children's freehand sketch maps of the world,
International Research in Geographical and Environmental Education , 10(1), 20-45. 10.1080/10382040108667422 - Hussein, M. A., Hassan, H., and Nassef, M., 2019, Automated language essay scoring systems: A literature review,
PeerJ Computer Science , 5, e208. 10.7717/peerj-cs.208 33816861 PMC7924549 - Jauhiainen, J. S. and Garagorry Guerra, A., 2024, Evaluating students’ open-ended written responses with LLMs: Using the RAG framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large,
Advances in Artificial Intelligence and Machine Learning , 4(4), 3097-3113. 10.54364/AAIML.2024.44177 - Jiang, N., Huang, Y., and Chen, J., 2025, Comparison of AI and human scoring on a visual arts assessment,
in Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress , 147-154. - Johnson, M. and Zhang, M., 2024, Examining the responsible use of zero-shot AI approaches to scoring essays,
Scientific Reports , 14, 30064. 10.1038/s41598-024-79208-2 39627285 PMC11614890 - Jung, J. Y., Tyack, L., and von Davier, M., 2022, Automated scoring of constructed-response items using artificial neural networks in international large-scale assessment,
Psychological Test and Assessment Modeling , 64(4), 471-494. - Koo, T. K. and Li, M. Y., 2016, A guideline of selecting and reporting intraclass correlation coefficients for reliability research,
Journal of Chiropractic Medicine , 15(2), 155-163. 10.1016/j.jcm.2016.02.012 27330520 PMC4913118 - Kortemeyer, G., 2024, Performance of the pre-trained large language model GPT-4 on automated short answer grading,
Discover Artificial Intelligence , 4, 47. 10.1007/s44163-024-00147-y - Lakens, D., 2013, Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs,
Frontiers in Psychology , 4, 863. 10.3389/fpsyg.2013.00863 24324449 PMC3840331 - Landauer, T., Laham, D., and Foltz, P., 2003, Automatic essay assessment,
Assessment in Education , 10(3), 295-308. 10.1080/0969594032000148154 - Lane, R. and Bourke, T., 2019, Assessment in geography education: A systematic review, International
Research in Geographical and Environmental Education , 28(1), 22-36. 10.1080/10382046.2017.1385348 - Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H.,
et al. , 2020, Retrieval-augmented generation for knowledge-intensive NLP tasks, in NIPS'20: Proceedings of the 34th International Conference on Neural Information Processing Systems, 793, 9459-9473. - Li, Z., Wang, Z., Wang, W., Hung, K., Xie, H., and Wang, F. L., 2025, Retrieval-augmented generation for educational application: A systematic survey,
Computers and Education: Artificial Intelligence , 8, 100417. 10.1016/j.caeai.2025.100417 - McMillan, J. H., 2024,
Classroom Assessment: Principles and Practice That Enhance Student Learning and Motivation , 8th ed., Pearson, Hoboken, NJ. - Mizumoto, A. and Eguchi, M., 2023, Exploring the potential of using an AI language model for automated essay scoring, Research Methods in Applied Linguistics, 2(2), 100050. 10.1016/j.rmal.2023.100050
- Morris, W., Holmes, L., Choi, J. S., and Crossley, S., 2025, Automated scoring of constructed response items in math assessment using large language models,
International Journal of Artificial Intelligence in Education , 35, 559-586. 10.1007/s40593-024-00418-w - Muhammad, L. N., 2023, Guidelines for repeated measures statistical analysis approaches with basic science research considerations,
The Journal of Clinical Investigation , 133(11), e171058. 10.1172/JCI171058 37259921 PMC10231988 - Munowenyu, E., 2007, Assessing the quality of essays using the SOLO taxonomy: Effects of field and classroom-based experiences by ‘A’ level geography students,
International Research in Geographical and Environmental Education , 16(1), 21-43. 10.2167/irg204.0 - Pack, A., Barrett, A., and Escalante, J., 2024, Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability,
Computers and Education: Artificial Intelligence , 6, 100234. 10.1016/j.caeai.2024.100234 - Page, E. B., 1966, The imminence of grading essays by computer,
The Phi Delta Kappan , 47(5), 238-243. - Popham, W. J., 2025,
Classroom Assessment: What Teachers Need to Know , 10th ed., Pearson, Hoboken, NJ. - Quah, B., Zheng, L., Sng, T. J. H., Yong, C. W., and Islam, I., 2024, Reliability of ChatGPT in automated essay scoring for dental undergraduate examinations,
BMC Medical Education , 24, 962. 10.1186/s12909-024-05881-6 39227811 PMC11373238 - Ramineni, C. and Williamson, D., 2018, Understanding mean score differences between the e-raterⓇ automated scoring engine and humans for demographically based groups in the GREⓇ general test,
ETS Research Report Series , 2018(1), 1-31. 10.1002/ets2.12192 - Schober, P., Boer, C., and Schwarte, L. A., 2018, Correlation coefficients: Appropriate use and interpretation,
Anesthesia and Analgesia , 126(5), 1763-1768. 10.1213/ANE.0000000000002864 - Shrout, P. E. and Fleiss, J. L., 1979, Intraclass correlations: Uses in assessing rater reliability,
Psychological Bulletin , 86(2), 420-428. 10.1037/0033-2909.86.2.420 - Sil, P., Bhattacharyya, P., Goyal, P., and Ramakrishnan, G., 2026, Can MLLMs generate human-like feedback in grading multimodal short answers?,
arXiv , arXiv:2412.19755v4 [cs.AI]. - Smith, A., Leeman-Munk, S., Shelton, A., Mott, B., Wiebe, E., and Lester, J., 2019, A multimodal assessment framework for integrating student writing and drawing in elementary science learning,
IEEE Transactions on Learning Technologies , 12(1), 3-15. 10.1109/TLT.2018.2799871 - Taghipour, K. and Ng, H. T., 2016, A neural approach to automated essay scoring,
in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , 1882-1891. 10.18653/v1/D16-1193 - van der Schee, J., Scholten, N., and Caldis, S., 2024, An introduction to geography education, in Bednarz, S. W. and Mitchell, J. T. (eds.),
Handbook of Geography Education , Springer, Cham, 13-34. 10.1007/978-3-031-72366-7_2 - Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I., 2017, Attention is all you need,
in Advances in Neural Information Processing Systems , 30, 5998-6008. - Waugh, C. K. and Gronlund, N. E., 2013,
Assessment of Student Achievement , 10th ed., Pearson, Hoboken, NJ. - Wetzler, E. L., Cassidy, K. S., Jones, M. J., Frazier, C. R., Korbut, N. A., Sims, C. M., Bowen, S. S., and Wood, M., 2025, Grading the graders: Comparing generative AI and human assessment in essay evaluation,
Teaching of Psychology , 52(3), 298-304. 10.1177/00986283241282696 - Williamson, D. M., Xi, X., and Breyer, F. J., 2012, A framework for evaluation and use of automated scoring,
Educational Measurement: Issues and Practice , 31(1), 2-13. 10.1111/j.1745-3992.2011.00223.x - Wise, N. and Kon, J. H., 1990, Assessing geographic knowledge with sketch maps,
Journal of Geography , 89(3), 123-129. 10.1080/00221349008979612 - Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E., 2024, A survey on multimodal large language models,
National Science Review , 11(12), nwae403. 10.1093/nsr/nwae403 39679213 PMC11645129 - Young, J. E., 1994, Reexamining the role of maps in geographic education: Images, analysis, and evaluation,
Cartographic Perspectives , 17, 10-20. 10.14714/CP17.943 - Zhao, W. X., Zhou, K., Li, J., Tang, T., Dong, Z., Hou, Y., Zhang, B.,
et al. , 2026, A survey of large language models,Frontiers of Computer Science , 20(12), 2012627. 10.1007/s11704-026-60308-3
- Publisher :The Korean Association Of Geographic And Environmental Education
- Publisher(Ko) :한국지리환경교육학회
- Journal Title :The Journal of The Korean Association of Geographic and Environmental Education
- Journal Title(Ko) :한국지리환경교육학회지
- Volume : 34
- No :3
- Pages :189~211
- DOI :https://doi.org/10.17279/jkagee.2026.34.3.189


The Journal of The Korean Association of Geographic and Environmental Education






