Inter national J our nal of Ev aluation and Resear ch in Education (IJERE) V ol. 15, No. 4, August 2026, pp. 3215 3227 ISSN: 2252-8822, DOI: 10.11591/ijere.v15i4.38456 3215 Image-nati v e automated scoring of hand written mathematical r esponses: r eliability e vidence and teacher –AI collaboration JiEun J anet Song 1 , Y oung-Seok Oh 2 , Dong J oong Kim 3 1 Department of Curriculum and Instruction, K orea Uni v ersity , Seoul, Republic of K orea 2 Di vision of Articial Intelligence and Data Science, K orea Cyber Uni v ersity , Seoul, Republic of K orea 3 Department of Mathematics Education, K orea Uni v ersity , Seoul, Republic of K orea Article Inf o Article history: Recei v ed Dec 31, 2025 Re vised Jun 26, 2026 Accepted Jul 6, 2026 K eyw ords: Automated scoring Handwritten mathematics Image-nati v e grading Multimodal lar ge language model Reliability Rubric-based scoring T eacher –AI collaboration ABSTRA CT This study e xamines the reliability of an image-nati v e multimodal AI system for automated scoring of handwritten respons es to Adv anced Placement (AP) Calculus free-response items wit hout requiring optical character recognition (OCR) preprocessing. Using inter -rater agreement indices and test–retest reliability ana lyses, we found substantial to almost perfect agreement between articial i ntelligence (AI)-generated scores and calibrated human ratings, as well as almost perfect stability across repeated scoring sessions. These results suggest that the observ ed reliability of the AI scoring system w arrants further in v estig ation of v alidity-related e vidence and inferences. As a practical implication for assessment practice, we propose a human-in-the-loop teacher –AI collaborati v e (T A C) frame w ork in which automated scoring operates under teacher o v ersight. T ak en together , these ndings pro vide initial e vidence of reliability supporting the responsible use of AI-based scoring as a measurement instrument in high-stak es educational assessment. This is an open access article under the CC BY -SA license . Corresponding A uthor: Dong Joong Kim Department of Mathematics Education, K orea Uni v ersity 145 Anam-ro, Seongb uk-gu, Seoul 02841, Republic of K orea Email: dongjoongkim@k orea.ac.kr 1. INTR ODUCTION A central concern in educational e v aluation is whether assessment scores pro vide a v alid and reliable basis for interpreting student performance [1], [2]. In lar ge-scale ass essment conte xts, this concern becomes especially important because score interpretations often inform consequential decisions about student achie v ement, placement, and instructional decision making. F or this reason, an y scoring system—whether human or automated—must be e xamined not only as a scoring mechanism, b ut also as part of a broader measurement frame w ork. From this perspecti v e, the use of articial intelligence (AI) in assessment raises a psychometric question as much as a technological one: whether AI-generated scores demonstrate suf cient reliability to support defensible interpretations of student performance. In recent years, standardized testing en vironments ha v e increasingly transitioned to digital formats. In 2024, the Colle ge Board announced an accelerated transition to digital administration of Adv anced Placement (AP) e xams, and by the May 2025 administration man y AP e xams were deli v ered through either fully digital or h ybrid digital formats [3], [4]. AP Calculus, including both AB and BC, continues to emplo y a h ybrid digital J ournal homepage: http://ijer e .iaescor e .com Evaluation Warning : The document was created with Spire.PDF for Python.
3216 ISSN: 2252-8822 administration format in which students vie w free-response items in Bluebook, handwrite their responses in paper booklets, and the booklets are subsequently returned for scoring [5], [6]. Colle ge Board materials note that, unlik e multiple-choice sections that are computer -scored, free-response sections are scored by AP teachers and colle ge f aculty [7]. This format preserv es handwritten mathematical e xpression within a broader digital testing en vironment, b ut lea v es scoring w orko ws only partially aligned with the digital transition. From this perspecti v e, image-nati v e automated scoring is rele v ant not because students should enter mathematical w ork digitally , b ut because it may help to connect digital e xam deli v ery with more scalable scoring of handwritten responses while preserving human o v ersight of consequential score interpretations. Automated scoring systems ha v e a long history of de v elopment [8], yet e xtending these approaches to mathematics pr esents distinct challenges. T raditional optical character recogniti on (OCR)-based pipelines ha v e struggled to interpret the spatial and symbolic comple xities of mathematical notation [9], [10], and the recognition errors the y introduce may produce construct-irr ele v ant v ariance that threatens the v alidity of score-based inferences [11]. Recent adv ances in multimodal lar ge language models (LLMs) of fer a potential methodological shift [12], [13]. Models capable of processing both te xtual and visual inputs can bypass the OCR preprocessing step [14]. In particular , OpenAI’ s o1 reasoning model, with enhanced mathematical reasoning through e xtended chain-of-thought process ing [15], of fers a promising approach for image-nati v e automated scoring of handwritten mathematical responses. Ho we v er , empirical e vidence on the reliability of image-nati v e scoring systems for mathe matics assessment remains limited. Although recent studies ha v e e xplored LLM-based grading of handwritten mathematics [16]–[18], most still rely on OCR as an intermediate step and ha v e reported limited grading accurac y for handwritten science, technology , engineering, and mathematics (STEM) responses [19]. As the broader use of LLMs in education e xpands [20], researchers ha v e also emphasized the importance of e xplainable AI (XAI) frame w orks to ensure that algorithmic decisions remain transparent and interpretable [21]. T o our kno wledge, no pre vious study has systematically e xamined whether image-nati v e AI scoring produces stable scores across repeated e v aluations of the same response. This study mak es tw o related contrib utions. First, it pro vides initial empirical e vidence for considering image-nati v e AI scoring in educational assessment and for framing the v alidity questions that follo w from its use. Second, it introduces a teacher –AI collaborati v e (T A C) frame w ork in which AI functions as a rst-pass scoring instrument while educators retain authority o v er nal score interpretation. Accordingly , this study e xamines the reliability of an image-nati v e automated scoring system based on OpenAI’ s o1 multimodal reasoning model for handwritten AP Calculus free-response items and considers its implications for assessment practice. T o what e xtent do AI-generated scores using direct image input demonstrate acceptable inter -rater agreement with human-assigned scores on AP Calculus AB/BC free-response items? (RQ1) T o what e xtent does the image-nati v e AI scoring system demonstrate score stability (test–retest reliability) across repeated e v aluations of the same student response images? (RQ2) Ho w does measurement error in AI-generated scores relati v e to human ratings v ary across student procienc y le v els? (RQ3) The d e v el op m ent of automated scoring systems traces back to P age’ s Project Essay Grade [8] in 1966, which demonstrated the potential for computer -based e v aluation of student writing. Subsequent decades brought substantial adv ances, including operational systems such as ETS’ s e-rater , which combines statistical modeling with natural language processing to e v aluate essay quality [22]. The criterion online writing service further e xtended these capabilities by inte grating automated scoring with diagnostic feedback to support formati v e assessment [23]. More recent studies ha v e applied deep learning approaches to automated essay scoring, achie ving le v els of agreement with human raters in man y conte xts [24], [25]. Ho we v er , automated scoring is fundamentally an assessment v alidity issue rather than merely a technical problem. Bennett and Bejar [11] emphasized that score agreement alone is ins uf cient e vidence of v alidity; the scoring process must represent the intended construct and support v alid score interpretations. This perspecti v e is particularly important for mathematics assessment, where the construct e xtends be yond nal answers to include reasoning processes, problem-solving strate gies, and mathematical communication [26]. As a prerequisite for an y v alidity ar gument, reliability must rst be established in e v aluating automated scoring systems [1]. In educational assessment, reliability is typically e xamined along tw o dimensions. The rst, inter -rater reliability , refers to the de gree to which dif ferent raters—in this conte xt, Int J Ev al & Res Educ, V ol. 15, No. 4, August 2026: 3215-3227 Evaluation Warning : The document was created with Spire.PDF for Python.
Int J Ev al & Res Educ ISSN: 2252-8822 3217 AI and human scorers—produce consistent scores for the same response. I t is commonly assessed using indices such as Cohen’ s κ , quadratic weighted κ (QWK), and e xact match (EM) rate [24], [25]. According to the benchmarks proposed by Landis and K och [27], κ v alues of .61–.80 indicate substantial agreement, whereas v alues abo v e .81 indicate almost perfect agreement. The second dimension, test–retest reliability , refers to the consistenc y of scores when the same response is e v aluated across repeated scoring trials. This dimension is particularly important in AI-based scoring because generati v e models may produce stochastic outputs. As emphasized in prior research on open-ended mathematics tasks [26], stable scores across repeated e v aluations are essential for defensible score interpretation. W ithout reliability e vidence, subsequent v alidity claims lack a sound empirical foundation. Mathematical responses present distinct challenges compared with essay scoring [28]. Student w ork often combines natural language, symbolic notation, diagrams, and graphs in spatiall y comple x arrangements that resist linear te xt processing [28]. Surv e ys of mathematical e xpression recognition emphasize that ef fecti v e systems must address tw o-dimensional layout, symbol ambiguity , and st ructural parsing, which are not easily handled by standard te xt recognition alone [9]. Benchmark competitions such as competition on recognition of online handwritten mathemati cal e xpressions (CR OHME) ha v e repeatedly sho wn that handwritten mathematical e xpression recognition remains challenging despite adv ances in deep learning [29]. T raditional pipelines that con v ert images to symbolic te xt through OCR before scoring f ace compounding errors where recognition mistak es propag ate into scoring decisions [10]. From an educational measurement perspecti v e, such OCR-induced errors introduce construct-irrele v ant v ariance—meas urement noise attrib utable to the recognition process rather than to the student’ s mathematical procienc y—thereby threatening the v alidity of score-based inferences deri v ed from automated scoring systems. This limitation has moti v ated the search for alternati v e approaches that can e v aluate mathematical w ork more holistically . Recent adv ances in multimodal LLMs ha v e e xpanded the methodological possibilities for educati on a l applications [12], [13]. Models such as GPT -4V can process images t og e ther with te xt, making it possible to analyze handwritten w ork without rst con v erting it through OCR [14]. OpenAI’ s technical report on GPT -4 describes both the multimodal capabilities and the safety considerations associated with vision-enabled deplo yment [30]. The o1 series is particularly rele v ant in this conte xt because the task in v olv es more than vis ual recognition alone. Scoring handwritten mathematics also requires judgments about logical structure, the suf cienc y of the e vidence pro vided, and whether a solution actually satises the rubric [15]. Recent studies ha v e be gun to e xplore LLM-based grading in educational settings. Prior w ork has e xamined handwritten ph ysics and mathematics e xams [16], [18], [19], while other studies ha v e in v estig ated written assignments in other disciplines [17]. Related research on chain-of-thought prompting suggests that e xplicit reasoning traces may impro v e scoring transparenc y and consistenc y [31]. The emer gence of multimodal and reasoning-oriented models has not eliminated the need for psychometric scrutin y , b ut it has made image-nati v e scori ng a more plausible object of in v estig ation. Earlier automated scoring pipelines for handwritten w ork typically relied on OCR or handcrafted feature e xtraction. By contrast, GPT -4 and GPT -4V made i t more realistic to in v esti g at e whether handwritten mathematical responses could be scored directly as images [14]. The later introduction of the o1 series further strengthened this possibility for tasks whose e v aluation depends on more than surf ace-le v el recognition [15]. That technological shift does not establish v alidity on its o wn, b ut it does mak e image-nati v e scoring a more credible object of psychometric in v estig ation. Rather than positioning AI as a replacement for human judgment , emer ging frame w orks emphasize collaborati v e approaches where AI assists teachers while humans retain o v ersight authority [20], [21]. Myint et al. [32] proposed an AI-instructor collaborati v e grading approach that incorporates instructor - rened marking schemes and emphasizes e xplainability and f airnes s. This perspecti v e also aligns with broader concerns that trustw orth y multimodal AI systems require transparenc y , f airness, and meaningful human o v ersight, particularl y in consequential decision conte xts [33]. Such collaborati v e frame w orks are particularly important gi v en ongoing concerns about AI reliability , f airness, and v alidity in educational assessment [34], [35]. While AI systems may pro vide ef cienc y g ains, maintaining teacher in v olv ement ensures that scoring decisions remain grounded in pe d a go gi cal e xpertise and can be adapted to indi vidual student circumstances that automated systems might not fully capture. Ima g e-native automated scoring of handwritten mathematical r esponses: r eliability ... (JiEun J anet Song) Evaluation Warning : The document was created with Spire.PDF for Python.
3218 ISSN: 2252-8822 2. RESEARCH METHOD 2.1. Resear ch design This study emplo yed a quantitati v e reliability design to e v aluate the performance of an i mage-nati v e automated scoring system for handwritten mathematical responses. The design compared AI-generated scores with of cial Colle ge Board scores across multiple response samples and repeated grading trials. Figure 1 illustrates the o v erall approach, which combines a public data collection pipeline with a direct image-nati v e scoring architecture. T o establish reliability , we e v aluated inter -rater agreement between AI-generated scores and of cial Colle ge Board scores, as well as test-retest consistenc y across repeated grading trials. Data sour ce Sampling AI scoring Reliability analysis Colle ge Board AP central (2022–2024) 136 FRQ items (AB + BC) 3 samples per item (A, B, C) 408 r esponse images OpenAI o1 Image-nati v e (No OCR) Blind scoring × 7 trials 2,856 total Grading runs Inter -rater AI vs. CB of cial (EM, κ , QWK, r ) T est-r etest 7 trials consistenc y (Cronbach’ s α ) Note: e xaminee identity acr oss items within a year is un veriable for eac h Sample A, B, or C. Colle g e Boar d Of cial Scor es (Gold Standar d) Figure 1. Ov ervie w of the image-nati v e AI scoring design for 408 responses e v aluated across se v en trials 2.2. System ar chitectur e and experimental setup T w o system components were de v eloped: an interacti v e user interf ace for indi vidual grading s essions and an automated batch grading system for lar ge-scale reliability e v aluation. 2.2.1. A utomated batch grading system T o conduct the 2,856 grading trials, we implemented an automated batch grading system that operated independently of a user interf ace. This system e x ecuted blind scoring, meaning that the AI grader had no access to human-assigned scores during the grading process. The design ensured that scoring decisions were based solely on rubric criteria, pre v enting score anchoring or conrmation bias. The system accessed a structured repository containing: i) student response images or g anized by e xamination year , course, and question number , and ii) of cial scoring rubrics specifying point allocation criteria for each item. Importantly , the system did not recei v e of cial w ork ed solutions or model answers. Instead, the AI e v aluated responses e xclusi v ely ag ainst the rubric criteria, mirroring the rubric-based e v aluation process used by human AP readers. F or each trial, the system generated a prompt containing the response image and the corresponding rubric and submitted it to the OpenAI application programming interf ace (API) using the o1 model (resolving to the o1-2024-12-17 snapshot). All grading trials were conducted on March 12–13, 2025. The API returned structured outputs includi ng : i) criterion-le v el scoring decisions, ii) the total score deri v ed from those decisions, and iii) e xplanatory scoring commentary . All outputs were automatically parsed and e xported to CSV format for subsequent statistical analysis. Figure 2 illustrates this blind scoring pipeline used in this study . 2.2.2. Scope of the pr esent study The present study focuses e xclusi v ely on score-based reliability metrics deri v ed from the automated batch grading outputs. The criterion-le v el commentary generated by the system pro vides rich qualitati v e data for v alidity a nalysis—specically , whet h e r the AI’ s scoring rational es accurately reect the rubric criteria and mathematical concepts. Ho we v er , v al idity auditing of AI reasoning processes lies be yond the scope of the present study and remains a ne xt step b uilt on this reliability e vidence. Int J Ev al & Res Educ, V ol. 15, No. 4, August 2026: 3215-3227 Evaluation Warning : The document was created with Spire.PDF for Python.
Int J Ev al & Res Educ ISSN: 2252-8822 3219 Student response image (handwritten) Of cial rubric (point criteria) Output schema (structured format) OpenAI o1 Multimodal reasoning model T otal score Criterion-le v el scoring decisions Scoring commentary (rationale) Human-assigned scores Model solutions Not pro vided to the scoring system Figure 2. Blind automated scoring architecture 2.3. Pr ompt structur e and output enf or cement A k e y des ign goal w as to maximize scoring transparenc y and reduce stochastic v ariabil ity across repeated grading. The prompt includes: i) the student’ s handwritten response and ii) the of cial rubric specifying point allocation criteria, follo wed by iii) a required output structure. This structure forces the model to report criterion-le v el decisions (e.g., whether each point is earned) and to compute the nal total score from those decisions. The design w as informed by the nding that prompting strate gies can elicit more e xplicit reasoning beha viors [31], while still constraining outputs to rubric-based justicat ions suitable for teacher re vie w . 2.4. P opulation and sampling The tar get population for this study comprises student responses to AP Calculus AB/BC FRQs. The sample w as dra wn from publicly a v ailable materials pro vided on the Colle ge Board’ s AP central e xam pages for AP Calculus AB and AP Calculus BC [36], [37]. These materials include of cial e xam quest ions, scoring guidelines, sample responses, and scoring distrib utions for recent administrations. 2.4.1. Human r efer ence scor es A k e y methodological consideration in this study is the source of human reference scores used for AI–human agreement analysis. Rather than relying on researcher -assigned scores, which coul d introduce in v estig ator bias, we used the of cial scores published by the Colle ge Board in its annual scoring guidelines and commentary for each sample response. These scores are assigned by trained AP readers—e xperienced educators who under go standardized calibration procedures to ensure consistent application of scoring rubrics across lar ge-scale e xaminations. F or each of the 408 sample respons e images, the AI-generated score w as compared with the corresponding of cial score published for that item. This approach w as adopted f o r tw o reasons. First, of cial Colle ge Board scores pro vide an e xternally calibrated reference standard based on the systematic training of AP readers. Second, using publicly a v ailable scoring materials enhances the transparenc y and reproducibility of the study while minimizing potential in v estig ator bias that might arise if the responses were rescored by the researchers. 2.4.2. Sample selection criteria No e xclusion criteria were applied. All publicly a v ailable sample responses from the 2022–2024 AP Calculus AB and BC e xaminations were included in t he study . Because the Colle ge Board publishes a x ed set of three sample responses per item, no additional researcher selection or ltering w as required. This complete inclusion approach w as adopted to maximize reproducibili ty and to a v oid researcher -introduced selection bias, ensuring that the dataset reects the full range of publicly a v ailable materials without subjecti v e screening. Ima g e-native automated scoring of handwritten mathematical r esponses: r eliability ... (JiEun J anet Song) Evaluation Warning : The document was created with Spire.PDF for Python.
3220 ISSN: 2252-8822 2.4.3. Sampling method F or each FRQ item, the Colle ge Board publishes three sample student responses in its annual scoring commentary , labeled Sample A, Sampl e B, and Sample C. These samples are selected to represent high-, medium-, and lo w-performing responses, respecti v ely , based on the student’ s o v erall e xamination performance across FRQ items. Because e xaminee identities are not disclosed, it cannot be determined whether samples labeled A, B, or C across dif ferent items originate from the same student. Each published response w as theref o r e treated as an independent observ ation in this study . W e adopted a complete enumeration approach. F or e v ery FRQ item across the 2022–2024 administrations of AP Calculus AB and BC, all t h r ee published sample responses (A, B, and C) were collected. This procedure yielded the full set of publicly a v ailable sample responses without additional researcher selection, thereby ensuring reproducibility . 2.4.4. Sample size The nal dataset included 136 AP Calculus AB/BC FRQ items from the 2022–2024 e xaminations. W ith three sample responses per item (A, B, and C), this yielded 408 handwritten responses ( 136 × 3 ). Each response w as scored se v en times, resulting in 2,856 AI grading runs ( 408 × 7 ). T abl e 1 summarizes the dataset composition by year , course, and e xperimental runs. T able 1. Summary of AP Calculus AB/BC FRQ dataset and e xperimental runs (2022–2024) [36], [37] Y ear Course FRQ items Samples per item Repetitions T otal trials 2022 AB 24 3 7 504 2023 AB 24 3 7 504 2024 AB 21 3 7 441 2022 BC 23 3 7 483 2023 BC 22 3 7 462 2024 BC 22 3 7 462 T otal 136 408 2,856 2.5. Student pr ociency gr ouping T o e v aluate whether grading reliabil ity v aries by student procienc y , responses were analyzed using the Colle ge Board’ s o wn sample designations. Published responses are cate gorized into three performance tiers—Sample A, Sample B, and Sample C—based on total e xamination scores, as in T able 2. In this study , Group A corresponds to Sample A responses, Group B to Sample B responses, and Group C to Sample C responses. T able 2. Student performance group denitions based on Colle ge Board sample designations and total e xamination score (out of 54) Group Sample designation Score range Description A Sample A High High-performing students B Sample B Medium Medium-performing students C Sample C Lo w Lo w-performing students 2.6. Statistical analysis Inter -rater reliability between AI-generated scores and of cial Colle ge Board scores w as e v aluated using four complementary metrics: i) EM rate, representing the proportion of responses recei ving identical scores; ii) adjacent agreement rate, capturing agreement within ± 1 point; iii) Cohen’ s κ (chance-corrected agreement); and i v) QWK, which weights disagreements by their magnitude. The selection of these indices follo ws the standards for educational and psychological testing, which emphasize reliability e vidence as a prerequisite for v alid score interpretations in educational assessment [1]. T est-retest reliability across the se v en repeated trials w as assessed using the intraclass correlation coef cient ( ICC), with Cronbach’ s alpha reported as a supplementary inde x of score consistenc y across tria ls. Group dif ferences in agreement metrics were analyzed using one-w ay analysis of v ariance (ANO V A) with T uk e y’ s honestly signicant dif ference (HSD) post hoc tests, and chi-square tests were used to e xamine group dif ferences in EM rates. Agreement analys es were conducte d at the sub-item le v el using the median AI score a cross se v en tri als for each of the 408 responses, while repeated-score consistenc y analyses used all se v en trial-specic scores. Although AP FRQs v a ry in total point v alues, all analyses were performed using the of cial sub-item score scales dened in the Colle ge Board rubrics. AP Calculus FRQs consist of six questions (maximum of 9 points Int J Ev al & Res Educ, V ol. 15, No. 4, August 2026: 3215-3227 Evaluation Warning : The document was created with Spire.PDF for Python.
Int J Ev al & Res Educ ISSN: 2252-8822 3221 each; total =54) with each question di vided into three to four sub-items scored from 0 to 5 points. Reliability analyses were conducted at the sub-item le v el because total scores can obscure compensatory scoring errors, whereas sub-item analysis pro vides a more precise e v aluation of scoring agreement in rubric-based assessment. 3. RESUL TS AND DISCUSSION 3.1. Inter -rater r eliability: AI–human agr eement Using the median AI score across se v en trials for each of the 408 responses, the image-nati v e AI scoring system demonstrated strong agreement with of cial human reference scores. As summarized in T able 3, chance-corrected agreement reached le v els con v entionally interpreted as substantial to almost perfect according to the Landis and K och [27] benchmarks, indicating that AI-generated scores closely approximated e xpert human judgment under rubric-based scoring conditions. The confusion m atrix in Figure 3, further supports this pattern, with e xact agreement in the majority of responses and nearly all remaining discrepancies limited to one-point dif ferences. T able 3. Ov erall AI–human agreement and repeated-score consistenc y indices Statistical metric V alue interpretation EM rate 83% Exact AI-human score agreement in most responses Adjacent agreement ( ± 1 point) > 95% Most discrepancies were minor Cohen’ s Kappa ( κ ) .77 Substantial chance-corrected agreement QWK .88 Strong agreement with hea vier penalty for lar ge discrepancies Pearson correlation ( r ) .88 Strong linear association between scores Intraclass correlation (ICC(3,1)) .88 High absolute agreement Cronbach’ s alpha ( α ) .99 Almost perfect consistenc y across repeated AI grading trials 62 18 8 0 1 0 8 87 18 2 0 0 0 8 126 5 0 0 0 1 0 49 0 0 0 0 0 0 12 0 0 0 0 0 0 3 0 0 1 1 2 2 3 3 4 4 5 5 AI pr edicted scor e Human actual scor e 0 63 126 Figure 3. Confusion matrix of median AI scores v ersus of cial reference scores (0–5 scale). Diagonal cells represent e xact agreement (339/408 = 83%) 3.2. T est-r etest r eliability: stability acr oss r epeated trials Despite the probabilistic nature of generati v e language models, repeated scoring trials sho wed almost perfect stability , indicating minimal stochast ic v ariability in the AI scoring process. The image-nati v e AI scoring system therefore demonstrated strong test–retest reliability across repeated e v aluations of the same responses. T able 3 summarizes the high le v el of agreement and score consistenc y observ ed across trials. Ima g e-native automated scoring of handwritten mathematical r esponses: r eliability ... (JiEun J anet Song) Evaluation Warning : The document was created with Spire.PDF for Python.
3222 ISSN: 2252-8822 3.3. P erf ormance differ ences by student pr ociency Although o v erall agreement and score stability were strong, agreement v aried across student procienc y le v els. Agreement metrics were therefore anal yzed by procienc y tier . As sho wn in Figure 4, AI–human agreement dif fered systematically by group. Group A (high-performing) demonstrated almost perfect agreement across indices, whereas Groups B and C sho wed lo wer b ut still substantial agreement. Agreement declined modestly at lo wer procienc y le v els, consistent with greater scoring ambiguity near performance thresholds. A one-w ay ANO V A on absolute error scores re v ealed a signicant ef fect of student procienc y group ( F (2 , 405) = 17 . 21 , p < . 001 , η 2 = . 08 ). T uk e y HSD post-hoc tests indicated that Group A dif fered signicantly from both Groups B and C ( p < . 001 ), whereas Groups B and C did not dif fer signicantly ( p . 99 ). A chi-square test on EM rates also conrmed signicant group dif ferences ( χ 2 (2) = 37 . 99 , p < . 001 ). These ndings suggest that the AI system performs more consistent ly when student responses are well-structured and aligned with e xpected solution approaches. Responses with incomplete reasoning, uncon v entional methods, or substantial errors present great er challenges for automated scoring, reecting the inherent dif culty of interpreting ambiguous mathematical w ork [9], [29]. Group A (High) Group B (Medium) Group C (Lo w) 0 20 40 60 80 100 99 . 3 75 . 0 75 . 0 98 . 9 64 . 5 61 . 5 99 . 6 76 . 0 65 . 5 V alue (% or × 100) EM (%) Cohen’ s κ ( × 100) QWK ( × 100) Figure 4. AI–human scoring agreement by student procienc y group 3.4. Qualitati v e err or analysis T o identify the sources of AI–human scoring dis crepancies, we conducted a qualitati v e analysis of responses where AI scores di v er ged from of cial Colle ge Board scores. This analysis focused on Group B (medium-performing) and Group C (lo w-performing) responses, where disagreement rates were highest. T w o recurring error patterns emer ged as the primary sources of these scoring discrepancies. 3.4.1. Outlier identication Across all 2,856 grading trials, v e cases sho wed lar ge discrepancies (3–4 points) between AI and human scores in at least one trial. One of these—a Group B response that matched the of cial score in six of se v en trials—represented an isolated stochastic de viation rather than a systematic scoring error and w as e xcluded from further qualitati v e analysis. The remaining four cases in v olv ed three distinct items, as one item appeared i n both t h e AB and BC e xaminat ions. All four occurred i n lo w-performing responses (Group C) that recei v ed a human-assigned score of 0 b ut were a w arded points by the AI system in multiple trials. As sho wn in T able 4, Cases 1 and 2 in v olv ed the same item (2022 AP Question 1-d), which appeared in both the AB and BC e xam inations. Case 1 sho wed persistent o v er -scoring, with the AI a w arding 4 points in 6 of 7 trials despite the of cial score of 0. Case 2 displayed an alternating pattern (0, 4, 0, 4, 0, 4, 0), suggesting ambiguity in the AI’ s interpretation. Cases 3 and 4 sho wed intermittent lar ge de viations, with the AI occasionally a w arding 2–3 points. 3.4.2. P atter n 1: inf ormal cancellation marks (scrib ble-outs) In Case 3 (2023 BC Question 6-c), the student’ s response contained e xtraneous markings, including scratch w ork that the student attempted to cancel using rough strik ethroughs and zigzag cross-outs rather than Int J Ev al & Res Educ, V ol. 15, No. 4, August 2026: 3215-3227 Evaluation Warning : The document was created with Spire.PDF for Python.
Int J Ev al & Res Educ ISSN: 2252-8822 3223 clean erasure. Human AP readers recognized these markings as cancellation indicators and e xcluded them from e v aluation, scoring only the int ended nal answer as a 0. In contrast, the AI system interpreted these cross-out marks as v alid mathematical content, occasionally assigning inated scores. This pattern appeared more frequently among lo wer -performing students, who often re vised their w ork without fully erasing earlier attempts. T able 4. AI scores for outlier cases. Se v en grading trials for Group C responses with of cial scores of 0 Case T est Item 1st 2nd 3rd 4th 5th 6th 7th CB score 1 2022 AB 1-d 0 4 4 4 4 4 4 0 2 2022 BC 1-d 0 4 0 4 0 4 0 0 3 2023 BC 6-c 0 0 3 0 0 2 0 0 4 2024 BC 2-c 3 0 0 0 0 0 0 0 3.4.3. P atter n 2: notation ambiguity and non-linear spatial lay out F or Cases 1, 2, and 4, qualitati v e analysis re v ealed a s econd error pattern in v olving ambiguous notation and non-linear spatial or g anization. These responses contained non-standard notation that human graders interpreted as incorrect b ut that the AI system occasionally interpreted more f a v orably . The tw o- dimensional layout appeared to hinder the AI’ s ability to reconstruct the logical structure of the response. Human graders therefore assigned zero points despite the presence of isolated v alid symbols, whereas the AI system sometimes credited such fragments without fully e v aluating their role in the o v erall reasoning. 3.4.4. Implications f or err or r emediation The identied error patterns suggest se v eral directions for remediation. First, prompts can be re vised to encourage holistic e v aluation rather than crediting isolated notation fragments. In follo w-up prompt tests, adding e xplicit holistic e v aluation instructions restored correct 0-point scoring for Cases 1, 2, and 4 without reducing agreement on well-formed responses. Second, impro v ed image preprocessing may help lter visual noise such as scribble-outs and o v erwritten w ork before scoring. These ndings also suggest that some errors arise from the visual and structural characteristics of handwritten responses rather than from mathematical reasoning alone. Problematic responses often contained partially erased w ork, o v erwritten notation, or intermediate steps that remained visually salient despite not representing the intended solution. Such features appear to increase the lik elihood that the model interprets residual notation as e vidence for partial credit. Future impro v ements in vision-language models may reduce some of these limitations, although this remains an empirical question [15], [16]. As these models become better at distinguishing intended mathematical content from visual artif acts, the need for e xternal preprocessing may diminish. Enhanced reasoning capabilities may also impro v e the e v aluation of incomplete or uncon v entional responses by helping the model distinguish responses that genuinely merit partial credit from those that merely resemble mathematically rele v ant notation. Another notable nding is the directional asymmetry in AI–human score discrepancies acr o s s procienc y le v els. Agreement w as almost perfect for Group A, whereas Groups B and C sho wed a greater tendenc y for the AI system to assign scores abo v e the human reference. This pattern is consistent with a possible lenienc y ef fect, particularly for responses containing incomplete reasoning, ambiguous notat ion, or uncon v entional solution strate gies. Ho we v er , this interpretation should be treated cautiously because score range restriction may ha v e reduced the visibility of o v er -scoring among high-performing responses near the rubric ceili ng. The present ndings therefore do not determine whether the observ ed asymmetry reec ts a lo wer -procienc y-specic issue or a broader scoring tendenc y mask ed by ceiling ef fects. 3.5. Inter pr etation of r eliability e vidence The central issue in this study is whether AI-generated scores are reliable enough to support educational interpretation, not merely whether the y appear plausible on the surf ace. In classical measurement theory , reliability is a necessary , though not suf cient, condition for v alidity [11]. The results indicate substantial to almost perfect agreement with e xpert human raters and high score stability across repeated trials, pro viding an empirical foundation for further v alidation. At the same time, these ndings should be interpreted as reliability e vidence rather than as proof of construct v alidity or operational readiness. Ima g e-native automated scoring of handwritten mathematical r esponses: r eliability ... (JiEun J anet Song) Evaluation Warning : The document was created with Spire.PDF for Python.
3224 ISSN: 2252-8822 In this study , image-nati v e AI scoring is considered not merely as a technical inno v ation b ut as part of the measurement process itself. In operational assessment conte xts, an y scoring mechanism—human or automated—must demonstrate adequate consistenc y , transparenc y , and interpretability [11]. Remo ving OCR preprocessing may reduce a source of construct-irrele v ant v ariance introduced by transcription errors [9], [10], and may therefore strengthen score consistenc y for handwritten responses. T reating the system as a scoring instrument rather than as a producti vity tool mak es it possible to e v aluate it using established psychometric criteria, including inter -rater reliability , score stability , and error patterns across performance le v els. The reliability metrics reported here compare f a v orably with prior w ork on automated scoring of mathematical responses. K ortem e yer et al. [16] reported substantially lo wer agreem ent using GPT -4-based w orko ws for handwritten ph ysics e xam grading, whereas Liu et al. [18] achie v ed comparable agreement b ut relied on OCR preprocessing t hat introduced additional error sources. In high-sta k es assessment conte xts, this distinction is consequential because scores may inuence placement, certication, and instructional decisions. 3.6. Implications f or teacher –AI collaboration and policy The proposed T A C frame w ork is intended not as a detailed w orko w b ut as a w ay of concept ualizing o v ersight in AI-assisted scoring. In this frame w ork, AI serv es as a rst-pass scoring instrument, whil e educators retain responsibility for nal score interpretation through tar geted re vie w of discrepant or lo w-condence cases. This arrangement preserv es professional judgment while le v eraging algorithmic consistenc y to reduce part of the scoring w orkload [33]. From a measurement perspecti v e, its v alue l ies in treating AI as a supervised component within the scoring process rather than as a subst itute for human raters. At the instructional le v el, the same approach may reduce teachers’ w orkload whil e maintaining their authority o v er score re vie w and interpretation. The k e y issue is therefore not whether human judgment remains in v olv ed, b ut ho w it is inte grated into an AI-assisted scoring process [34], [35]. At present, T A C should be re g arded as a conceptual frame w ork rather than a tested impl ementation model. Questions re g arding teacher acceptance, agging thresholds, w orko w ef cienc y , and the allocation of re vie w responsibility remain empirical issues for future study . F or no w , image-nati v e AI scoring is best vie wed as a rst-pass support mechanism within a human-supervised quality-control process rather than as a replacement for e xpert raters. The k e y question is not whether AI scoring is technically feasi ble, b ut under what conditions its educational use can be justied. In classroom assessment, consistent AI-assisted scoring may help reduce grading w orkload while preserving defensible score interpretation through supervised frame w orks such as T A C. F or e x a mination boards, image-nati v e scoring may pro vide a more scalable approach for e v aluating handwritten mathematical reasoning while a v oiding some transcription errors associated with OCR-based w orko ws. At the polic y le v el, operational use of AI scoring should depend on clear psychometric e vidence. Reliability should be treated as a baseline requirement rather than an assumed consequence of technological inno v ation. 3.7. Limitations and futur e dir ections Se v eral limitations should be ackno wledged. First, the study relied on publicly a v ailable AP sample responses selected for instructional purposes, which may not fully represent the di v ersity of handwriting styles, response patterns, and error types found in li v e e xamination settings. The absence of e xaminee-le v el information also limits inferences about within-student consistenc y across responses. Second, the reported reliability es timates are contingent on the stability of the underlying AI model v ersion. Because model updates may alter scoring beha vior e v en under the same rubric and prompting structure, future research should e xamine score comparability across model v ersions, particularly in high-stak es conte xts. Third, the almost perfect agreement observ ed for high-performing responses should be interpreted cautiously because score range restriction may reduce the visibility of scoring err o r s. In addition, treating all one-point discrepancies as equi v alent may obscure dif ferences in se v erity across sub-items with dif ferent point v alues. Finally , this study pro vides reliability e vidence rather than full v alidation. Future research should therefore e xamine construct alignment, response-process e vidence, f airness across student groups, and the consequences of score use in practice. 4. CONCLUSION This study e xamined whether an image-nati v e AI scoring system based on a multimodal reasoning model can function as a reliable measurement inst rument for handwritten mathematical responses. Across Int J Ev al & Res Educ, V ol. 15, No. 4, August 2026: 3215-3227 Evaluation Warning : The document was created with Spire.PDF for Python.