Inter
national
J
our
nal
of
Ev
aluation
and
Resear
ch
in
Education
(IJERE)
V
ol.
15,
No.
4,
August
2026,
pp.
3215
∼
3227
ISSN:
2252-8822,
DOI:
10.11591/ijere.v15i4.38456
❒
3215
Image-nati
v
e
automated
scoring
of
hand
written
mathematical
r
esponses:
r
eliability
e
vidence
and
teacher
–AI
collaboration
JiEun
J
anet
Song
1
,
Y
oung-Seok
Oh
2
,
Dong
J
oong
Kim
3
1
Department
of
Curriculum
and
Instruction,
K
orea
Uni
v
ersity
,
Seoul,
Republic
of
K
orea
2
Di
vision
of
Articial
Intelligence
and
Data
Science,
K
orea
Cyber
Uni
v
ersity
,
Seoul,
Republic
of
K
orea
3
Department
of
Mathematics
Education,
K
orea
Uni
v
ersity
,
Seoul,
Republic
of
K
orea
Article
Inf
o
Article
history:
Recei
v
ed
Dec
31,
2025
Re
vised
Jun
26,
2026
Accepted
Jul
6,
2026
K
eyw
ords:
Automated
scoring
Handwritten
mathematics
Image-nati
v
e
grading
Multimodal
lar
ge
language
model
Reliability
Rubric-based
scoring
T
eacher
–AI
collaboration
ABSTRA
CT
This
study
e
xamines
the
reliability
of
an
image-nati
v
e
multimodal
AI
system
for
automated
scoring
of
handwritten
respons
es
to
Adv
anced
Placement
(AP)
Calculus
free-response
items
wit
hout
requiring
optical
character
recognition
(OCR)
preprocessing.
Using
inter
-rater
agreement
indices
and
test–retest
reliability
ana
lyses,
we
found
substantial
to
almost
perfect
agreement
between
articial
i
ntelligence
(AI)-generated
scores
and
calibrated
human
ratings,
as
well
as
almost
perfect
stability
across
repeated
scoring
sessions.
These
results
suggest
that
the
observ
ed
reliability
of
the
AI
scoring
system
w
arrants
further
in
v
estig
ation
of
v
alidity-related
e
vidence
and
inferences.
As
a
practical
implication
for
assessment
practice,
we
propose
a
human-in-the-loop
teacher
–AI
collaborati
v
e
(T
A
C)
frame
w
ork
in
which
automated
scoring
operates
under
teacher
o
v
ersight.
T
ak
en
together
,
these
ndings
pro
vide
initial
e
vidence
of
reliability
supporting
the
responsible
use
of
AI-based
scoring
as
a
measurement
instrument
in
high-stak
es
educational
assessment.
This
is
an
open
access
article
under
the
CC
BY
-SA
license
.
Corresponding
A
uthor:
Dong
Joong
Kim
Department
of
Mathematics
Education,
K
orea
Uni
v
ersity
145
Anam-ro,
Seongb
uk-gu,
Seoul
02841,
Republic
of
K
orea
Email:
dongjoongkim@k
orea.ac.kr
1.
INTR
ODUCTION
A
central
concern
in
educational
e
v
aluation
is
whether
assessment
scores
pro
vide
a
v
alid
and
reliable
basis
for
interpreting
student
performance
[1],
[2].
In
lar
ge-scale
ass
essment
conte
xts,
this
concern
becomes
especially
important
because
score
interpretations
often
inform
consequential
decisions
about
student
achie
v
ement,
placement,
and
instructional
decision
making.
F
or
this
reason,
an
y
scoring
system—whether
human
or
automated—must
be
e
xamined
not
only
as
a
scoring
mechanism,
b
ut
also
as
part
of
a
broader
measurement
frame
w
ork.
From
this
perspecti
v
e,
the
use
of
articial
intelligence
(AI)
in
assessment
raises
a
psychometric
question
as
much
as
a
technological
one:
whether
AI-generated
scores
demonstrate
suf
cient
reliability
to
support
defensible
interpretations
of
student
performance.
In
recent
years,
standardized
testing
en
vironments
ha
v
e
increasingly
transitioned
to
digital
formats.
In
2024,
the
Colle
ge
Board
announced
an
accelerated
transition
to
digital
administration
of
Adv
anced
Placement
(AP)
e
xams,
and
by
the
May
2025
administration
man
y
AP
e
xams
were
deli
v
ered
through
either
fully
digital
or
h
ybrid
digital
formats
[3],
[4].
AP
Calculus,
including
both
AB
and
BC,
continues
to
emplo
y
a
h
ybrid
digital
J
ournal
homepage:
http://ijer
e
.iaescor
e
.com
Evaluation Warning : The document was created with Spire.PDF for Python.
3216
❒
ISSN:
2252-8822
administration
format
in
which
students
vie
w
free-response
items
in
Bluebook,
handwrite
their
responses
in
paper
booklets,
and
the
booklets
are
subsequently
returned
for
scoring
[5],
[6].
Colle
ge
Board
materials
note
that,
unlik
e
multiple-choice
sections
that
are
computer
-scored,
free-response
sections
are
scored
by
AP
teachers
and
colle
ge
f
aculty
[7].
This
format
preserv
es
handwritten
mathematical
e
xpression
within
a
broader
digital
testing
en
vironment,
b
ut
lea
v
es
scoring
w
orko
ws
only
partially
aligned
with
the
digital
transition.
From
this
perspecti
v
e,
image-nati
v
e
automated
scoring
is
rele
v
ant
not
because
students
should
enter
mathematical
w
ork
digitally
,
b
ut
because
it
may
help
to
connect
digital
e
xam
deli
v
ery
with
more
scalable
scoring
of
handwritten
responses
while
preserving
human
o
v
ersight
of
consequential
score
interpretations.
Automated
scoring
systems
ha
v
e
a
long
history
of
de
v
elopment
[8],
yet
e
xtending
these
approaches
to
mathematics
pr
esents
distinct
challenges.
T
raditional
optical
character
recogniti
on
(OCR)-based
pipelines
ha
v
e
struggled
to
interpret
the
spatial
and
symbolic
comple
xities
of
mathematical
notation
[9],
[10],
and
the
recognition
errors
the
y
introduce
may
produce
construct-irr
ele
v
ant
v
ariance
that
threatens
the
v
alidity
of
score-based
inferences
[11].
Recent
adv
ances
in
multimodal
lar
ge
language
models
(LLMs)
of
fer
a
potential
methodological
shift
[12],
[13].
Models
capable
of
processing
both
te
xtual
and
visual
inputs
can
bypass
the
OCR
preprocessing
step
[14].
In
particular
,
OpenAI’
s
o1
reasoning
model,
with
enhanced
mathematical
reasoning
through
e
xtended
chain-of-thought
process
ing
[15],
of
fers
a
promising
approach
for
image-nati
v
e
automated
scoring
of
handwritten
mathematical
responses.
Ho
we
v
er
,
empirical
e
vidence
on
the
reliability
of
image-nati
v
e
scoring
systems
for
mathe
matics
assessment
remains
limited.
Although
recent
studies
ha
v
e
e
xplored
LLM-based
grading
of
handwritten
mathematics
[16]–[18],
most
still
rely
on
OCR
as
an
intermediate
step
and
ha
v
e
reported
limited
grading
accurac
y
for
handwritten
science,
technology
,
engineering,
and
mathematics
(STEM)
responses
[19].
As
the
broader
use
of
LLMs
in
education
e
xpands
[20],
researchers
ha
v
e
also
emphasized
the
importance
of
e
xplainable
AI
(XAI)
frame
w
orks
to
ensure
that
algorithmic
decisions
remain
transparent
and
interpretable
[21].
T
o
our
kno
wledge,
no
pre
vious
study
has
systematically
e
xamined
whether
image-nati
v
e
AI
scoring
produces
stable
scores
across
repeated
e
v
aluations
of
the
same
response.
This
study
mak
es
tw
o
related
contrib
utions.
First,
it
pro
vides
initial
empirical
e
vidence
for
considering
image-nati
v
e
AI
scoring
in
educational
assessment
and
for
framing
the
v
alidity
questions
that
follo
w
from
its
use.
Second,
it
introduces
a
teacher
–AI
collaborati
v
e
(T
A
C)
frame
w
ork
in
which
AI
functions
as
a
rst-pass
scoring
instrument
while
educators
retain
authority
o
v
er
nal
score
interpretation.
Accordingly
,
this
study
e
xamines
the
reliability
of
an
image-nati
v
e
automated
scoring
system
based
on
OpenAI’
s
o1
multimodal
reasoning
model
for
handwritten
AP
Calculus
free-response
items
and
considers
its
implications
for
assessment
practice.
–
T
o
what
e
xtent
do
AI-generated
scores
using
direct
image
input
demonstrate
acceptable
inter
-rater
agreement
with
human-assigned
scores
on
AP
Calculus
AB/BC
free-response
items?
(RQ1)
–
T
o
what
e
xtent
does
the
image-nati
v
e
AI
scoring
system
demonstrate
score
stability
(test–retest
reliability)
across
repeated
e
v
aluations
of
the
same
student
response
images?
(RQ2)
–
Ho
w
does
measurement
error
in
AI-generated
scores
relati
v
e
to
human
ratings
v
ary
across
student
procienc
y
le
v
els?
(RQ3)
The
d
e
v
el
op
m
ent
of
automated
scoring
systems
traces
back
to
P
age’
s
Project
Essay
Grade
[8]
in
1966,
which
demonstrated
the
potential
for
computer
-based
e
v
aluation
of
student
writing.
Subsequent
decades
brought
substantial
adv
ances,
including
operational
systems
such
as
ETS’
s
e-rater
,
which
combines
statistical
modeling
with
natural
language
processing
to
e
v
aluate
essay
quality
[22].
The
criterion
online
writing
service
further
e
xtended
these
capabilities
by
inte
grating
automated
scoring
with
diagnostic
feedback
to
support
formati
v
e
assessment
[23].
More
recent
studies
ha
v
e
applied
deep
learning
approaches
to
automated
essay
scoring,
achie
ving
le
v
els
of
agreement
with
human
raters
in
man
y
conte
xts
[24],
[25].
Ho
we
v
er
,
automated
scoring
is
fundamentally
an
assessment
v
alidity
issue
rather
than
merely
a
technical
problem.
Bennett
and
Bejar
[11]
emphasized
that
score
agreement
alone
is
ins
uf
cient
e
vidence
of
v
alidity;
the
scoring
process
must
represent
the
intended
construct
and
support
v
alid
score
interpretations.
This
perspecti
v
e
is
particularly
important
for
mathematics
assessment,
where
the
construct
e
xtends
be
yond
nal
answers
to
include
reasoning
processes,
problem-solving
strate
gies,
and
mathematical
communication
[26].
As
a
prerequisite
for
an
y
v
alidity
ar
gument,
reliability
must
rst
be
established
in
e
v
aluating
automated
scoring
systems
[1].
In
educational
assessment,
reliability
is
typically
e
xamined
along
tw
o
dimensions.
The
rst,
inter
-rater
reliability
,
refers
to
the
de
gree
to
which
dif
ferent
raters—in
this
conte
xt,
Int
J
Ev
al
&
Res
Educ,
V
ol.
15,
No.
4,
August
2026:
3215-3227
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Ev
al
&
Res
Educ
ISSN:
2252-8822
❒
3217
AI
and
human
scorers—produce
consistent
scores
for
the
same
response.
I
t
is
commonly
assessed
using
indices
such
as
Cohen’
s
κ
,
quadratic
weighted
κ
(QWK),
and
e
xact
match
(EM)
rate
[24],
[25].
According
to
the
benchmarks
proposed
by
Landis
and
K
och
[27],
κ
v
alues
of
.61–.80
indicate
substantial
agreement,
whereas
v
alues
abo
v
e
.81
indicate
almost
perfect
agreement.
The
second
dimension,
test–retest
reliability
,
refers
to
the
consistenc
y
of
scores
when
the
same
response
is
e
v
aluated
across
repeated
scoring
trials.
This
dimension
is
particularly
important
in
AI-based
scoring
because
generati
v
e
models
may
produce
stochastic
outputs.
As
emphasized
in
prior
research
on
open-ended
mathematics
tasks
[26],
stable
scores
across
repeated
e
v
aluations
are
essential
for
defensible
score
interpretation.
W
ithout
reliability
e
vidence,
subsequent
v
alidity
claims
lack
a
sound
empirical
foundation.
Mathematical
responses
present
distinct
challenges
compared
with
essay
scoring
[28].
Student
w
ork
often
combines
natural
language,
symbolic
notation,
diagrams,
and
graphs
in
spatiall
y
comple
x
arrangements
that
resist
linear
te
xt
processing
[28].
Surv
e
ys
of
mathematical
e
xpression
recognition
emphasize
that
ef
fecti
v
e
systems
must
address
tw
o-dimensional
layout,
symbol
ambiguity
,
and
st
ructural
parsing,
which
are
not
easily
handled
by
standard
te
xt
recognition
alone
[9].
Benchmark
competitions
such
as
competition
on
recognition
of
online
handwritten
mathemati
cal
e
xpressions
(CR
OHME)
ha
v
e
repeatedly
sho
wn
that
handwritten
mathematical
e
xpression
recognition
remains
challenging
despite
adv
ances
in
deep
learning
[29].
T
raditional
pipelines
that
con
v
ert
images
to
symbolic
te
xt
through
OCR
before
scoring
f
ace
compounding
errors
where
recognition
mistak
es
propag
ate
into
scoring
decisions
[10].
From
an
educational
measurement
perspecti
v
e,
such
OCR-induced
errors
introduce
construct-irrele
v
ant
v
ariance—meas
urement
noise
attrib
utable
to
the
recognition
process
rather
than
to
the
student’
s
mathematical
procienc
y—thereby
threatening
the
v
alidity
of
score-based
inferences
deri
v
ed
from
automated
scoring
systems.
This
limitation
has
moti
v
ated
the
search
for
alternati
v
e
approaches
that
can
e
v
aluate
mathematical
w
ork
more
holistically
.
Recent
adv
ances
in
multimodal
LLMs
ha
v
e
e
xpanded
the
methodological
possibilities
for
educati
on
a
l
applications
[12],
[13].
Models
such
as
GPT
-4V
can
process
images
t
og
e
ther
with
te
xt,
making
it
possible
to
analyze
handwritten
w
ork
without
rst
con
v
erting
it
through
OCR
[14].
OpenAI’
s
technical
report
on
GPT
-4
describes
both
the
multimodal
capabilities
and
the
safety
considerations
associated
with
vision-enabled
deplo
yment
[30].
The
o1
series
is
particularly
rele
v
ant
in
this
conte
xt
because
the
task
in
v
olv
es
more
than
vis
ual
recognition
alone.
Scoring
handwritten
mathematics
also
requires
judgments
about
logical
structure,
the
suf
cienc
y
of
the
e
vidence
pro
vided,
and
whether
a
solution
actually
satises
the
rubric
[15].
Recent
studies
ha
v
e
be
gun
to
e
xplore
LLM-based
grading
in
educational
settings.
Prior
w
ork
has
e
xamined
handwritten
ph
ysics
and
mathematics
e
xams
[16],
[18],
[19],
while
other
studies
ha
v
e
in
v
estig
ated
written
assignments
in
other
disciplines
[17].
Related
research
on
chain-of-thought
prompting
suggests
that
e
xplicit
reasoning
traces
may
impro
v
e
scoring
transparenc
y
and
consistenc
y
[31].
The
emer
gence
of
multimodal
and
reasoning-oriented
models
has
not
eliminated
the
need
for
psychometric
scrutin
y
,
b
ut
it
has
made
image-nati
v
e
scori
ng
a
more
plausible
object
of
in
v
estig
ation.
Earlier
automated
scoring
pipelines
for
handwritten
w
ork
typically
relied
on
OCR
or
handcrafted
feature
e
xtraction.
By
contrast,
GPT
-4
and
GPT
-4V
made
i
t
more
realistic
to
in
v
esti
g
at
e
whether
handwritten
mathematical
responses
could
be
scored
directly
as
images
[14].
The
later
introduction
of
the
o1
series
further
strengthened
this
possibility
for
tasks
whose
e
v
aluation
depends
on
more
than
surf
ace-le
v
el
recognition
[15].
That
technological
shift
does
not
establish
v
alidity
on
its
o
wn,
b
ut
it
does
mak
e
image-nati
v
e
scoring
a
more
credible
object
of
psychometric
in
v
estig
ation.
Rather
than
positioning
AI
as
a
replacement
for
human
judgment
,
emer
ging
frame
w
orks
emphasize
collaborati
v
e
approaches
where
AI
assists
teachers
while
humans
retain
o
v
ersight
authority
[20],
[21].
Myint
et
al.
[32]
proposed
an
AI-instructor
collaborati
v
e
grading
approach
that
incorporates
instructor
-
rened
marking
schemes
and
emphasizes
e
xplainability
and
f
airnes
s.
This
perspecti
v
e
also
aligns
with
broader
concerns
that
trustw
orth
y
multimodal
AI
systems
require
transparenc
y
,
f
airness,
and
meaningful
human
o
v
ersight,
particularl
y
in
consequential
decision
conte
xts
[33].
Such
collaborati
v
e
frame
w
orks
are
particularly
important
gi
v
en
ongoing
concerns
about
AI
reliability
,
f
airness,
and
v
alidity
in
educational
assessment
[34],
[35].
While
AI
systems
may
pro
vide
ef
cienc
y
g
ains,
maintaining
teacher
in
v
olv
ement
ensures
that
scoring
decisions
remain
grounded
in
pe
d
a
go
gi
cal
e
xpertise
and
can
be
adapted
to
indi
vidual
student
circumstances
that
automated
systems
might
not
fully
capture.
Ima
g
e-native
automated
scoring
of
handwritten
mathematical
r
esponses:
r
eliability
...
(JiEun
J
anet
Song)
Evaluation Warning : The document was created with Spire.PDF for Python.
3218
❒
ISSN:
2252-8822
2.
RESEARCH
METHOD
2.1.
Resear
ch
design
This
study
emplo
yed
a
quantitati
v
e
reliability
design
to
e
v
aluate
the
performance
of
an
i
mage-nati
v
e
automated
scoring
system
for
handwritten
mathematical
responses.
The
design
compared
AI-generated
scores
with
of
cial
Colle
ge
Board
scores
across
multiple
response
samples
and
repeated
grading
trials.
Figure
1
illustrates
the
o
v
erall
approach,
which
combines
a
public
data
collection
pipeline
with
a
direct
image-nati
v
e
scoring
architecture.
T
o
establish
reliability
,
we
e
v
aluated
inter
-rater
agreement
between
AI-generated
scores
and
of
cial
Colle
ge
Board
scores,
as
well
as
test-retest
consistenc
y
across
repeated
grading
trials.
Data
sour
ce
Sampling
AI
scoring
Reliability
analysis
Colle
ge
Board
AP
central
(2022–2024)
136
FRQ
items
(AB
+
BC)
3
samples
per
item
(A,
B,
C)
408
r
esponse
images
OpenAI
o1
Image-nati
v
e
(No
OCR)
Blind
scoring
×
7
trials
2,856
total
Grading
runs
Inter
-rater
AI
vs.
CB
of
cial
(EM,
κ
,
QWK,
r
)
T
est-r
etest
7
trials
consistenc
y
(Cronbach’
s
α
)
Note:
e
xaminee
identity
acr
oss
items
within
a
year
is
un
veriable
for
eac
h
Sample
A,
B,
or
C.
Colle
g
e
Boar
d
Of
cial
Scor
es
(Gold
Standar
d)
Figure
1.
Ov
ervie
w
of
the
image-nati
v
e
AI
scoring
design
for
408
responses
e
v
aluated
across
se
v
en
trials
2.2.
System
ar
chitectur
e
and
experimental
setup
T
w
o
system
components
were
de
v
eloped:
an
interacti
v
e
user
interf
ace
for
indi
vidual
grading
s
essions
and
an
automated
batch
grading
system
for
lar
ge-scale
reliability
e
v
aluation.
2.2.1.
A
utomated
batch
grading
system
T
o
conduct
the
2,856
grading
trials,
we
implemented
an
automated
batch
grading
system
that
operated
independently
of
a
user
interf
ace.
This
system
e
x
ecuted
blind
scoring,
meaning
that
the
AI
grader
had
no
access
to
human-assigned
scores
during
the
grading
process.
The
design
ensured
that
scoring
decisions
were
based
solely
on
rubric
criteria,
pre
v
enting
score
anchoring
or
conrmation
bias.
The
system
accessed
a
structured
repository
containing:
i)
student
response
images
or
g
anized
by
e
xamination
year
,
course,
and
question
number
,
and
ii)
of
cial
scoring
rubrics
specifying
point
allocation
criteria
for
each
item.
Importantly
,
the
system
did
not
recei
v
e
of
cial
w
ork
ed
solutions
or
model
answers.
Instead,
the
AI
e
v
aluated
responses
e
xclusi
v
ely
ag
ainst
the
rubric
criteria,
mirroring
the
rubric-based
e
v
aluation
process
used
by
human
AP
readers.
F
or
each
trial,
the
system
generated
a
prompt
containing
the
response
image
and
the
corresponding
rubric
and
submitted
it
to
the
OpenAI
application
programming
interf
ace
(API)
using
the
o1
model
(resolving
to
the
o1-2024-12-17
snapshot).
All
grading
trials
were
conducted
on
March
12–13,
2025.
The
API
returned
structured
outputs
includi
ng
:
i)
criterion-le
v
el
scoring
decisions,
ii)
the
total
score
deri
v
ed
from
those
decisions,
and
iii)
e
xplanatory
scoring
commentary
.
All
outputs
were
automatically
parsed
and
e
xported
to
CSV
format
for
subsequent
statistical
analysis.
Figure
2
illustrates
this
blind
scoring
pipeline
used
in
this
study
.
2.2.2.
Scope
of
the
pr
esent
study
The
present
study
focuses
e
xclusi
v
ely
on
score-based
reliability
metrics
deri
v
ed
from
the
automated
batch
grading
outputs.
The
criterion-le
v
el
commentary
generated
by
the
system
pro
vides
rich
qualitati
v
e
data
for
v
alidity
a
nalysis—specically
,
whet
h
e
r
the
AI’
s
scoring
rational
es
accurately
reect
the
rubric
criteria
and
mathematical
concepts.
Ho
we
v
er
,
v
al
idity
auditing
of
AI
reasoning
processes
lies
be
yond
the
scope
of
the
present
study
and
remains
a
ne
xt
step
b
uilt
on
this
reliability
e
vidence.
Int
J
Ev
al
&
Res
Educ,
V
ol.
15,
No.
4,
August
2026:
3215-3227
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Ev
al
&
Res
Educ
ISSN:
2252-8822
❒
3219
Student
response
image
(handwritten)
Of
cial
rubric
(point
criteria)
Output
schema
(structured
format)
OpenAI
o1
Multimodal
reasoning
model
T
otal
score
Criterion-le
v
el
scoring
decisions
Scoring
commentary
(rationale)
Human-assigned
scores
Model
solutions
Not
pro
vided
to
the
scoring
system
Figure
2.
Blind
automated
scoring
architecture
2.3.
Pr
ompt
structur
e
and
output
enf
or
cement
A
k
e
y
des
ign
goal
w
as
to
maximize
scoring
transparenc
y
and
reduce
stochastic
v
ariabil
ity
across
repeated
grading.
The
prompt
includes:
i)
the
student’
s
handwritten
response
and
ii)
the
of
cial
rubric
specifying
point
allocation
criteria,
follo
wed
by
iii)
a
required
output
structure.
This
structure
forces
the
model
to
report
criterion-le
v
el
decisions
(e.g.,
whether
each
point
is
earned)
and
to
compute
the
nal
total
score
from
those
decisions.
The
design
w
as
informed
by
the
nding
that
prompting
strate
gies
can
elicit
more
e
xplicit
reasoning
beha
viors
[31],
while
still
constraining
outputs
to
rubric-based
justicat
ions
suitable
for
teacher
re
vie
w
.
2.4.
P
opulation
and
sampling
The
tar
get
population
for
this
study
comprises
student
responses
to
AP
Calculus
AB/BC
FRQs.
The
sample
w
as
dra
wn
from
publicly
a
v
ailable
materials
pro
vided
on
the
Colle
ge
Board’
s
AP
central
e
xam
pages
for
AP
Calculus
AB
and
AP
Calculus
BC
[36],
[37].
These
materials
include
of
cial
e
xam
quest
ions,
scoring
guidelines,
sample
responses,
and
scoring
distrib
utions
for
recent
administrations.
2.4.1.
Human
r
efer
ence
scor
es
A
k
e
y
methodological
consideration
in
this
study
is
the
source
of
human
reference
scores
used
for
AI–human
agreement
analysis.
Rather
than
relying
on
researcher
-assigned
scores,
which
coul
d
introduce
in
v
estig
ator
bias,
we
used
the
of
cial
scores
published
by
the
Colle
ge
Board
in
its
annual
scoring
guidelines
and
commentary
for
each
sample
response.
These
scores
are
assigned
by
trained
AP
readers—e
xperienced
educators
who
under
go
standardized
calibration
procedures
to
ensure
consistent
application
of
scoring
rubrics
across
lar
ge-scale
e
xaminations.
F
or
each
of
the
408
sample
respons
e
images,
the
AI-generated
score
w
as
compared
with
the
corresponding
of
cial
score
published
for
that
item.
This
approach
w
as
adopted
f
o
r
tw
o
reasons.
First,
of
cial
Colle
ge
Board
scores
pro
vide
an
e
xternally
calibrated
reference
standard
based
on
the
systematic
training
of
AP
readers.
Second,
using
publicly
a
v
ailable
scoring
materials
enhances
the
transparenc
y
and
reproducibility
of
the
study
while
minimizing
potential
in
v
estig
ator
bias
that
might
arise
if
the
responses
were
rescored
by
the
researchers.
2.4.2.
Sample
selection
criteria
No
e
xclusion
criteria
were
applied.
All
publicly
a
v
ailable
sample
responses
from
the
2022–2024
AP
Calculus
AB
and
BC
e
xaminations
were
included
in
t
he
study
.
Because
the
Colle
ge
Board
publishes
a
x
ed
set
of
three
sample
responses
per
item,
no
additional
researcher
selection
or
ltering
w
as
required.
This
complete
inclusion
approach
w
as
adopted
to
maximize
reproducibili
ty
and
to
a
v
oid
researcher
-introduced
selection
bias,
ensuring
that
the
dataset
reects
the
full
range
of
publicly
a
v
ailable
materials
without
subjecti
v
e
screening.
Ima
g
e-native
automated
scoring
of
handwritten
mathematical
r
esponses:
r
eliability
...
(JiEun
J
anet
Song)
Evaluation Warning : The document was created with Spire.PDF for Python.
3220
❒
ISSN:
2252-8822
2.4.3.
Sampling
method
F
or
each
FRQ
item,
the
Colle
ge
Board
publishes
three
sample
student
responses
in
its
annual
scoring
commentary
,
labeled
Sample
A,
Sampl
e
B,
and
Sample
C.
These
samples
are
selected
to
represent
high-,
medium-,
and
lo
w-performing
responses,
respecti
v
ely
,
based
on
the
student’
s
o
v
erall
e
xamination
performance
across
FRQ
items.
Because
e
xaminee
identities
are
not
disclosed,
it
cannot
be
determined
whether
samples
labeled
A,
B,
or
C
across
dif
ferent
items
originate
from
the
same
student.
Each
published
response
w
as
theref
o
r
e
treated
as
an
independent
observ
ation
in
this
study
.
W
e
adopted
a
complete
enumeration
approach.
F
or
e
v
ery
FRQ
item
across
the
2022–2024
administrations
of
AP
Calculus
AB
and
BC,
all
t
h
r
ee
published
sample
responses
(A,
B,
and
C)
were
collected.
This
procedure
yielded
the
full
set
of
publicly
a
v
ailable
sample
responses
without
additional
researcher
selection,
thereby
ensuring
reproducibility
.
2.4.4.
Sample
size
The
nal
dataset
included
136
AP
Calculus
AB/BC
FRQ
items
from
the
2022–2024
e
xaminations.
W
ith
three
sample
responses
per
item
(A,
B,
and
C),
this
yielded
408
handwritten
responses
(
136
×
3
).
Each
response
w
as
scored
se
v
en
times,
resulting
in
2,856
AI
grading
runs
(
408
×
7
).
T
abl
e
1
summarizes
the
dataset
composition
by
year
,
course,
and
e
xperimental
runs.
T
able
1.
Summary
of
AP
Calculus
AB/BC
FRQ
dataset
and
e
xperimental
runs
(2022–2024)
[36],
[37]
Y
ear
Course
FRQ
items
Samples
per
item
Repetitions
T
otal
trials
2022
AB
24
3
7
504
2023
AB
24
3
7
504
2024
AB
21
3
7
441
2022
BC
23
3
7
483
2023
BC
22
3
7
462
2024
BC
22
3
7
462
T
otal
–
136
408
–
2,856
2.5.
Student
pr
ociency
gr
ouping
T
o
e
v
aluate
whether
grading
reliabil
ity
v
aries
by
student
procienc
y
,
responses
were
analyzed
using
the
Colle
ge
Board’
s
o
wn
sample
designations.
Published
responses
are
cate
gorized
into
three
performance
tiers—Sample
A,
Sample
B,
and
Sample
C—based
on
total
e
xamination
scores,
as
in
T
able
2.
In
this
study
,
Group
A
corresponds
to
Sample
A
responses,
Group
B
to
Sample
B
responses,
and
Group
C
to
Sample
C
responses.
T
able
2.
Student
performance
group
denitions
based
on
Colle
ge
Board
sample
designations
and
total
e
xamination
score
(out
of
54)
Group
Sample
designation
Score
range
Description
A
Sample
A
High
High-performing
students
B
Sample
B
Medium
Medium-performing
students
C
Sample
C
Lo
w
Lo
w-performing
students
2.6.
Statistical
analysis
Inter
-rater
reliability
between
AI-generated
scores
and
of
cial
Colle
ge
Board
scores
w
as
e
v
aluated
using
four
complementary
metrics:
i)
EM
rate,
representing
the
proportion
of
responses
recei
ving
identical
scores;
ii)
adjacent
agreement
rate,
capturing
agreement
within
±
1
point;
iii)
Cohen’
s
κ
(chance-corrected
agreement);
and
i
v)
QWK,
which
weights
disagreements
by
their
magnitude.
The
selection
of
these
indices
follo
ws
the
standards
for
educational
and
psychological
testing,
which
emphasize
reliability
e
vidence
as
a
prerequisite
for
v
alid
score
interpretations
in
educational
assessment
[1].
T
est-retest
reliability
across
the
se
v
en
repeated
trials
w
as
assessed
using
the
intraclass
correlation
coef
cient
(
ICC),
with
Cronbach’
s
alpha
reported
as
a
supplementary
inde
x
of
score
consistenc
y
across
tria
ls.
Group
dif
ferences
in
agreement
metrics
were
analyzed
using
one-w
ay
analysis
of
v
ariance
(ANO
V
A)
with
T
uk
e
y’
s
honestly
signicant
dif
ference
(HSD)
post
hoc
tests,
and
chi-square
tests
were
used
to
e
xamine
group
dif
ferences
in
EM
rates.
Agreement
analys
es
were
conducte
d
at
the
sub-item
le
v
el
using
the
median
AI
score
a
cross
se
v
en
tri
als
for
each
of
the
408
responses,
while
repeated-score
consistenc
y
analyses
used
all
se
v
en
trial-specic
scores.
Although
AP
FRQs
v
a
ry
in
total
point
v
alues,
all
analyses
were
performed
using
the
of
cial
sub-item
score
scales
dened
in
the
Colle
ge
Board
rubrics.
AP
Calculus
FRQs
consist
of
six
questions
(maximum
of
9
points
Int
J
Ev
al
&
Res
Educ,
V
ol.
15,
No.
4,
August
2026:
3215-3227
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Ev
al
&
Res
Educ
ISSN:
2252-8822
❒
3221
each;
total
=54)
with
each
question
di
vided
into
three
to
four
sub-items
scored
from
0
to
5
points.
Reliability
analyses
were
conducted
at
the
sub-item
le
v
el
because
total
scores
can
obscure
compensatory
scoring
errors,
whereas
sub-item
analysis
pro
vides
a
more
precise
e
v
aluation
of
scoring
agreement
in
rubric-based
assessment.
3.
RESUL
TS
AND
DISCUSSION
3.1.
Inter
-rater
r
eliability:
AI–human
agr
eement
Using
the
median
AI
score
across
se
v
en
trials
for
each
of
the
408
responses,
the
image-nati
v
e
AI
scoring
system
demonstrated
strong
agreement
with
of
cial
human
reference
scores.
As
summarized
in
T
able
3,
chance-corrected
agreement
reached
le
v
els
con
v
entionally
interpreted
as
substantial
to
almost
perfect
according
to
the
Landis
and
K
och
[27]
benchmarks,
indicating
that
AI-generated
scores
closely
approximated
e
xpert
human
judgment
under
rubric-based
scoring
conditions.
The
confusion
m
atrix
in
Figure
3,
further
supports
this
pattern,
with
e
xact
agreement
in
the
majority
of
responses
and
nearly
all
remaining
discrepancies
limited
to
one-point
dif
ferences.
T
able
3.
Ov
erall
AI–human
agreement
and
repeated-score
consistenc
y
indices
Statistical
metric
V
alue
interpretation
EM
rate
83%
Exact
AI-human
score
agreement
in
most
responses
Adjacent
agreement
(
±
1
point)
>
95%
Most
discrepancies
were
minor
Cohen’
s
Kappa
(
κ
)
.77
Substantial
chance-corrected
agreement
QWK
.88
Strong
agreement
with
hea
vier
penalty
for
lar
ge
discrepancies
Pearson
correlation
(
r
)
.88
Strong
linear
association
between
scores
Intraclass
correlation
(ICC(3,1))
.88
High
absolute
agreement
Cronbach’
s
alpha
(
α
)
.99
Almost
perfect
consistenc
y
across
repeated
AI
grading
trials
62
18
8
0
1
0
8
87
18
2
0
0
0
8
126
5
0
0
0
1
0
49
0
0
0
0
0
0
12
0
0
0
0
0
0
3
0
0
1
1
2
2
3
3
4
4
5
5
AI
pr
edicted
scor
e
Human
actual
scor
e
0
63
126
Figure
3.
Confusion
matrix
of
median
AI
scores
v
ersus
of
cial
reference
scores
(0–5
scale).
Diagonal
cells
represent
e
xact
agreement
(339/408
=
83%)
3.2.
T
est-r
etest
r
eliability:
stability
acr
oss
r
epeated
trials
Despite
the
probabilistic
nature
of
generati
v
e
language
models,
repeated
scoring
trials
sho
wed
almost
perfect
stability
,
indicating
minimal
stochast
ic
v
ariability
in
the
AI
scoring
process.
The
image-nati
v
e
AI
scoring
system
therefore
demonstrated
strong
test–retest
reliability
across
repeated
e
v
aluations
of
the
same
responses.
T
able
3
summarizes
the
high
le
v
el
of
agreement
and
score
consistenc
y
observ
ed
across
trials.
Ima
g
e-native
automated
scoring
of
handwritten
mathematical
r
esponses:
r
eliability
...
(JiEun
J
anet
Song)
Evaluation Warning : The document was created with Spire.PDF for Python.
3222
❒
ISSN:
2252-8822
3.3.
P
erf
ormance
differ
ences
by
student
pr
ociency
Although
o
v
erall
agreement
and
score
stability
were
strong,
agreement
v
aried
across
student
procienc
y
le
v
els.
Agreement
metrics
were
therefore
anal
yzed
by
procienc
y
tier
.
As
sho
wn
in
Figure
4,
AI–human
agreement
dif
fered
systematically
by
group.
Group
A
(high-performing)
demonstrated
almost
perfect
agreement
across
indices,
whereas
Groups
B
and
C
sho
wed
lo
wer
b
ut
still
substantial
agreement.
Agreement
declined
modestly
at
lo
wer
procienc
y
le
v
els,
consistent
with
greater
scoring
ambiguity
near
performance
thresholds.
A
one-w
ay
ANO
V
A
on
absolute
error
scores
re
v
ealed
a
signicant
ef
fect
of
student
procienc
y
group
(
F
(2
,
405)
=
17
.
21
,
p
<
.
001
,
η
2
=
.
08
).
T
uk
e
y
HSD
post-hoc
tests
indicated
that
Group
A
dif
fered
signicantly
from
both
Groups
B
and
C
(
p
<
.
001
),
whereas
Groups
B
and
C
did
not
dif
fer
signicantly
(
p
≈
.
99
).
A
chi-square
test
on
EM
rates
also
conrmed
signicant
group
dif
ferences
(
χ
2
(2)
=
37
.
99
,
p
<
.
001
).
These
ndings
suggest
that
the
AI
system
performs
more
consistent
ly
when
student
responses
are
well-structured
and
aligned
with
e
xpected
solution
approaches.
Responses
with
incomplete
reasoning,
uncon
v
entional
methods,
or
substantial
errors
present
great
er
challenges
for
automated
scoring,
reecting
the
inherent
dif
culty
of
interpreting
ambiguous
mathematical
w
ork
[9],
[29].
Group
A
(High)
Group
B
(Medium)
Group
C
(Lo
w)
0
20
40
60
80
100
99
.
3
75
.
0
75
.
0
98
.
9
64
.
5
61
.
5
99
.
6
76
.
0
65
.
5
V
alue
(%
or
×
100)
EM
(%)
Cohen’
s
κ
(
×
100)
QWK
(
×
100)
Figure
4.
AI–human
scoring
agreement
by
student
procienc
y
group
3.4.
Qualitati
v
e
err
or
analysis
T
o
identify
the
sources
of
AI–human
scoring
dis
crepancies,
we
conducted
a
qualitati
v
e
analysis
of
responses
where
AI
scores
di
v
er
ged
from
of
cial
Colle
ge
Board
scores.
This
analysis
focused
on
Group
B
(medium-performing)
and
Group
C
(lo
w-performing)
responses,
where
disagreement
rates
were
highest.
T
w
o
recurring
error
patterns
emer
ged
as
the
primary
sources
of
these
scoring
discrepancies.
3.4.1.
Outlier
identication
Across
all
2,856
grading
trials,
v
e
cases
sho
wed
lar
ge
discrepancies
(3–4
points)
between
AI
and
human
scores
in
at
least
one
trial.
One
of
these—a
Group
B
response
that
matched
the
of
cial
score
in
six
of
se
v
en
trials—represented
an
isolated
stochastic
de
viation
rather
than
a
systematic
scoring
error
and
w
as
e
xcluded
from
further
qualitati
v
e
analysis.
The
remaining
four
cases
in
v
olv
ed
three
distinct
items,
as
one
item
appeared
i
n
both
t
h
e
AB
and
BC
e
xaminat
ions.
All
four
occurred
i
n
lo
w-performing
responses
(Group
C)
that
recei
v
ed
a
human-assigned
score
of
0
b
ut
were
a
w
arded
points
by
the
AI
system
in
multiple
trials.
As
sho
wn
in
T
able
4,
Cases
1
and
2
in
v
olv
ed
the
same
item
(2022
AP
Question
1-d),
which
appeared
in
both
the
AB
and
BC
e
xam
inations.
Case
1
sho
wed
persistent
o
v
er
-scoring,
with
the
AI
a
w
arding
4
points
in
6
of
7
trials
despite
the
of
cial
score
of
0.
Case
2
displayed
an
alternating
pattern
(0,
4,
0,
4,
0,
4,
0),
suggesting
ambiguity
in
the
AI’
s
interpretation.
Cases
3
and
4
sho
wed
intermittent
lar
ge
de
viations,
with
the
AI
occasionally
a
w
arding
2–3
points.
3.4.2.
P
atter
n
1:
inf
ormal
cancellation
marks
(scrib
ble-outs)
In
Case
3
(2023
BC
Question
6-c),
the
student’
s
response
contained
e
xtraneous
markings,
including
scratch
w
ork
that
the
student
attempted
to
cancel
using
rough
strik
ethroughs
and
zigzag
cross-outs
rather
than
Int
J
Ev
al
&
Res
Educ,
V
ol.
15,
No.
4,
August
2026:
3215-3227
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Ev
al
&
Res
Educ
ISSN:
2252-8822
❒
3223
clean
erasure.
Human
AP
readers
recognized
these
markings
as
cancellation
indicators
and
e
xcluded
them
from
e
v
aluation,
scoring
only
the
int
ended
nal
answer
as
a
0.
In
contrast,
the
AI
system
interpreted
these
cross-out
marks
as
v
alid
mathematical
content,
occasionally
assigning
inated
scores.
This
pattern
appeared
more
frequently
among
lo
wer
-performing
students,
who
often
re
vised
their
w
ork
without
fully
erasing
earlier
attempts.
T
able
4.
AI
scores
for
outlier
cases.
Se
v
en
grading
trials
for
Group
C
responses
with
of
cial
scores
of
0
Case
T
est
Item
1st
2nd
3rd
4th
5th
6th
7th
CB
score
1
2022
AB
1-d
0
4
4
4
4
4
4
0
2
2022
BC
1-d
0
4
0
4
0
4
0
0
3
2023
BC
6-c
0
0
3
0
0
2
0
0
4
2024
BC
2-c
3
0
0
0
0
0
0
0
3.4.3.
P
atter
n
2:
notation
ambiguity
and
non-linear
spatial
lay
out
F
or
Cases
1,
2,
and
4,
qualitati
v
e
analysis
re
v
ealed
a
s
econd
error
pattern
in
v
olving
ambiguous
notation
and
non-linear
spatial
or
g
anization.
These
responses
contained
non-standard
notation
that
human
graders
interpreted
as
incorrect
b
ut
that
the
AI
system
occasionally
interpreted
more
f
a
v
orably
.
The
tw
o-
dimensional
layout
appeared
to
hinder
the
AI’
s
ability
to
reconstruct
the
logical
structure
of
the
response.
Human
graders
therefore
assigned
zero
points
despite
the
presence
of
isolated
v
alid
symbols,
whereas
the
AI
system
sometimes
credited
such
fragments
without
fully
e
v
aluating
their
role
in
the
o
v
erall
reasoning.
3.4.4.
Implications
f
or
err
or
r
emediation
The
identied
error
patterns
suggest
se
v
eral
directions
for
remediation.
First,
prompts
can
be
re
vised
to
encourage
holistic
e
v
aluation
rather
than
crediting
isolated
notation
fragments.
In
follo
w-up
prompt
tests,
adding
e
xplicit
holistic
e
v
aluation
instructions
restored
correct
0-point
scoring
for
Cases
1,
2,
and
4
without
reducing
agreement
on
well-formed
responses.
Second,
impro
v
ed
image
preprocessing
may
help
lter
visual
noise
such
as
scribble-outs
and
o
v
erwritten
w
ork
before
scoring.
These
ndings
also
suggest
that
some
errors
arise
from
the
visual
and
structural
characteristics
of
handwritten
responses
rather
than
from
mathematical
reasoning
alone.
Problematic
responses
often
contained
partially
erased
w
ork,
o
v
erwritten
notation,
or
intermediate
steps
that
remained
visually
salient
despite
not
representing
the
intended
solution.
Such
features
appear
to
increase
the
lik
elihood
that
the
model
interprets
residual
notation
as
e
vidence
for
partial
credit.
Future
impro
v
ements
in
vision-language
models
may
reduce
some
of
these
limitations,
although
this
remains
an
empirical
question
[15],
[16].
As
these
models
become
better
at
distinguishing
intended
mathematical
content
from
visual
artif
acts,
the
need
for
e
xternal
preprocessing
may
diminish.
Enhanced
reasoning
capabilities
may
also
impro
v
e
the
e
v
aluation
of
incomplete
or
uncon
v
entional
responses
by
helping
the
model
distinguish
responses
that
genuinely
merit
partial
credit
from
those
that
merely
resemble
mathematically
rele
v
ant
notation.
Another
notable
nding
is
the
directional
asymmetry
in
AI–human
score
discrepancies
acr
o
s
s
procienc
y
le
v
els.
Agreement
w
as
almost
perfect
for
Group
A,
whereas
Groups
B
and
C
sho
wed
a
greater
tendenc
y
for
the
AI
system
to
assign
scores
abo
v
e
the
human
reference.
This
pattern
is
consistent
with
a
possible
lenienc
y
ef
fect,
particularly
for
responses
containing
incomplete
reasoning,
ambiguous
notat
ion,
or
uncon
v
entional
solution
strate
gies.
Ho
we
v
er
,
this
interpretation
should
be
treated
cautiously
because
score
range
restriction
may
ha
v
e
reduced
the
visibility
of
o
v
er
-scoring
among
high-performing
responses
near
the
rubric
ceili
ng.
The
present
ndings
therefore
do
not
determine
whether
the
observ
ed
asymmetry
reec
ts
a
lo
wer
-procienc
y-specic
issue
or
a
broader
scoring
tendenc
y
mask
ed
by
ceiling
ef
fects.
3.5.
Inter
pr
etation
of
r
eliability
e
vidence
The
central
issue
in
this
study
is
whether
AI-generated
scores
are
reliable
enough
to
support
educational
interpretation,
not
merely
whether
the
y
appear
plausible
on
the
surf
ace.
In
classical
measurement
theory
,
reliability
is
a
necessary
,
though
not
suf
cient,
condition
for
v
alidity
[11].
The
results
indicate
substantial
to
almost
perfect
agreement
with
e
xpert
human
raters
and
high
score
stability
across
repeated
trials,
pro
viding
an
empirical
foundation
for
further
v
alidation.
At
the
same
time,
these
ndings
should
be
interpreted
as
reliability
e
vidence
rather
than
as
proof
of
construct
v
alidity
or
operational
readiness.
Ima
g
e-native
automated
scoring
of
handwritten
mathematical
r
esponses:
r
eliability
...
(JiEun
J
anet
Song)
Evaluation Warning : The document was created with Spire.PDF for Python.
3224
❒
ISSN:
2252-8822
In
this
study
,
image-nati
v
e
AI
scoring
is
considered
not
merely
as
a
technical
inno
v
ation
b
ut
as
part
of
the
measurement
process
itself.
In
operational
assessment
conte
xts,
an
y
scoring
mechanism—human
or
automated—must
demonstrate
adequate
consistenc
y
,
transparenc
y
,
and
interpretability
[11].
Remo
ving
OCR
preprocessing
may
reduce
a
source
of
construct-irrele
v
ant
v
ariance
introduced
by
transcription
errors
[9],
[10],
and
may
therefore
strengthen
score
consistenc
y
for
handwritten
responses.
T
reating
the
system
as
a
scoring
instrument
rather
than
as
a
producti
vity
tool
mak
es
it
possible
to
e
v
aluate
it
using
established
psychometric
criteria,
including
inter
-rater
reliability
,
score
stability
,
and
error
patterns
across
performance
le
v
els.
The
reliability
metrics
reported
here
compare
f
a
v
orably
with
prior
w
ork
on
automated
scoring
of
mathematical
responses.
K
ortem
e
yer
et
al.
[16]
reported
substantially
lo
wer
agreem
ent
using
GPT
-4-based
w
orko
ws
for
handwritten
ph
ysics
e
xam
grading,
whereas
Liu
et
al.
[18]
achie
v
ed
comparable
agreement
b
ut
relied
on
OCR
preprocessing
t
hat
introduced
additional
error
sources.
In
high-sta
k
es
assessment
conte
xts,
this
distinction
is
consequential
because
scores
may
inuence
placement,
certication,
and
instructional
decisions.
3.6.
Implications
f
or
teacher
–AI
collaboration
and
policy
The
proposed
T
A
C
frame
w
ork
is
intended
not
as
a
detailed
w
orko
w
b
ut
as
a
w
ay
of
concept
ualizing
o
v
ersight
in
AI-assisted
scoring.
In
this
frame
w
ork,
AI
serv
es
as
a
rst-pass
scoring
instrument,
whil
e
educators
retain
responsibility
for
nal
score
interpretation
through
tar
geted
re
vie
w
of
discrepant
or
lo
w-condence
cases.
This
arrangement
preserv
es
professional
judgment
while
le
v
eraging
algorithmic
consistenc
y
to
reduce
part
of
the
scoring
w
orkload
[33].
From
a
measurement
perspecti
v
e,
its
v
alue
l
ies
in
treating
AI
as
a
supervised
component
within
the
scoring
process
rather
than
as
a
subst
itute
for
human
raters.
At
the
instructional
le
v
el,
the
same
approach
may
reduce
teachers’
w
orkload
whil
e
maintaining
their
authority
o
v
er
score
re
vie
w
and
interpretation.
The
k
e
y
issue
is
therefore
not
whether
human
judgment
remains
in
v
olv
ed,
b
ut
ho
w
it
is
inte
grated
into
an
AI-assisted
scoring
process
[34],
[35].
At
present,
T
A
C
should
be
re
g
arded
as
a
conceptual
frame
w
ork
rather
than
a
tested
impl
ementation
model.
Questions
re
g
arding
teacher
acceptance,
agging
thresholds,
w
orko
w
ef
cienc
y
,
and
the
allocation
of
re
vie
w
responsibility
remain
empirical
issues
for
future
study
.
F
or
no
w
,
image-nati
v
e
AI
scoring
is
best
vie
wed
as
a
rst-pass
support
mechanism
within
a
human-supervised
quality-control
process
rather
than
as
a
replacement
for
e
xpert
raters.
The
k
e
y
question
is
not
whether
AI
scoring
is
technically
feasi
ble,
b
ut
under
what
conditions
its
educational
use
can
be
justied.
In
classroom
assessment,
consistent
AI-assisted
scoring
may
help
reduce
grading
w
orkload
while
preserving
defensible
score
interpretation
through
supervised
frame
w
orks
such
as
T
A
C.
F
or
e
x
a
mination
boards,
image-nati
v
e
scoring
may
pro
vide
a
more
scalable
approach
for
e
v
aluating
handwritten
mathematical
reasoning
while
a
v
oiding
some
transcription
errors
associated
with
OCR-based
w
orko
ws.
At
the
polic
y
le
v
el,
operational
use
of
AI
scoring
should
depend
on
clear
psychometric
e
vidence.
Reliability
should
be
treated
as
a
baseline
requirement
rather
than
an
assumed
consequence
of
technological
inno
v
ation.
3.7.
Limitations
and
futur
e
dir
ections
Se
v
eral
limitations
should
be
ackno
wledged.
First,
the
study
relied
on
publicly
a
v
ailable
AP
sample
responses
selected
for
instructional
purposes,
which
may
not
fully
represent
the
di
v
ersity
of
handwriting
styles,
response
patterns,
and
error
types
found
in
li
v
e
e
xamination
settings.
The
absence
of
e
xaminee-le
v
el
information
also
limits
inferences
about
within-student
consistenc
y
across
responses.
Second,
the
reported
reliability
es
timates
are
contingent
on
the
stability
of
the
underlying
AI
model
v
ersion.
Because
model
updates
may
alter
scoring
beha
vior
e
v
en
under
the
same
rubric
and
prompting
structure,
future
research
should
e
xamine
score
comparability
across
model
v
ersions,
particularly
in
high-stak
es
conte
xts.
Third,
the
almost
perfect
agreement
observ
ed
for
high-performing
responses
should
be
interpreted
cautiously
because
score
range
restriction
may
reduce
the
visibility
of
scoring
err
o
r
s.
In
addition,
treating
all
one-point
discrepancies
as
equi
v
alent
may
obscure
dif
ferences
in
se
v
erity
across
sub-items
with
dif
ferent
point
v
alues.
Finally
,
this
study
pro
vides
reliability
e
vidence
rather
than
full
v
alidation.
Future
research
should
therefore
e
xamine
construct
alignment,
response-process
e
vidence,
f
airness
across
student
groups,
and
the
consequences
of
score
use
in
practice.
4.
CONCLUSION
This
study
e
xamined
whether
an
image-nati
v
e
AI
scoring
system
based
on
a
multimodal
reasoning
model
can
function
as
a
reliable
measurement
inst
rument
for
handwritten
mathematical
responses.
Across
Int
J
Ev
al
&
Res
Educ,
V
ol.
15,
No.
4,
August
2026:
3215-3227
Evaluation Warning : The document was created with Spire.PDF for Python.