Inter
national
J
our
nal
of
Electrical
and
Computer
Engineering
(IJECE)
V
ol.
16,
No.
4,
August
2026,
pp.
2074
∼
2086
ISSN:
2088-8708,
DOI:
10.11591/ijece.v16i4.pp2074-2086
❒
2074
Integrating
principal
component
analysis
in
spatial-spectral
fusion
models
f
or
h
yperspectral
image
segmentation
Alexander
Calvin,
Laksmita
Rahadianti
F
aculty
of
Computer
Science,
Uni
v
ersitas
Indonesia,
Depok,
Indonesia
Article
Inf
o
Article
history:
Recei
v
ed
Dec
6,
2025
Re
vised
Apr
12,
2026
Accepted
Apr
26,
2026
K
eyw
ords:
Hyperspectral
imaging
Land-use
spatial-spectral
fusion
Principal
component
analysis
Semantic
se
gmentation
U
A
V
remote
sensing
ABSTRA
CT
Hyperspectral
imaging
(HSI)
from
unmanned
aerial
v
ehicles
(U
A
Vs)
pro
vides
rich
spatial-spectral
data,
b
ut
its
high
dimensionality
presents
signicant
com-
putational
challenges
for
semantic
se
gmentation.
While
state-of-the-art
models
lik
e
the
transformer
-based
HSI-T
ransUnet
are
often
emplo
yed,
the
y
introduce
massi
v
e
computational
o
v
erhead.
This
study
adapts
a
lightweight,
dual-tunnel
deep
con
v
olutional
neural
netw
ork
(DCNN)
frame
w
ork
for
land-use
se
gmenta-
tion
on
h
yperspectral
images
by
inte
grating
PCA-based
spatial
reduction
in
the
spatial
branch,
and
benchmarks
it
on
the
U
A
V
-HSI-Crop
dataset
ag
ainst
HSI-
T
ransUnet.
F
or
further
analysis,
an
ablation
study
compares
principal
compo-
nent
analysis
(PCA)
and
local
similarity
projection
(LSP)
as
spatial
feature
e
x-
tractors.
The
results
demonstrate
a
signicant
performance
and
ef
cienc
y
adv
an-
tage.
Our
proposed
PCA-based
model
(271.1K
parameters)
obtained
a
Kappa
(
κ
)
of
0.8582,
o
v
erall
accurac
y
(O
A)
of
0.8800,
and
a
v
erage
accurac
y
(AA)
of
0.4918,
outperforming
the
LSP-based
model
by
0.65%
in
κ
,
0.51%
in
O
A,
and
2.16%
in
AA
and
the
HSI-T
ransUnet
baseline
by
2.35%
in
κ
,
1.95%
in
O
A,
and
8.10%
in
AA.
On
our
e
xperimental
setup,
this
res
ult
w
as
achie
v
ed
with
a
152.7-fold
reduction
in
model
size,
a
14.2-fold
decrease
in
training
time,
and
a
4.6-fold
speedup
in
inference
relati
v
e
to
the
reported
HSI-T
ransUnet
baseline.
These
ndings
sho
w
that
the
PCA-based
dual-tunnel
DCNN
pro
vides
a
f
a
v
or
-
able
trade-of
f
between
class-balanced
accurac
y
and
computational
ef
cienc
y
for
this
HSI
se
gmentation
task.
This
is
an
open
access
article
under
the
CC
BY
-SA
license
.
Corresponding
A
uthor:
Ale
xander
Calvin
F
aculty
of
Computer
Science,
Uni
v
ersitas
Indonesia
UI
Depok
Campus,
Depok,
Indonesia
Email:
ale
xander
.calvin@ui.ac.id
1.
INTR
ODUCTION
Hyperspectral
imaging
(HSI)
in
remote
sensing
pro
vides
the
capability
to
record
high-resolution
spec-
tral
information
in
a
lar
ge
number
of
narro
w
and
contiguous
spectral
bands.
The
use
of
unmanned
aerial
v
ehi-
cles
(U
A
Vs)
with
h
yperspectral
imaging
technology
has
a
lso
enabled
ne-scale
observ
ation
of
surf
ace
materials
for
lar
ge
re
gions
[1].
Unlik
e
re
gular
RGB
images,
HSI
pro
vides
recorded
reectance
information
in
hundreds
of
narro
w
spectral
bands
to
acquire
infor
mation
related
to
material
composition
[2].
This
spectral
detail
allo
ws
for
precise
identication
of
surf
ace
materials,
supporting
applications
such
as
land
use
mapping
[3],
en
vironmental
monitoring
[4],
precision
agriculture
[5],
and
urban
analysis
[6].
Ho
we
v
er
,
the
se
gmentation
of
h
yperspectral
aerial
imagery
remains
a
non-tri
vial
task
due
to
the
high
dimensionality
of
the
data,
redundanc
y
among
spectral
bands,
limited
a
v
ailability
of
annotated
samples,
and
the
frequent
presence
of
noise
and
en
vironmental
artif
acts
[7].
J
ournal
homepage:
http://ijece
.iaescor
e
.com
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Elec
&
Comp
Eng
ISSN:
2088-8708
❒
2075
The
high-dimensionality
of
HSI
data
poses
considerable
computational
challenges.
A
lar
ge
number
of
spectral
bands
causes
high
computational
comple
xity
with
the
potential
for
model
o
v
ertting.
This
is
particu-
larly
problematic
gi
v
en
the
scarcity
of
labeled
data
in
the
conte
xt
of
agriculture
[8].
T
o
address
such
problems,
se
v
eral
methods
of
dimensionality
reduction
and
feature
e
xtraction
ha
v
e
been
proposed
to
eliminate
redundant
information
with
the
intention
of
retaining
spectral-spatial
discriminati
v
e
features.
These
include
methods
such
as
principal
component
analysis
(PCA)
[9]
and
linear
discriminant
analysis
(LD
A)
[10].
T
raditional
machine
learning
(ML)
techniques
initially
impro
v
ed
classication
accurac
y
by
reduc-
ing
features
dimensionality
.
Ho
we
v
er
,
these
methods
relied
hea
vily
on
manual
feature
engineering
and
often
struggled
to
model
the
comple
x,
strongly
nonlinear
spectral
relationships
inherent
in
h
yperspectral
imagery
[11]-[13].
While
deep
learning
(
DL)
techniques
using
con
v
olutional
neural
netw
orks
(CNNs)
o
v
ercame
earlier
limitations
by
automating
feature
learning
[14],
these
spectrally
focused
methods
often
ignore
spatial
conte
xt
[15],
leading
to
nois
y
salt-and-pepper
classication
maps
[16].
T
o
address
these
limitati
ons,
recent
literature
emphasizes
spatial-spectral
feature
e
xtraction
methods
[17],
e
xploiting
high-dimensional
spectral
features
and
neighboring
spatial
information
to
balance
spectral
redundanc
y
with
spatial
coherence.
Recent
studies
ha
v
e
further
rened
this
frame
w
ork
by
combining
spatial
re
gularization
and
ltering
processes,
including
the
use
of
guided
or
edge-a
w
are
ltering
[18],
[19],
te
xture-based
fusion
with
Gabor
lters
[20],
and
morphological
or
state-space
models
to
ef
ciently
capture
spatial-spectral
dependencies
[21].
It
is
important
to
note
that
these
methods
do
not
solely
focus
on
DL-based
pipelines
,
b
ut
also
on
inte-
gration
or
comparison
with
traditional
ML
methods.
T
raditional
ML
methods
generally
struggle
to
model
com-
ple
x
spatial-spectral
relationships
in
h
yperspectral
data,
perform
poorly
on
patch-based
schemes,
and
f
ace
high
computational
demands
for
full-scene
processing.
In
comparison,
patch-wise
CNN
approaches
e
xtract
local
features
from
o
v
erlapping
patches
centered
on
each
pix
el,
capturing
ne
spatial-spectral
patterns
b
ut
introduc-
ing
redundant
computation
and
high
training
costs
[22].
Meanwhile,
U-Net-style
architectures
perform
pix
el-
wise
prediction
on
non-o
v
erlapping
spectral–spatial
patches,
reducing
redundanc
y
while
maintaining
broader
conte
xtual
a
w
areness,
achie
ving
a
more
ef
cient
balance
between
accurac
y
and
comple
xity
[23].
This
article
focuses
on
e
v
aluating
spatial-spectral
feature
e
xtraction
approaches
that
inte
grate
spatial
conte
xt
within
deep
learning
frame
w
orks
implem
ented
on
crop
h
yperspectral
data,
due
to
the
scene’
s
distinct
area
cate
gories
and
te
xtural
characterist
ics
that
reect
v
ariations
in
v
e
getation
type,
gro
wth
stage,
and
canop
y
structure,
which
collecti
v
ely
inuence
the
spatial-spectral
patterns
captured
by
the
imaging
sensor
[24],
[25].
Simply
adding
spatial
cues
to
spectral
features
is
insuf
cient.
High-dimensional
HSI
data
can
cause
spatial
lters
to
be
o
v
ershado
wed
by
spectral
redundanc
y
,
while
purely
spectral
learning
often
o
v
erlooks
important
local
consistenc
y
[26].
While
r
ecent
T
ransformer
models
are
po
werful,
hea
vy
computational
costs
often
limit
their
real-w
orld
appli
cation
[27].
This
highlights
an
important
res
earch
g
ap
for
lightweight,
spatial-spectral
architectures
that
can
accurately
process
comple
x
agricultural
data
without
the
massi
v
e
o
v
erhead.
A
balanced
approach
is
therefore
needed
to
ensure
that
both
components
contrib
ute
meaningfully
.
T
o
address
this,
the
study
adopts
a
dual-tunnel
st
rate
gy
in
which
the
spatial
pathw
ay
acts
as
a
spectrally
a
w
are
lter
.
Dimensionality
reduction
is
applied
before
spatial
ltering,
allo
wing
the
model
to
reduce
redundanc
y
and
e
xtract
clean,
te
xture-ori
ented
information.
This
design
ensures
that
spa
tial
conte
xt
becomes
a
rened
e
xtension
of
the
spectral
representation,
ef
fecti
v
ely
complementing
the
deeper
spectral
features
learned
through
the
parallel
pathw
ay
.
The
main
contrib
utions
of
this
article
are
summarized
as
follo
ws.
−
Ev
aluation
of
the
spatial-spectral
DCNN
frame
w
ork:
This
study
implements
and
tests
a
spatial-spectral
feature
e
xtraction
method
inte
grated
within
a
dual-tunnel
deep
con
v
olutional
neural
netw
ork
(DCNN)
[20].
The
frame
w
ork
is
e
v
aluated
for
its
ability
to
enhance
se
gmentation
accurac
y
on
comple
x
U
A
V
h
yperspectral
crop
data
by
e
xplicitly
separating
and
fusing
spectral
and
spatial
process
path.
−
Comparati
v
e
study
of
dimensionality
reduction
methods:
Ablation
study
is
conducted
to
e
v
aluate
dimen-
sionality
reduction
techniques
within
the
spatial
lter
tunnel.
Specically
,
the
study
compares
the
ef
fec-
ti
v
eness
of
local
similarity
projection
(LSP)
[20]
and
PCA
[9]
to
determine
which
method
best
preserv
es
discriminati
v
e
spatial-spectral
features.
The
rest
of
this
article
is
or
g
anized
as
follo
ws.
Section
2
introduces
the
e
v
aluated
spatial-spect
ral
methods,
the
dataset
description,
parameter
settings,
e
v
aluation
metrics,
and
the
proposed
e
xperimental
frame-
w
ork.
Section
3
discusses
the
results
and
comparati
v
e
analysis
with
e
xisting
benchmarks.
Finally
,
section
4
concludes
the
study
and
outlines
future
directions.
Inte
gr
ating
principal
component
analysis
in
spatial-spectr
al
...
(Ale
xander
Calvin)
Evaluation Warning : The document was created with Spire.PDF for Python.
2076
❒
ISSN:
2088-8708
2.
METHOD
This
study
adopts
an
e
xperimental
design
to
e
v
aluate
the
ef
fect
of
spatial-spectral
feature
e
xtraction
for
crop
se
gmentation
of
U
A
V
-based
h
yperspectral
images.
The
o
v
erall
w
orko
w
follo
ws
the
general
spatial-
spectral
fusion
frame
w
ork
inspired
by
recent
local
similarity-based
methods,
b
ut
adapts
it
for
semantic
se
gmen-
tation
instead
of
patch-based
classication
[20].
The
processing
pipeline
inte
grates
dimensionality
reduction,
spatial
ltering,
feature
fusion,
and
semantic
se
gmentation
within
a
unied
structure.
2.1.
Ov
erall
framew
ork
of
the
spatial-spectral
pipeline
The
o
v
erall
frame
w
ork
of
the
method
is
illustrated
in
Figure
1.
The
frame
w
ork
is
di
vided
into
tw
o
parts,
namely
the
spatial
tunnel
and
the
spectral
tunnel.
The
spatial
tunnel
e
xtracts
structural
and
te
xture
information
using
LSP
and
Gabor
ltering,
which
together
enhance
spatial
coherence
and
edge
representa-
tion.
The
spectral
tunnel
emplo
ys
a
DCNN
to
learn
discrimi
nati
v
e
spectral
characteristics
directly
from
pix
el-
wise
reectance
v
ectors.
The
outputs
from
both
tunnels
are
fused
via
concatenation
to
form
a
joint
repre-
sentation
which
is
processed
by
a
dual-optimized
clas
sier
head
to
generate
the
nal
pix
el-wise
se
gmentation
[20].
Figure
1.
The
dual-tunnel
DCNN
architecture
for
spatial-spectral
feature
learning.
A
ra
w
h
yperspectral
patch
(
96
×
96
×
200
)
is
processed
through
parallel
branches:
a
spatial
tunnel
(dimensionality
reduction
and
Gabor
ltering)
and
a
spectral
tunnel
(2-D
con
v
olutions).
The
e
xtracted
spatial
and
spectral
features
are
concatenated
into
a
joint
representation
and
fed
into
a
dual
classier
to
output
the
nal
96
×
96
predicted
patch
2.1.1.
Local
similarity
and
Gabor
-based
spatial
featur
e
extraction
The
spatial
tunnel,
sho
wn
in
white
in
Figure
1,
combines
LSP
and
2-D
Gabor
ltering
to
deri
v
e
spatially
coherent
and
te
xture-sensiti
v
e
representations
from
h
yperspectral
data.
The
method
addresses
the
challenges
posed
by
the
high
dimensionality
and
limited
labeled
samples
of
h
yperspectral
data
by
rst
apply-
ing
LSP
to
reduce
dimensionality
while
preserving
neighborhood
relationships.
Spatial
features
are
e
xtracted
through
2-D
Gabor
ltering,
which
captures
edge-
and
frequenc
y-oriented
te
xtures,
while
spectral
features
are
learned
using
a
CNN
applied
to
the
original
h
yperspectral
cube.
These
tw
o
feature
streams
are
then
fused
and
input
into
a
deeper
CNN,
follo
wed
by
classication
using
a
dual-optimization
classier
.
LSP
focuses
on
projecting
the
local
similarity
of
HSI
data
by
ensuring
that
spectrally
similar
neigh-
boring
pix
els
remain
close
in
the
reduced
feature
space
[20].
This
mak
es
LSP
particularly
suitable
for
HSIs,
where
class
distrib
utions
are
often
non-Gaussian,
multi-modal,
and
spatially
dependent.
LSP
operates
under
the
assumption
that
adjacent
pix
els,
especially
those
belonging
to
the
same
material
class,
share
strong
spectral
similarity
and
should
therefore
be
represented
closely
in
the
projected
subspace.
F
ormally
,
gi
v
en
a
set
of
train-
ing
samples
{
x
i
∈
R
d
}
n
i
=1
with
corresponding
class
labels
y
i
∈
{
1
,
.
.
.
,
c
}
,
where
c
represents
the
number
of
classes.
Int
J
Elec
&
Comp
Eng,
V
ol.
16,
No.
4,
August
2026:
2074-2086
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Elec
&
Comp
Eng
ISSN:
2088-8708
❒
2077
LSP
denes
the
similarity
weight
between
tw
o
samples
x
i
and
x
j
as
in
(1).
A
i,j
=
exp
−
∥
x
i
−
x
j
∥
2
γ
i
·
γ
j
(1)
where
γ
i
and
γ
j
denote
the
local
scaling
f
actors
determined
from
distances
to
their
respecti
v
e
m
-nearest
neigh-
bors.
This
adapti
v
e
k
ernel
ensures
that
local
neighborhoods
are
preserv
ed,
e
v
en
in
re
gions
with
v
arying
class.
LSP
uses
local
inter
-class
and
intra-class
scatter
matrices
to
construct
a
transformation
matrix
T
LSP
that
si-
multaneously
maxi
mizes
class
separability
and
minimizes
intra-class
v
ariance.
This
is
achie
v
ed
through
the
optimization
of
the
Fisher
criterion
in
(2).
T
LSP
=
arg
max
T
tr
h
T
⊤
L
l
w
T
−
1
T
⊤
L
lb
T
i
(2)
where
L
lb
and
L
l
w
represent
the
local
inter
-
and
intra-class
scatter
matrices.
These
matrices
are
constructed
using
the
similarity
weights
A
i,j
and
are
dened
as
in
(3)
and
(4)
respecti
v
ely
.
L
lb
=
1
2
X
i,j
W
i,j
(
x
i
−
x
j
)(
x
i
−
x
j
)
⊤
(3)
L
l
w
=
1
2
X
i,j
W
′
i,j
(
x
i
−
x
j
)(
x
i
−
x
j
)
⊤
(4)
By
preserving
these
local
relationships,
LSP
ensures
that
adjacent
pix
els
of
the
same
cl
ass
are
tightly
clustered,
while
those
of
dif
ferent
classes
remain
wel
l-separated
in
the
reduced
space.
This
struc-
ture
leads
to
lo
wer
classication
error
,
especially
in
challenging
scenarios
with
o
v
erlapping
classes
and
spatial
noise.
Meanwhile,
the
Gabor
lter
serv
es
as
a
spatial
tunnel
that
encodes
directional
te
xture
patterns
around
each
pix
el,
pro
viding
spatial
information
that
enhances
the
spectral
tunnel
deri
v
ed
from
the
original
h
yperspec-
tral
cube
[20].
The
2-D
Gabor
transform
function
is
dened
as
a
sinusoidal
w
a
v
e
modulated
by
a
Gaussian
en
v
elope,
e
xpressed
as
(5).
g
(
x,
y
;
.
.
.
)
=
exp
−
x
′
2
+
γ
2
y
′
2
2
σ
2
exp
j
2
π
x
′
ϑ
+
ϕ
(5)
where
the
rotated
coordinates
x
′
and
y
′
are
dened
in
(6).
x
′
=
x
cos
θ
+
y
sin
θ
y
′
=
y
cos
θ
−
x
sin
θ
(6)
In
this
formulation,
(
x,
y
)
represent
the
horizontal
and
v
ertical
spatial
coordinates,
and
θ
denotes
the
orientation
of
the
lter
in
radians.
The
parameter
ϑ
controls
the
w
a
v
elength
of
the
sinusoidal
f
actor
,
while
ϕ
species
the
phase
of
fset.
Additionally
,
σ
denes
the
standard
de
viation
of
the
Gaussian
en
v
elope,
and
γ
represents
the
spatial
aspect
ratio
of
the
Gaussian
function.
This
conguration
allo
ws
the
Gabor
lter
to
selecti
v
ely
enhance
spatial
structures
aligned
with
a
specic
direction
and
scale,
ef
fecti
v
ely
encoding
local
te
xtures
and
edges.
By
applying
this
ltering
on
the
LSP-reduced
data,
the
model
preserv
es
local
feature
consistenc
y
while
reducing
redundanc
y
across
bands.
The
resulting
Gabor
-ltered
output
is
passed
into
the
spatial
tunnel
for
fusion
with
spectral
features.
2.1.2.
PCA-based
spatial
r
eduction
T
o
e
v
aluate
a
simpler
alternati
v
e
to
the
original
LSP
module
[20],
a
PCA-based
v
ariant
is
introduced
in
the
spatial
tunnel.
In
this
v
ariant,
PCA
replaces
only
the
dimensionality-reduction
stage
[28],
while
the
subsequent
Gabor
ltering,
feature
fusion,
and
se
gmentation
components
remain
unchanged.
Gi
v
en
an
input
ra
w
HSI
with
B
spectral
bands,
PCA
is
rst
applied
as
a
pre-processing
step
to
project
the
ra
w
HSI
into
k
principal
components.
The
s
ame
Gabor
-based
spatial
ltering
module
then
processes
the
resulting
principal
component
images
to
form
the
spatial
feature
representation,
which
is
concatenated
with
the
spectral
features
Inte
gr
ating
principal
component
analysis
in
spatial-spectr
al
...
(Ale
xander
Calvin)
Evaluation Warning : The document was created with Spire.PDF for Python.
2078
❒
ISSN:
2088-8708
from
the
parallel
DCNN
branch
before
pix
el-wise
classication.
The
procedural
o
w
for
this
PCA-based
spatial-spectral
classication
approach
is
summarized
in
Algorithm
1.
Algorithm
1
.
PCA-Based
Spatial-Spectral
Classication
Requir
e:
Ra
w
HSI
X
with
B
spectral
bands,
tar
get
components
k
Ensur
e:
Pix
el-wise
classication
map
M
1:
//
1.
Spatial
T
unnel
Pr
ocessing
2:
X
pca
←
PCA
(
X
,
k
)
{
Project
B
bands
into
k
principal
components
}
3:
F
spatial
←
GaborFilter
(
X
pca
)
{
Extract
spatial
features
}
4:
//
2.
Spectral
T
unnel
Pr
ocessing
(in
parallel)
5:
F
spectr
al
←
DCNN
(
X
)
{
Extract
deep
spectral
features
}
6:
//
3.
F
eatur
e
Fusion
&
Classication
7:
F
f
used
←
Concatenate
(
F
spatial
,
F
spectr
al
)
{
Combine
spatial
and
spectral
features
}
8:
M
←
Classify
(
F
f
used
)
{
Perform
pix
el-wise
classication
}
9:
r
etur
n
M
2.1.3.
Spectral
featur
e
extraction
tunnel
T
o
impro
v
e
the
utilization
of
spectral
data
features,
a
spectral
feature
e
xtraction
tunnel,
sho
wn
in
blue
in
Figure
1,
is
designed
to
run
in
parallel
with
the
spatial
process
ing.
This
tunnel
focuses
on
the
central
pix
el
Z
ij
at
position
P
ij
and
its
immediate
neighborhood,
taking
a
3-D
patch
of
size
k
×
k
×
B
as
input,
where
B
is
the
number
of
spectral
bands
and
k
=
2
r
+
1
is
determined
by
the
radius
r
.
The
e
xtraction
process
in
v
olv
es
a
series
of
2-D
con
v
olutions
and
non-linear
transformations
designed
to
compress
the
spectral
depth
while
preserving
discriminati
v
e
signatures.
The
original
data
Z
ij
is
rst
processed
by
a
2-D
con
v
olutional
layer
follo
wed
by
batch
normalization
(BN)
to
standardize
features.
This
is
follo
wed
by
a
rectied
linear
unit
(R
eLU)
acti
v
ation
function.
A
second
con
v
olutional
layer
further
renes
these
features.
2.1.4.
DCNN
f
or
spectral-spatial
featur
e
extraction
The
DCNN
architecture
[20],
sho
wn
in
purple
in
Figure
1,
follo
ws
an
encoder
-decoder
structure
de-
signed
to
e
xtract
features
and
subsequently
restore
spatial
resolution.
The
encoding
path
be
gins
with
the
C1
layer
,
which
applies
a
5
×
5
con
v
olution
k
ernel
to
the
input
image
patch,
producing
50
feature
maps.
These
feature
maps
are
do
wnsampled
by
the
S1
layer
using
a
2
×
2
sampling
windo
w
for
maximum
pooling.
F
ollo
w-
ing
this,
t
h
e
C2
layer
applies
another
5
×
5
con
v
olution
k
ernel
to
the
output
of
S1,
ag
ain
resulting
in
50
feature
maps.
Finally
,
the
S2
layer
performs
a
second
do
wnsampling
operation
with
a
2
×
2
maximum
pooling
windo
w
,
outputting
50
feature
maps
with
encoded
semantic
information.
T
o
generate
the
dense
pix
el-wise
classication
map,
the
decoding
path
emplo
ys
upsampling
layers
via
transposed
con
v
olutions.
The
encoded
features
are
rst
upsampled
by
a
transposed
con
v
olutional
layer
using
a
3
×
3
k
ernel
with
a
stride
of
2,
doubling
the
spatial
di-
mension.
A
subsequent
transposed
con
v
olution
layer
res
tores
the
feature
maps
to
the
original
input
resolution,
mitig
ating
the
loss
of
spatial
details
caused
by
the
pooling
operations.
2.1.5.
Dual-optimized
classier
f
or
Pixel-wise
segmentation
The
features
e
xtracted
by
the
DCNN
are
then
concatenated
to
form
t
he
input
for
the
crop
se
gmentation
step
through
the
dual
classier
step.
This
step
is
sho
wn
in
green
in
Figure
1.
T
o
adapt
the
dual-classi
er
concept
for
pix
el-le
v
el
se
gmentation
output,
a
stack
ed
feature
renement
strate
gy
is
emplo
yed:
−
First
layer:
feature
reconstruction.
The
DCNN
output
feature
maps
are
reconstructed
for
each
pix
el,
forming
a
rened
spectral–spatial
representation.
−
Second
layer:
pix
el-wise
classication
head.
The
reconstructed
features
are
concatenated
with
the
original
DCNN
outputs
and
passed
through
a
1×1
con
v
olutional
layer
with
softmax
acti
v
ation
to
generate
per
-pix
el
class
probabilities.
This
ef
fecti
v
ely
e
xpands
the
feature
representation
and
impro
v
es
se
gmentation
accurac
y
,
especially
for
minority
classes.
This
approach
allo
ws
the
netw
ork
to
maintain
the
adv
anta
ges
of
dual
optimization,
including
enhanced
feature
representation
and
impro
v
ed
discrimination,
while
producing
dense
pix
el-le
v
el
outputs
suitable
for
se
gmenta-
tion
rather
than
center
-patch
labels
used
in
classical
h
yperspectral
classication.
Int
J
Elec
&
Comp
Eng,
V
ol.
16,
No.
4,
August
2026:
2074-2086
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Elec
&
Comp
Eng
ISSN:
2088-8708
❒
2079
2.2.
Dataset
The
U
A
V
-HSI-Crop
dataset
is
a
publicly
a
v
ailable
datase
t
which
s
erv
es
as
a
standardized
be
n
c
hmark
for
e
v
aluating
h
yperspectral
se
gmentation
methods
proposed
by
Niu
et
al.
[23].
The
dataset
w
as
collected
in
Shenzhou
City
,
China.
The
dataset
consists
of
tw
o
scene
s:
Area
A
(
864
×
1618
pix
els)
and
Area
B
(
2332
×
959
pix
els),
captured
using
a
Pika
L
sensor
with
200
spectral
bands
(400–1000
nm)
and
a
0.1
m
ground
sampling
distance
(GSD).
As
illustrated
in
Figure
2,
the
dataset
reects
a
comple
x
real-w
orld
agricultural
scenario
containing
27
crop
cate
gories
(
e
.g
.
,
corn,
cabbage)
planted
in
irre
gular
,
smallholder
plots.
This
spatial
fragmentation
and
high
spectral
similarity
among
classes
pro
vide
a
rigorous
en
vironment
for
benchmarking
spatial
feature
e
xtraction
ag
ainst
models
lik
e
HSI-T
ransUNet.