IAES
Inter
national
J
our
nal
of
Robotics
and
A
utomation
(IJRA)
V
ol.
15,
No.
3,
September
2026,
pp.
544
∼
552
ISSN:
2722-2586,
DOI:
10.11591/ijra.v15i3.pp544-552
❒
544
Spatial-channel
r
econstruction
f
or
efcient
multiscale
attention
in
r
obotic
object
detection
Mohammed
Maiza
1
,
Chahira
Cherif
2
,
Samira
Chouraqui
3
,
Abdelmalik
T
aleb-Ahmed
4
1
LITIO
Laboratory
,
F
aculty
of
Exact
and
Applied
Sciences,
Uni
v
ersity
of
Oran
1
Ahmed
Ben
Bella,
Oran,
Algeria
2
RIIR
Laboratory
,
F
aculty
of
Medicine,
Uni
v
ersity
of
Oran
1
Ahmed
Ben
Bella,
Oran,
Algeria
3
F
aculty
of
Mathematics
and
Computer
Science,
Uni
v
ersity
of
Sciences
and
T
echnology
of
Oran,
Oran,
Algeria
4
IEMN,
Polytechnic
Uni
v
ersity
of
Hauts-de-France,
Uni
v
ersity
of
Lille,
V
alenciennes,
France
Article
Inf
o
Article
history:
Recei
v
ed
Feb
17,
2026
Re
vised
Apr
24,
2026
Accepted
May
22,
2026
K
eyw
ords:
Autonomous
robots
Lightweight
attention
Object
detection
Real-time
perception
Y
OLO
ABSTRA
CT
Real-time
object
detection
is
a
core
capability
for
autonomous
robots,
un-
manned
aerial
v
ehicles
(U
A
Vs),
and
self-dri
ving
systems
operating
in
resource-
constrained
en
vironments.
This
paper
presents
spatial-channel
enhanced
mul-
tiscale
attention
(SCEMA),
a
no
v
el
l
ightweight
attention
module
designed
to
enhance
robotic
perception
while
minimi
zing
computational
o
v
erhead
for
em-
bedded
deplo
yment.
SCEMA
emplo
ys
a
parallel
dual-branch
architecture
that
syner
gistically
combines
spatial-channel
reconstruction
with
multiscale
atten-
tion
mechanisms.
When
inte
grated
into
the
Y
OLOv8n
frame
w
ork
(3.01M
base-
line
parameters),
the
proposed
Y
OLO-SCEMA
model
achie
v
es
signicant
per
-
formance
g
ains
across
multiple
challenging
benchmarks
rele
v
ant
to
robotics
au-
tomation.
Experiments
on
an
NVIDIA
R
TX
4080
GPU
demonstrate
that
on
the
ExDark
datase
t,
Y
OLO-SCEMA
impro
v
es
mAP@50
by
7.37%
o
v
er
the
baseline
(69.07%
to
76.44%)
while
reducing
parameters
by
36.88%
(3.01M
to
1.90M)
and
computational
cost
by
8.64%
(8.1
to
7.4
GFLOPs).
Consistent
im-
pro
v
ements
are
also
observ
ed
on
V
isDrone2019
(+3.24%
mAP@50)
and
FYP
(+1.50%
mAP@50)
datasets.
Comparati
v
e
analysis
demonstrates
that
Y
OLO-
SCEMA
achie
v
es
superior
accurac
y-ef
cienc
y
trade-of
fs,
making
it
particularly
suitable
for
deplo
yment
in
lo
w-light
conditions,
dense
scenes,
and
comple
x
structural
en
vironments
for
autonomous
na
vig
ation,
robotic
surv
eillance,
and
industrial
automation
applications.
This
is
an
open
access
article
under
the
CC
BY
-SA
license
.
Corresponding
A
uthor:
Mohammed
Maiza
LITIO
Laboratory
,
F
aculty
of
Exact
and
Applied
Sciences,
Uni
v
ersity
of
Oran
1
Ahmed
Ben
Bella
Oran,
Algeria
Email:
maiza.mohammed@uni
v-oran1.dz
1.
INTR
ODUCTION
Object
detection
constitutes
a
fundamental
perception
capability
for
autonomous
robotic
s
ystems,
including
ground
robots,
unmanned
aerial
v
ehicles
(U
A
Vs),
and
self-dri
ving
v
ehicles.
These
systems
must
reli-
ably
detect
and
localize
objects
in
real-time
while
operating
under
se
v
ere
computational
and
ener
gy
constraints
inherent
to
embedded
robotic
platforms
[1],
[2].
Con
v
olutional
neural
netw
orks
(CNNs)
ha
v
e
become
the
dom-
inant
paradigm
for
visual
perception
in
robotics;
ho
we
v
er
,
modern
detectors
achi
e
v
e
high
accurac
y
at
the
cost
of
increased
depth
and
width,
signicantly
raising
computational
and
memory
requirements
that
limit
deplo
yment
on
resource-constrained
robotic
hardw
are
[3].
In
robotic
applications
such
as
autonomous
na
vig
ation,
surv
eil-
lance,
and
industrial
automation,
perception
systems
f
ace
three
critical
challenges:
i)
lo
w-light
conditions
(e.g.,
J
ournal
homepage:
http://ijr
a.iaescor
e
.com
Evaluation Warning : The document was created with Spire.PDF for Python.
IAES
Int
J
Rob
&
Autom
ISSN:
2722-2586
❒
545
nighttime
autonomous
dri
ving),
ii)
dense
object
distrib
utions
with
occlusions
(e.g.,
U
A
V
-based
cro
wd
mon-
itoring),
and
iii)
comple
x
structural
scenes
(e.g.,
manuf
acturing
en
vironments).
As
netw
ork
depth
increases,
feature
maps
e
xhibit
substantial
spatial
and
channel
redundanc
y
[4],
de
grading
computational
ef
cienc
y
.
While
attention
mechanisms
ha
v
e
been
introduced
to
emphasize
informati
v
e
features,
man
y
e
xisting
approaches
incur
non-ne
gligible
computational
o
v
erhead
or
focus
on
either
spatial
or
channel
dependencies
in
isolation.
Con-
sequently
,
designing
lightweight
attention
modules
that
jointly
e
xploit
spatial
and
channel
interactions
while
preserving
multiscale
conte
xtual
information
remains
a
critical
challenge
for
real-time
robotic
perception
[5],
[6].
Recent
Y
OLO-based
detectors
emphasize
ef
cienc
y
and
real-ti
me
performance,
yet
their
lightweight
v
ari-
ants
often
suf
fer
from
de
graded
ac
curac
y
in
challenging
robotic
en
vironments.
T
o
address
this
g
ap,
this
paper
proposes
SCEMA,
a
no
v
el
lightweight
attention
module
that
enhances
robotic
perception
without
introducing
e
xcessi
v
e
computational
cost.
This
w
ork
adv
ances
robotics
and
automation
by
enabling
lightweight,
ef
cient,
and
rob
ust
per
ception
for
autonomous
systems.
The
paper
is
or
g
anized
as
follo
ws.
Section
2
re
vie
ws
related
w
ork
on
object
detection
architectures
and
attention
mechanisms
in
r
ob
ot
ics
conte
xts.
Section
3
presents
the
SCEMA
methodology
,
in-
cluding
inte
gration
into
robotic
vision
pipelines.
Section
4
details
e
xperimental
results
across
three
benchmark
datasets,
with
analysis
framed
in
terms
of
robotic
metrics.
Finally
,
section
5
concludes
with
implications
for
autonomous
robotics
and
future
directions.
2.
RELA
TED
W
ORK
Attention
mechanisms
enhance
feature
discrimination
by
selecti
v
ely
emphasizing
informati
v
e
aspects
of
input
data,
which
is
particularly
v
aluable
for
robotic
perception
in
cluttered
or
lo
w-visi
bility
en
vironments.
Channel
attention
methods,
pioneered
by
SENet
[7],
emplo
y
global
a
v
erage
pooling
follo
wed
by
fully
con-
nected
layers
to
recalibrate
feature
responses.
Spatial
attention
approaches
compute
importance
maps
across
spatial
dimensions,
highlighting
re
gions
with
high
semantic
rele
v
ance.
Hybrid
attention
modules,
such
as
CB
AM
[8]
and
ECA
[9],
combine
spatial
and
channel
cues
sequentially
or
jointly
.
Ho
we
v
er
,
man
y
attention
modules
introduce
computational
o
v
erhead
through
parameter
-intensi
v
e
operations,
hindering
deplo
yment
on
robotic
hardw
are
with
limited
processing
capabilities.
Furthermore,
e
xisting
mechanisms
frequently
f
ail
to
ef-
ciently
model
multiscale
spatial
dependencies,
limiting
their
ef
fecti
v
eness
for
robotic
applications
in
v
olving
objects
at
v
arying
distances
and
aspect
ratios.
The
demand
for
real-time
object
detection
in
robotics
has
dri
v
en
research
into
lightweight
detectors
that
balance
computational
ef
cienc
y
and
detection
accurac
y
.
The
Y
OLO
f
amily
has
emer
ged
as
a
dominant
paradigm
for
real-time
robotic
perception
due
to
its
single-stage
design
and
ef
cient
feature
e
xtraction
backbone
[10].
Recent
v
ariants,
incl
uding
Y
OLOv8
[11],
Y
OLOv9
[12],
Y
OLOv10
[13],
Y
OLOv11
[14],
and
Y
OLOv12
[15],
ha
v
e
progres
si
v
ely
impro
v
ed
the
speed-accurac
y
trade-of
f.
Ho
we
v
er
,
aggressi
v
e
model
compression
can
signicantly
de
grade
performance
in
challenging
scenarios
characteristic
of
robotic
applications:
lo
w
lighting,
dense
object
distrib
utions,
occlusions,
and
signicant
scale
v
ariations.
De-
v
eloping
lightweight
detectors
that
maintain
rob
ust
performance
across
di
v
erse
en
vironments
remains
critical
for
autonomous
na
vig
ation,
U
A
V
surv
eillance,
and
industrial
automation.
Multi
scale
feature
representation
is
essential
for
robotic
perception,
as
real-w
orld
scenes
contain
objects
with
substantial
size
v
ariations.
Feature
p
yramid-based
architectures
[16]
construct
hierarchical
representations
through
top-do
wn
fusion
or
bidirec-
tional
information
o
w
.
Ho
we
v
er
,
these
me
thods
often
rely
on
computationally
e
xpensi
v
e
fusion
strate
gies
that
increase
inference
latenc
y
and
memory
consumption,
challenging
real-time
robotic
operation.
The
in-
te
gration
of
attention
mechanis
ms
with
multiscale
feature
p
yramids
remains
undere
xplored,
limiting
potential
syner
gistic
benets.
Ef
cient
multiscale
attention
mechanisms
that
operate
across
dif
ferent
feature
scales
while
maintaining
real-time
performance
represent
a
signicant
open
research
problem
for
robotics
[17],
particularly
for
applications
requiring
long-range
dependenc
y
modeling
and
conte
xtual
inte
gration
from
dif
ferent
recepti
v
e
elds.
3.
METHOD
The
o
v
erall
architecture
of
SCEMA
is
illustrated
in
Figure
1.
SCEMA
adopts
a
parallel
proce
ssing
strate
gy
consisting
of
tw
o
complementary
branches.
In
the
l
eft
branch,
spatial-rened
features
X
w
are
e
xtracted
using
the
spatial
renement
unit
(SR
U)
and
spatial
attention,
follo
wed
by
channel
renement
through
the
channel
renement
unit
(CR
U)
to
produce
channel-rened
features
Y
.
The
right
branch
focuses
on
learning
ef
fecti
v
e
channel
representations
by
grouping
features
without
Spatial-c
hannel
r
econstruction
for
ef
cient
multiscale
attention
in
r
obotic
object
(Mohammed
Maiza)
Evaluation Warning : The document was created with Spire.PDF for Python.
546
❒
ISSN:
2722-2586
reducing
channel
dimensions,
enhancing
pix
el-le
v
el
attention
and
rob
ustness
to
multi-scale
v
ariations.
The
architecture
of
the
Y
OLO-SCEMA
model
is
sho
wn
in
Figure
2.
Figure
1.
Architecture
of
the
proposed
SCEMA
module
Figure
2.
Ov
erall
architecture
of
the
Y
OLO-SCEMA
model
SCEMA
is
inte
grated
into
the
Y
OLOv8
backbone,
forming
a
perception
module
suitable
for
robotic
autonomy
stacks.
W
ithin
the
con
v
olutional
block
(CBS),
batch
normalization
is
replaced
by
the
feature
re-
IAES
Int
J
Rob
&
Autom,
V
ol.
15,
No.
3,
September
2026:
544-552
Evaluation Warning : The document was created with Spire.PDF for Python.
IAES
Int
J
Rob
&
Autom
ISSN:
2722-2586
❒
547
nement
module
and
the
cross-spatial
learning
module.
In
the
bottleneck
structure,
the
second
con
v
olutional
layer
is
substituted
with
the
Feature
Renement
Module,
forming
Bottleneck-FR.
The
con
v
olutional
blocks
in
the
spatial
p
yramid
pooling
f
ast
(SPPF)
module
are
replaced
with
CBS-FR
and
CBS-CSL.
Through
SCEMA
inte
gration,
spatial
and
channel
redundancies
are
ef
fecti
v
ely
reduced,
enabling
enhanced
feature
representation
for
comple
x
robotic
scenes.
3.1.
F
eatur
e
r
enement
module
The
feature
renement
module
consists
of
a
spatial
reconstruction
unit
(SR
U),
spatial
attention
mod-
ule,
channel
reconstruction
unit
(CR
U),
and
channel
attention
module.
As
illustrated
in
Figure
3,
the
SR
U
performs
feature
separation
and
reconstruction
by
di
viding
the
input
feature
map
into
informati
v
e
and
less-
informati
v
e
components
using
a
scaling
f
actor
deri
v
ed
from
group
normalization
(GN)
[18].
Figure
3.
Architecture
of
the
SR
U
Gi
v
en
an
intermediate
feature
map
X
∈
R
N
×
C
×
H
×
W
,
GN
is
applied
as,
X
out
=
GN
(
X
)
=
γ
X
−
µ
√
σ
2
+
ε
+
β
,
(1)
where
µ
and
σ
denote
mean
and
standard
de
viation,
ε
ensures
numerical
stability
,
and
γ
,
β
are
learna
b
l
e
param-
eters.
Normalized
features
are
mapped
through
a
sigmoid-based
g
ating
mechanism
to
generate
information-
a
w
are
weights:
W
=
Gate
(
Sigmoid
(
GN
(
X
)))
.
(2)
W
eights
abo
v
e
a
predened
threshold
form
the
informati
v
e
mask
W
1
,
while
remaining
weights
constitute
the
non-informati
v
e
mask
W
2
.
The
input
feature
map
is
decomposed
as
X
W
1
=
W
1
⊗
X
and
X
W
2
=
W
2
⊗
X
,
representing
informati
v
e
and
redundant
features,
respecti
v
ely
.
Cross-reconstruction
produces
spatially
rened
feature
map
X
W
.
As
sho
wn
in
Figure
4,
the
CR
U
further
renes
X
W
by
reducing
channel
redundanc
y
through
splitting,
transformation,
and
fusion
operations
with
compression
ratio
r
=
2
and
group
size
g
=
2
,
producing
nal
channel-rened
features
Y
.
Figure
4.
Architecture
of
the
CR
U
3.2.
Cr
oss
spatial
lear
ning
module
The
CSL
module
emplo
ys
parallel
sub-structure
processing.
Initially
,
it
utilizes
1
×
1
con
v
olution
as
a
comm
on
element.
Subsequently
,
to
inte
grate
spatial
information
across
multiple
scales,
it
positions
3
×
3
Spatial-c
hannel
r
econstruction
for
ef
cient
multiscale
attention
in
r
obotic
object
(Mohammed
Maiza)
Evaluation Warning : The document was created with Spire.PDF for Python.
548
❒
ISSN:
2722-2586
con
v
olution
alongside
the
1
×
1
branch.
The
CSL
module
se
gments
input
feature
map
X
∈
R
C
×
H
×
W
into
three
distinct
sub-feature
groups,
each
focusing
on
distinct
semantics.
T
w
o
paths
use
1
×
1
con
v
olutions
with
one-
dimensional
global
a
v
erage
pooling
along
horizontal
and
v
ertical
directions,
follo
wing
Coordinate
Attention.
The
third
path
adopts
3
×
3
con
v
olution
to
capture
multi-scale
spatial
features.
Outputs
are
inte
grated
through
matrix
dot-product
operations:
A
=
Sigmoid
((
W
1
×
1
∗
X
)
⊙
(
W
3
×
3
∗
X
))
(3)
where
⊙
denotes
element-wise
multiplication
and
∗
denotes
con
v
olution.
This
cross-spatial
interaction
enables
ef
fecti
v
e
multi-scale
feature
fusion
for
robotic
perception.
3.3.
Integration
with
r
obotic
per
ception
pipelines
SCEMA
is
designed
for
seamless
inte
gration
into
robotic
vision
stacks,
includi
ng
R
OS-based
per
-
ception
systems
.
The
module
accepts
standard
tensor
inputs
(
C
×
H
×
W
)
and
outputs
rened
feature
maps
with
identical
dimensions,
enabling
drop-in
replacement
for
e
xisting
con
v
olutional
blocks.
W
ith
only
1.90M
parameters
and
7.4
GFLOPs,
SCEMA
is
suitable
for
deplo
yment
on
edge
de
vices
such
as
NVIDIA
Jetson,
Raspberry
Pi,
and
U
A
V
onboard
computers.
The
computational
ef
cienc
y
enables
real-time
operation
at
30+
FPS
on
embedded
hardw
are,
supporting
autonomous
na
vig
ation
loops.
4.
RESUL
TS
AND
DISCUSSION
This
section
presents
comprehensi
v
e
e
v
aluation
of
Y
OLO-SCEMA
across
three
challenging
robotic
scenarios:
lo
w-light
object
detection
(nighttime
autonomous
dri
ving),
dense
scene
detection
(U
A
V
cro
wd
surv
eillance),
and
comple
x
image
structure
understanding
(industrial
automation).
Experiments
were
con-
ducted
on
an
NVIDIA
R
TX
4080
GPU
with
Intel
Core
Ultra
7
155H
CPU.
Hyperparameters:
initial
learning
rate
0.01,
momentum
0.937,
weight
decay
0.0005,
input
resolution
640
×
640
,
training
epochs
300,
SGD
optimizer
with
cosine
annealing
schedul
er
.
Batch
sizes:
16
for
ExDark
and
V
isDrone2019,
8
for
FYP
.
Data
augmentation:
mosaic,
mixup,
random
horizontal
ip,
HSV
color
jittering.
Results
represent
mean
±
standard
de
viation
o
v
er
three
independent
runs.
Ev
aluation
on
the
ExDark
dataset
[19]
(5,891
training,
1,472
test
images,
12
cate
gories,
10
illumi-
nation
conditions)
simulates
nighttime
autonomous
dri
ving
and
lo
w-light
robotic
surv
eillance.
T
able
1
sho
ws
SCEMA
consistently
outperforms
Y
OLOv8n
and
recent
attention
methods
including
MCA
[20],
GAM
[21],
and
CB
AM
[8].
SCEMA
achie
v
es
mAP@50
of
76.44%
±
0.28%,
a
7.37%
impro
v
ement
o
v
er
Y
OLOv8n
(69.07%
±
0.31%),
while
reducing
parameters
by
36.88%
and
computational
cost
by
8.64%.
Compared
to
Y
OLOv12n
[15],
SCEMA
achie
v
es
0.5%
higher
mAP@50
using
0.66M
fe
wer
para
meters.
From
a
robotics
perspecti
v
e,
these
impro
v
ements
translate
to
enhanced
autonomous
v
ehicle
safety
during
nighttime
operation
and
more
reliable
lo
w-light
U
A
V
surv
eillance.
T
able
1.
Performance
comparison
on
ExDark
dataset
for
lo
w-light
robotic
perception
Model
mAP@50
(%)
mAP@95
(%)
P
arams
(M)
GFLOPs
Y
OLOv8n
[11]
69.07
±
0.31
43.15
±
0.28
3.01
8.1
MCA
[20]
72.07
±
0.29
44.36
±
0.26
1.87
7.1
GAM
[21]
73.39
±
0.33
46.80
±
0.30
3.94
9.5
CB
AM
[8]
75.58
±
0.27
48.65
±
0.25
3.07
8.2
Y
OLOv9t
[12]
74.74
±
0.30
48.08
±
0.27
1.97
7.6
Y
OLOv10n
[13]
75.47
±
0.28
48.38
±
0.26
2.69
8.2
Y
OLOv11n
[14]
75.64
±
0.29
48.73
±
0.27
2.58
6.3
Y
OLOv12n
[15]
75.94
±
0.27
49.06
±
0.25
2.56
6.3
SCEMA
76.44
±
0.28
49.38
±
0.26
1.90
7.4
Ev
aluation
on
V
isDrone2019
[22]
(6,471
training,
548
v
alidation,
1,580
test
images)
simulates
U
A
V
-
based
cro
wd
surv
eillance
and
dense
scene
understanding.
T
able
2
sho
ws
SCEMA
achie
v
es
mAP@50
of
34.98%
±
0.32%,
surpassing
Y
OLOv8n
by
3.24%
and
outperforming
PCon
v
[23],
ECA
[9],
and
CB
AM
[8]
by
3.21%,
2.59%,
and
2.24%,
respecti
v
ely
.
Compared
to
Y
OLOv12n,
SCEMA
impro
v
es
mAP@50
by
0.33%
while
re-
ducing
parameters
by
25.8%,
at
the
cost
of
a
modest
17.5%
increase
in
GFLOPs.
F
or
robotics
applications,
these
results
indicate
enhanced
capability
for
U
A
V
-based
cro
wd
monitoring,
autonomous
na
vig
ation
in
con-
gested
en
vironments,
and
reliable
detection
of
occluded
objects.
IAES
Int
J
Rob
&
Autom,
V
ol.
15,
No.
3,
September
2026:
544-552
Evaluation Warning : The document was created with Spire.PDF for Python.
IAES
Int
J
Rob
&
Autom
ISSN:
2722-2586
❒
549
T
able
2.
Performance
comparison
on
V
isDrone2019
dataset
for
dense
scene
robotic
perception
Model
mAP@50
(%)
mAP@95
(%)
P
arams
(M)
GFLOPs
Y
OLOv8n
[11]
31.74
±
0.34
18.28
±
0.30
3.01
8.1
PCon
v
[23]
31.77
±
0.33
18.32
±
0.31
1.86
7.1
ECA
[9]
32.39
±
0.31
18.53
±
0.29
2.65
8.4
CB
AM
[8]
32.74
±
0.30
18.75
±
0.28
3.07
8.2
Y
OLOv9t
[12]
34.58
±
0.29
20.15
±
0.27
1.97
7.6
Y
OLOv10n
[13]
34.30
±
0.31
20.01
±
0.28
2.69
8.2
Y
OLOv11n
[14]
34.54
±
0.30
20.24
±
0.27
2.58
6.3
Y
OLOv12n
[15]
34.65
±
0.29
20.36
±
0.26
2.56
6.3
SCEMA
34.98
±
0.32
20.58
±
0.29
1.90
7.4
Ev
aluation
on
FYP
dataset
[24]
simulates
industrial
automation
and
manuf
acturing
en
vironme
n
t
s
with
hea
vy
occlusion,
o
v
erlapping
objects,
and
comple
x
structural
conditions.
T
able
3
demonstrates
SCEMA
achie
v
es
mAP@50
of
69.51%
±
0.30%,
outperforming
Y
OLOv8n
by
1.50%
and
surpassing
EMA
[25],
MCA
[20],
and
CB
AM
[8]
by
0.45%,
0.98%,
and
0.27%,
respecti
v
ely
.
W
ith
only
1.90M
parameters
a
n
d
7.4
GFLOPs,
SCEMA
strik
es
an
ef
fecti
v
e
balance
for
industrial
robotic
applications
requiring
structural
understanding.
T
able
3.
Performance
comparison
on
FYP
dataset
for
industrial
automation
Model
mAP@50
(%)
mAP@95
(%)
P
arams
(M)
GFLOPs
Y
OLOv8n
[11]
68.01
±
0.33
51.65
±
0.30
3.01
8.1
EMA
[25]
69.06
±
0.31
53.16
±
0.28
3.02
8.2
MCA
[20]
68.53
±
0.32
53.48
±
0.29
1.87
7.1
CB
AM
[8]
69.24
±
0.30
53.24
±
0.28
3.07
8.2
Y
OLOv9t
[12]
68.58
±
0.31
53.53
±
0.29
1.97
7.6
Y
OLOv10n
[13]
68.67
±
0.32
53.61
±
0.30
2.69
8.2
Y
OLOv11n
[14]
68.84
±
0.31
53.74
±
0.28
2.58
6.3
Y
OLOv12n
[15]
68.95
±
0.30
53.83
±
0.27
2.56
6.3
SCEMA
69.51
±
0.30
54.01
±
0.28
1.90
7.4
Be
yond
standard
mAP
and
GFLOPs,
SCEMA
’
s
suitability
for
robotics
is
e
v
aluated
through
addit
ional
metrics.
Inference
latenc
y
on
NVIDIA
R
TX
4080:
SCEMA
achie
v
es
4.2
ms
per
frame
(238
FPS)
compared
to
Y
OLOv8n’
s
4.8
ms
(208
FPS),
representing
a
12.5%
speed
impro
v
ement.
Estimated
ener
gy
ef
cienc
y:
SCEMA
requires
approximately
0.38
J
per
inference
vs.
0.44
J
for
Y
OLOv8n
(13.6%
reduction),
critical
for
battery-po
wered
robots
and
U
A
Vs.
These
impro
v
ements
enable
higher
update
rates
for
robotic
control
loops
and
e
xtended
mission
duration
for
autonomous
systems.
T
able
4
analyzes
each
SCEMA
component’
s
contri-
b
ution
across
three
datasets.
F
our
v
ariants
are
e
v
aluated:
without
SCCon
v
(spatial-channel
reconstruction),
without
CSL,
without
CB
AM,
and
full
SCEMA.
T
able
4.
Ablation
study
results
across
robotic
perception
datasets
Model
SCCon
v
CSL
CB
AM
FYP
V
isDrone
ExDark
Y
OLOv8n
–
–
–
68.01
31.74
69.07
SCEMA
w/o
SCCon
v
–
✓
✓
68.69
31.95
72.95
SCEMA
w/o
CSL
✓
–
✓
68.54
32.28
73.63
SCEMA
w/o
CB
AM
✓
✓
–
68.95
32.55
74.74
Full
SCEMA
✓
✓
✓
69.51
34.98
76.44
Results
demonstrate
full
SCEMA
consistently
achie
v
es
highest
mAP@50.
Remo
ving
SCCon
v
ca
u
s
es
performance
drops
of
0.82%,
3.03%,
and
3.49%
on
FYP
,
V
isDrone,
and
ExDark,
respecti
v
ely
.
Remo
ving
CSL
leads
to
decreases
of
0.97%,
2.70%,
and
2.81%.
The
combined
impro
v
ement
o
v
er
baseline
reaches
7.37%
on
ExDark,
indicating
strong
syner
gistic
ef
fects
between
spatial-channel
reconstruction
and
cross-spatial
learning.
The
e
xperimental
results
demonstrate
SCEMA
’
s
ef
fecti
v
eness
for
robotic
perception.
Statistical
signif
-
icance
testing
using
paired
t-tests
conrms
all
impro
v
ements
are
s
ignicant
(
p
<
0
.
05
),
with
p
-v
alues
ranging
from
0.003
to
0.021.
Con
v
er
gence
analysis
re
v
eals
SCEMA
achie
v
es
st
able
training
approximately
25
epochs
earlier
than
Y
OLOv8n,
with
nal
loss
v
alues
0.18
lo
wer
,
attrib
uted
to
reduced
feature
redundanc
y
enabling
more
ef
cient
gradient
o
w
.
F
or
autonomous
v
ehicles
operating
at
night,
the
7.37%
mAP
impro
v
ement
on
ExDark
directly
addresses
safety
during
lo
w-light
operation.
F
or
drone-based
surv
eillance
in
cro
wded
en
viron-
ments,
the
3.24%
impro
v
ement
on
V
isDrone2019
enables
more
reliable
monitoring
of
occluded
objects.
F
or
Spatial-c
hannel
r
econstruction
for
ef
cient
multiscale
attention
in
r
obotic
object
(Mohammed
Maiza)
Evaluation Warning : The document was created with Spire.PDF for Python.
550
❒
ISSN:
2722-2586
industrial
robotics
requiring
structural
understanding,
the
1.50%
impro
v
ement
on
FYP
f
acilitates
better
scene
interpretation
and
na
vig
ation
in
comple
x
manuf
acturing
en
vironments.
5.
CONCLUSION
This
paper
introduced
SCEMA,
a
lightweight
and
ef
cient
multiscale
attention
fusion
module
de-
signed
to
enhance
object
detection
for
robotics
and
automation
under
comple
x
visual
conditions.
By
emplo
y-
ing
a
parallel
dual-branch
architecture,
SCEMA
reduces
spatial
and
channel
redundanc
y
while
strengthening
multiscale
feature
representation
through
cross-spatial
learning.
Inte
grated
into
Y
OLOv8,
Y
OLO-SCEMA
achie
v
es
superior
detection
accurac
y
with
signicantly
fe
wer
parameters
(1.90M
vs.
3.01M)
and
lo
wer
com-
putational
cost
(7.4
vs.
8.1
GFLOPs).
Extensi
v
e
e
v
aluations
on
ExDark,
V
isDrone2019,
and
FYP
datasets
demonstrate
consistent
outperformance
of
e
xisting
attention
methods
and
recent
Y
OLO
v
ariants,
particularly
in
lo
w-light,
dense,
and
structurally
comple
x
scenes.
The
broader
impact
lies
in
enabling
accurate
real-time
object
detection
in
resource-constrained
robotic
en
vironments,
with
applications
in
autonomous
dri
ving
(night-
time
operation
safety),
U
A
V
surv
eillance
(occluded
object
detection
in
cro
wded
en
vironments),
and
industrial
automation
(structural
underst
anding
in
comple
x
manuf
acturing
settings).
The
demonstrated
ability
to
main-
tain
high
performance
under
challenging
conditions
while
reducing
computational
requirements
represents
a
signicant
step
to
w
ard
practical
deplo
yment
of
dee
p
learning-based
perception
in
real-w
orld
robotic
systems.
Code
and
pre-trained
models
will
be
made
publicly
a
v
ailable.
FUNDING
INFORMA
TION
Author
state
no
funding
is
in
v
olv
ed.
A
UTHOR
CONTRIB
UTIONS
ST
A
TEMENT
This
journal
uses
the
C
on
t
rib
utor
Roles
T
axonomy
(CRediT)
to
recognize
indi
vidual
author
contrib
u-
tions,
reduce
authorship
disputes,
and
f
acilitate
collaboration.
Name
of
A
uthor
C
M
So
V
a
F
o
I
R
D
O
E
V
i
Su
P
Fu
Mohammed
Maiza
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
Chahira
Cherif
✓
✓
✓
✓
✓
✓
✓
Samira
Chouraqui
✓
✓
✓
✓
✓
✓
✓
Abdelmalik
T
aleb-Ahmed
✓
✓
✓
✓
✓
C
:
C
onceptualization
I
:
I
n
v
estig
ation
V
i
:
V
i
sualization
M
:
M
ethodology
R
:
R
esources
Su
:
Su
pervision
So
:
So
ftw
are
D
:
D
ata
Curation
P
:
P
roject
Administration
V
a
:
V
a
lidation
O
:
Writing
-
O
riginal
Draft
Fu
:
Fu
nding
Acquisition
F
o
:
F
o
rmal
Analysis
E
:
Writing
-
Re
vie
w
&
E
diting
CONFLICT
OF
INTEREST
ST
A
TEMENT
Authors
state
no
conict
of
interest.
D
A
T
A
A
V
AILABILITY
Deri
v
ed
data
support
ing
the
ndings
of
this
study
are
a
v
ailable
from
the
corresponding
author
upon
reasonable
request.
REFERENCES
[1]
N.
Ma,
X.
Zhang,
H.-T
.
Zheng,
and
S.
Sun,
“Shuf
eNet
V2:
Practical
guidelines
for
ef
cient
CNN
architecture
de-
sign,
”
in
Pr
oceedings
of
the
Eur
opean
Confer
ence
on
Computer
V
ision
(ECCV)
,
2018,
pp.
116–131,
doi:
10.1007/978-
3-030-01264-9
8.
IAES
Int
J
Rob
&
Autom,
V
ol.
15,
No.
3,
September
2026:
544-552
Evaluation Warning : The document was created with Spire.PDF for Python.
IAES
Int
J
Rob
&
Autom
ISSN:
2722-2586
❒
551
[2]
X.
Lu,
M.
Sug
anuma,
and
T
.
Okatani,
“SBCF
ormer:
Lightweight
netw
ork
capable
of
full-size
ImageNet
classica-
tion
at
1
FPS
on
single
board
computers,
”
in
Pr
oceedings
of
the
IEEE/CVF
W
inter
Confer
ence
on
Applications
of
Computer
V
ision
(W
A
CV)
,
2024,
pp.
1123–1133,
doi:
10.1109/W
A
CV57701.2024.00116.
[3]
J.
Redmon
and
A.
F
arhadi,
“Y
OLOv3:
An
incremental
impro
v
eme
nt,
”
arXiv:1804.02767
,
2018.
[4]
J.
Li,
Y
.
W
en,
and
L.
He,
“SCCon
v:
Spatial
and
channel
reconstruction
con
v
olution
for
feature
redundanc
y
,
”
in
Pr
oceedings
of
the
IEEE/CVF
Confer
ence
on
Computer
V
i
sion
and
P
attern
Reco
gnition
(CVPR)
,
2023,
pp.
6153–
6162,
doi:
10.1109/CVPR52729.2023.00596.
[5]
M.
Sohan,
T
.
S.
Ram,
and
C.
V
.
R.
Reddy
,
“
A
re
vie
w
on
Y
OLOv8
and
its
adv
ancements,
”
in
Pr
oceedings
of
the
Inter
-
national
Confer
ence
on
Data
Intellig
ence
and
Co
gnitive
Infor
matics
(ICDICI)
,
2024,
pp.
529–545,
doi:
10.1007/978-
981-99-7962-2
39.
[6]
Q.
Y
an,
Y
.
Feng,
C.
Zhang,
P
.
W
ang,
and
L.
Zhang,
“HVI:
A
ne
w
color
space
for
lo
w-light
image
enhancement,
”
in
Pr
oceedings
of
the
IEEE/CVF
Confer
ence
on
Computer
V
is
ion
and
P
attern
Reco
gnition
(CVPR)
,
2025,
pp.
5678–
5687,
doi:
10.1109/CVPR52734.2025.00533.
[7]
J.
Hu,
L.
Shen,
and
G.
Sun,
“Squeeze-and-e
xcitation
netw
orks,
”
in
Pr
oceedings
of
the
IEEE/CVF
Confer
ence
on
Computer
V
ision
and
P
attern
Reco
gnition
(CVPR)
,
2018,
pp.
7132–7141.
[8]
S.
W
oo,
J.
P
ark,
J.-Y
.
Lee,
and
I.
S.
Kweon,
“CB
AM:
Con
v
olutional
block
attention
module,
”
in
Pr
oceedings
of
the
Eur
opean
Confer
ence
on
Computer
V
ision
(ECCV)
,
2018,
pp.
3–19,
doi:
10.1007/978-3-030-01234-2
1.
[9]
Q.
W
ang,
B
.
W
u,
P
.
Zhu,
P
.
Li,
W
.
Zuo,
and
Q.
Hu,
“ECA-Net:
Ef
cient
channel
attention
for
deep
con
v
olutional
neural
net
w
orks,
”
in
Pr
oceedings
of
the
IEEE/CVF
Confer
ence
on
Computer
V
ision
and
P
attern
Reco
gnition
(CVPR)
,
2020,
pp.
11531–11539,
doi:
10.1109/CVPR42600.2020.01155.
[10]
J.
T
erv
en,
D.-M.
C
´
ordo
v
a-Esparza,
and
J.-A.
Romero-Gonz
´
alez,
“
A
comprehensi
v
e
re
vie
w
of
Y
OLO
architectures
in
computer
vision:
From
Y
OLOv1
to
Y
OLOv8
and
Y
OLO-N
AS,
”
Mac
hine
Learning
and
Knowledg
e
Extr
action
,
v
ol.
5,
no.
4,
pp.
1680–1716,
2023,
doi:
10.3390/mak
e5040083.
[11]
“Y
OLOv8:
A
state-of-the-art
object
detect
ion
model,
”
GitHub
,
2023.
https://roboo
w
.com/model/yolo
v8.
[12]
C.-Y
.
W
ang,
I.-H.
Y
eh,
and
H.-Y
.
M.
Liao,
“Y
OLOv9:
Learning
what
you
w
ant
to
learn
using
programmable
gra-
dient
i
nformation,
”
in
Pr
oceedings
of
the
Eur
opean
Confer
ence
on
Computer
V
ision
(ECCV)
,
ser
.
Lecture
Notes
in
Computer
Science,
v
ol.
15089,
2025,
pp.
1–18,
doi:
10.1007/978-3-031-72751-1
1.
[13]
A.
W
ang
et
al.
,
“Y
OLOv10:
Real-time
end-to-end
object
detection,
”
in
Pr
oceedings
of
the
International
Confer
ence
on
Neur
al
Information
Pr
ocessing
Systems
(NeurIPS)
,
v
ol.
38,
2024.
[14]
C.
W
ang,
X.
Song,
J.
W
ang,
and
Y
.
Zhang,
“
An
impro
v
ed
Y
OLOv11
algorithm
for
small
object
detection
in
U
A
V
images,
”
Signal,
Ima
g
e
and
V
ideo
Pr
ocessing
,
v
ol.
19,
p.
507,
2025,
doi:
10.1007/s11760-025-04157-w
.
[15]
Y
.
Qiao,
Y
.
W
ei,
and
K.
W
ang,
“F
ast-Y
OLOv12:
An
attention-guided
lightweight
netw
ork
for
real-time
steel
surf
ace
defect
detection,
”
J
ournal
of
Real-T
ime
Ima
g
e
Pr
ocessing
,
v
ol.
23,
p.
9,
2026,
doi:
10.1007/s11554-025-01807-7.
[16]
T
.-Y
.
Lin,
P
.
Doll
´
ar
,
R.
Girshick,
K.
He,
B.
Hariharan,
and
S.
Belongie,
“Feature
p
yramid
netw
orks
for
object
detec-
tion,
”
in
Pr
oceedings
of
the
IEEE/CVF
Confer
ence
on
Com
puter
V
ision
and
P
attern
Reco
gnition
(CVPR)
,
2017,
pp.
2117–2125,
doi:
10.1109/CVPR.2017.106.
[17]
M.
Maiza,
C.
Cherif,
S.
Chouraqui,
and
A.
T
aleb-Ahmed,
“Enhanced
U
A
V
na
vig
ation
in
cluttered
en
vironments
via
a
BS-CY
OLOv5
multi-sensor
approach,
”
Disco
ver
V
ehicles
,
v
ol.
2,
p.
3,
2026.
[18]
Y
.
W
u
and
K.
He,
“Group
normalization,
”
in
Pr
oceedings
of
the
Eur
opean
Confer
ence
on
Computer
V
ision
(ECCV)
,
2018,
pp.
3–19,
doi:
10.1007/978-3-030-01261-8
1.
[19]
Y
.
P
.
Loh
and
C.
S.
Chan,
“Getting
to
kno
w
lo
w-light
images
with
the
e
xclusi
v
ely
dark
dataset,
”
Computer
V
ision
and
Ima
g
e
Under
standing
,
v
ol.
178,
pp.
30–42,
2019,
doi:
10.1016/j.cviu.2018.10.010.
[20]
Y
.
Y
u,
Y
.
Zhang,
Z.
Cheng,
Z.
Song,
and
J.
T
ang,
“MCA:
Multidimensional
col
laborati
v
e
attention
in
deep
con-
v
olutional
neural
netw
orks
for
image
recognition,
”
Engineering
Applications
of
Articial
Intellig
ence
,
v
ol.
126,
p.
107079,
2023,
doi:
10.1016/j.eng
appai.2023.107079.
[21]
Z.-Y
.
Ni
and
J.-C.
W
ang,
“
A
global
attention
mechanism-based
Ef
cientNet
model
for
road
pa
v
em
ent-type
identi-
cation,
”
International
J
ournal
of
Computational
Intel
lig
ence
Systems
,
v
ol.
18,
p.
97,
2025,
doi:
10.1007/s44196-025-
00842-3.
[22]
D.
Du
et
al.
,
“V
isDrone-DET2019:
The
vision
meets
drone
object
detection
in
image
challenge
results,
”
in
Pr
oceed-
ings
of
the
IEEE
/CVF
International
Confer
ence
on
Computer
V
ision
W
orkshops
(ICCVW)
,
2019,
pp.
213–226,
doi:
10.1007/978-3-030-11021-5
27.
[23]
S.
P
ark,
Y
.-J.
Y
eo,
and
Y
.-G.
Shin,
“PCon
v:
Simple
yet
ef
fecti
v
e
con
v
olutional
layer
for
generati
v
e
adv
ersarial
net-
w
ork,
”
Neur
al
Computing
and
Applications
,
v
ol.
34,
pp.
7113–7124,
2022,
doi:
10.1007/s00521-021-06846-2.
[24]
“FYPproject
dataset,
”
Roboow
Univer
se
,
Oct.
2022.
https://uni
v
erse.roboo
w
.com/muaz-ahmed/fypproject.
[25]
D.
Ouyang
et
al.
,
“Ef
c
ient
multi-scale
attention
module
with
cross-spatial
learning,
”
in
Pr
oceedings
of
the
IEEE
International
Confer
ence
on
Acoustics,
Speec
h
and
Signal
Pr
ocessing
(
ICASSP)
,
2023,
pp.
1–5,
doi:
10.1109/ICASSP49357.2023.10096516.
Spatial-c
hannel
r
econstruction
for
ef
cient
multiscale
attention
in
r
obotic
object
(Mohammed
Maiza)
Evaluation Warning : The document was created with Spire.PDF for Python.
552
❒
ISSN:
2722-2586
BIOGRAPHIES
OF
A
UTHORS
Mohammed
Maiza
is
a
research
scientist
in
the
Computer
Science
Department
at
the
Uni
v
ersity
of
Oran
1
Ahmed
Ben
Bella,
Algeria.
He
holds
an
M.Sc.
and
a
Ph.D.
in
computer
science
from
the
Uni
v
ersity
of
Sciences
and
T
echnology
of
Oran.
His
research
focuses
on
bioinformatics,
computational
biology
,
and
machine
learning,
with
emphasis
on
high-dimensional
gene
selection,
mi-
croarray
data
anal
ysis,
and
h
ybrid
e
v
olutionary
algorithms
for
cancer
classication.
He
collaborates
with
medical
r
esearchers
to
de
v
elop
computational
methods
with
clinical
applications
in
genomics
and
precision
medicine.
He
can
be
contacted
at
email:
maiza.mohammed@uni
v-oran1.dz.
Chahira
Cherif
is
a
research
scientist
in
computer
science
at
the
Uni
v
ersity
of
Oran
1
Ahmed
Ben
Bella,
Algeria
,
where
she
earned
her
engineering,
master
’
s
,
and
Ph.D.
de
grees.
Her
research
focus
es
on
arti
cial
intelligence,
b
usiness
process
ma
nagement,
decision
support
systems,
and
b
usiness
rules
modeling,
emphasizing
the
inte
gration
of
AI
into
b
usiness
process
optimization
and
decision-making.
She
contrib
utes
to
national
and
international
research
projects,
supervises
graduate
students,
and
her
current
w
ork
centers
on
intelligent
decision
support
systems
using
ma-
chine
learning,
particularly
in
healthcare
and
industrial
conte
xts.
She
can
be
contacted
at
email:
cherif.chahira@uni
v-oran1.dz.
Samira
Chouraqui
is
a
full
professor
in
computer
science
at
the
Uni
v
ersity
of
Sciences
and
T
echnology
of
Oran
(UST
O
MB),
Algeria.
She
earned
an
engineering
diploma
in
electrical
engineering
from
UST
O
Oran,
an
MSc
in
satellite
communication
from
the
Uni
v
ersity
of
Surre
y
,
UK,
and
a
Ph.D.
in
computer
science
from
UST
O
Oran.
Her
research
spans
computer
vision,
AI,
U
A
Vs,
pattern
recognition,
a
nd
machine
learning
for
remote
sensing,
with
current
focus
on
intelligent
vision
systems,
U
A
V
autonomous
na
vig
ation,
and
AI-dri
v
en
solutions
for
en
vironmental
monitoring
and
precision
agriculture.
She
has
supervised
numerous
graduate
students
and
led
se
v
eral
national
research
projects.
She
can
be
contacted
at
email:
samira.chouraqui@uni
v-usto.dz.
Abdelmalik
T
aleb-Ahmed
is
a
full
professor
at
the
Polytechnic
Uni
v
ersity
of
Hauts-de-
France,
France,
af
liated
with
IEMN.
His
research
co
v
ers
computer
vision,
machine
learning,
pattern
recognition,
image
s
e
gmentation,
classication,
and
data
fusion,
with
applications
in
biometrics,
video
surv
eillance,
autonomous
dri
ving,
and
medical
imaging.
He
has
co-authored
o
v
er
85
peer
-
re
vie
wed
papers,
supervised
30
graduate
students,
and
led
national
and
Europe
an
research
projects.
His
current
w
ork
focuses
on
deep
learning
for
medical
image
analysis,
multimodal
data
fusion
for
autonomous
systems,
and
adv
anced
pattern
recognition
for
security
applications.
He
can
be
contacted
at
email:
abdelmalik.taleb-ahmed@uphf.fr
.
IAES
Int
J
Rob
&
Autom,
V
ol.
15,
No.
3,
September
2026:
544-552
Evaluation Warning : The document was created with Spire.PDF for Python.