IAES
Inter
national
J
our
nal
of
Robotics
and
A
utomation
(IJRA)
V
ol.
15,
No.
3,
September
2026,
pp.
621
∼
638
ISSN:
2722-2586,
DOI:
10.11591/ijra.v15i3.pp621-638
❒
621
Consensus-based
path
planning
f
or
U
A
V
swarms
under
multiple
constraints:
A
r
e
view
Y
ana
Lu,
Lianpeng
Li,
Hui
Zhao,
Xu
Zhao
School
of
Automation,
Beijing
Information
Science
and
T
echnology
Uni
v
ersity
,
Beijing,
China
Article
Inf
o
Article
history:
Recei
v
ed
Feb
9,
2026
Re
vised
Apr
27,
2026
Accepted
May
22,
2026
K
eyw
ords:
Classical
path
search
Consistenc
y
constraints
Deep
reinforcement
learning
Intelligent
optimization
algorithms
P
ath
planning
Unmanned
aerial
v
ehicle
sw
arm
ABSTRA
CT
Unmanned
aerial
v
ehicle
(U
A
V)
sw
arms
are
essential
for
emer
genc
y
response,
logistics,
reconnaissance,
and
en
vironmental
monitoring,
yet
achie
ving
safe
and
scalable
path
planning
under
dynamic
conditions
and
comple
x
constraints
re-
mains
challenging.
Unlik
e
e
xisting
surv
e
ys
that
cate
gorize
algorithms
by
theo-
retical
foundations,
this
paper
systematically
re
vie
ws
U
A
V
sw
arm
path
planning
through
the
lens
of
spatial,
temporal,
and
task-le
v
el
consistenc
y
constraints
.
W
e
classify
recent
adv
ances
into
classical
path
search,
intelligent
optimization,
and
deep
reinforcement
learning,
emphasizing
ho
w
each
addresses
geometric
con-
tinuity
,
beha
vioral
coordination,
and
full-chain
perception–decision–planning
consistenc
y
under
multi-constraint
coupling.
W
e
further
identify
critical
limita-
tions
in
scalability
,
dynamic
adaptability
,
and
heterogeneous
sw
arm
cooperation,
and
outline
future
directions,
including
distrib
uted
control,
multi-source
per
-
ception
fusion,
cross-platform
coll
aboration,
and
rob
ust
autonomous
decision-
making.
This
re
vie
w
pro
vides
a
unique,
application-centric
taxonomy
based
on
consensus
constraints,
of
fering
actionable
insights
for
de
v
eloping
consistenc
y-
a
w
are
U
A
V
sw
arm
path
planning
technologies.
This
is
an
open
access
article
under
the
CC
BY
-SA
license
.
Corresponding
A
uthor:
Li
Lianpeng
School
of
Automation,
Beijing
Information
Science
and
T
echnology
Uni
v
ersity
Beijing,
China
Email:
llp@bistu.edu.cn
1.
INTR
ODUCTION
W
ith
the
rapid
e
v
olution
of
autonomous
intelligent
systems,
Unmanned
aerial
v
ehicle
(U
A
V)
sw
ar
ms
—
composed
of
multiple
U
A
Vs
with
autonomous
perception,
distrib
uted
decision-making,
and
cooperati
v
e
interaction
capabilities
—
ha
v
e
sho
wn
transformati
v
e
potential
in
emer
genc
y
rescue,
logistics,
monitoring,
and
military
reconnaissance
[1],
[2].
Recent
breakthroughs
in
deep
learning
and
sw
arm
intelligence
ha
v
e
further
enhanced
their
situational
a
w
areness
and
cooperati
v
e
control
[3].
As
the
foundation
of
sw
arm
cooperation,
path
planning
directly
determines
mission
safety
and
has
become
a
k
e
y
indicator
of
system
intelligence.
Ho
we
v
er
,
research
on
U
A
V
sw
arm
path
planning
under
multiple
consensus
constraints
remains
in
its
early
stages,
with
only
a
fe
w
studies
ha
ving
just
be
gun
to
e
xplore
this
challenging
area
[4].
U
A
V
sw
arm
path
planning
aims
to
generate
a
set
of
feasible
trajectories
for
the
entire
sw
arm
in
c
om-
ple
x
en
vironments
while
satisfying
mission
requirements.
It
must
simultaneously
ensure
trajectory
continuity
,
temporal
synchronization,
and
task-le
v
el
decision
consistenc
y
,
t
hereby
enabling
coordinated
mission
e
x
ecution
[5]
and
a
v
oiding
ener
gy
inef
cienc
y
or
e
v
en
obstacle
a
v
oidance
f
ailures.
When
performing
comple
x
missions
in
dynamic
and
uncertain
en
vironments,
U
A
V
sw
arms
are
required
to
satisfy
consensus
constraints
across
mul-
tiple
dimensions,
including
spatial,
temporal,
and
cooperati
v
e
task
aspects.
J
ournal
homepage:
http://ijr
a.iaescor
e
.com
Evaluation Warning : The document was created with Spire.PDF for Python.
622
❒
ISSN:
2722-2586
Specically
,
spatial
constraints
[6]
require
planned
paths
to
be
continuously
e
x
ecutable
[7],
safe,
and
capable
of
ef
fecti
v
e
obstacle
a
v
oidance;
temporal
constraints
[8]
emphasize
mission
synchronization,
forma-
tion
maintenance
[9],
and
consistenc
y
in
action
timing
[10];
and
cooperati
v
e
task
constraints
[11]
are
reected
in
information
e
xchange
[12],
beha
vior
coordination,
and
strate
gy
unication
among
agents.
W
ith
the
increas-
ing
en
vi
ronmental
dynamics,
the
strengthening
coupling
among
heterogeneous
constraints,
and
t
he
continuous
gro
wth
of
sw
arm
scale
[13],
path
planning
under
multiple
constraints
has
progressi
v
ely
e
v
olv
ed
into
a
highly
comple
x
problem
characterized
by
high
dimensionality
,
nonlineari
ty
,
strong
coupling,
and
multi-objecti
v
e
op-
timization.
Consequently
,
the
applicability
of
traditional
path
planning
methods
in
such
scenarios
has
become
increasingly
limited.
T
o
address
the
abo
v
e
challenges,
research
ef
forts
ha
v
e
focused
on
problem
modeling,
constraint
inte-
gration
and
solution
strate
gies
[14].
In
classical
path
search,
impro
v
ed
heuristic
functions,
dynamic
cost
update
mechanisms,
and
local
conict
resolution
enhance
the
ability
to
handle
dynamic
obstacles,
spatiotem
po
r
al
con-
icts,
and
limited
cooperati
v
e
beha
viors.
Intelligent
optimization
algorithms,
le
v
eraging
strong
global
search
and
adaptability
,
enable
sw
arm-le
v
el
cooperati
v
e
optimization
under
comple
x
constraints,
mo
ving
path
plan-
ning
be
yond
geometric
feasibility
to
w
ard
multi
consensus
constraint
handling.
Meanwhile,
deep
reinforcement
learning
and
multi-agent
learning
allo
w
U
A
V
sw
arms
to
achie
v
e
continual
learning
and
distrib
uted
cooperation
in
dynamic,
uncertain
en
vironments,
of
fering
more
adapti
v
e
solutions
for
multi-constraint
path
planning.
Despite
signicant
progress
in
e
xisting
research,
maintai
ning
consensus
among
U
A
V
sw
arms
in
dy-
namic
en
vironments
remains
a
primary
challenge
for
path
planning.
F
actors
such
as
en
vironmental
dynam-
ics,
multi-source
perception
inconsistencies,
and
platform
heterogeneity
often
lea
d
to
de
g
r
aded
cooperati
v
e
performance
[15].
Therefore,
it
is
necessary
to
conduct
a
systematic
re
vie
w
from
the
perspecti
v
e
of
con-
sensus
constraints,
with
the
aim
of
pro
viding
important
theoretical
support
and
technical
references
for
the
de
v
elopment
of
highly
autonomous
U
A
V
sw
arm
systems.
Unlik
e
e
xisting
surv
e
ys
that
cate
gorize
algorithms
by
theoretical
foundations
[16],
this
paper
pro
vides
a
no
v
el
application-centric
taxonomy
based
on
spatial-
temporal-cooperati
v
e
task
consensus
constraints.
It
systematically
analyzes
ho
w
classi
cal
search,
intelligent
optimization,
and
deep
reinforcement
learning
methods
maintain
consensus
under
multiple
conicting
con-
straints,
critically
e
v
aluating
their
scalability
,
real-time
performance,
and
suitability
for
heterogeneous
sw
arms.
T
o
impro
v
e
the
structural
clarity
of
the
re
vie
wed
methodologies,
the
o
v
erall
path
planning
process
for
U
A
V
sw
arms
can
generally
be
summarized
into
se
v
eral
interconnected
stages,
including
en
vironment
percep-
tion,
constraint
modeling,
cooperati
v
e
task
allocation,
global
path
generation,
local
dynamic
replanning,
and
consistenc
y
maintenance.
Dif
ferent
planning
approaches
re
vie
wed
in
this
paper
mainly
dif
fer
in
their
imple-
mentation
strate
gies
for
these
stages,
particularly
in
balancing
global
optimality
,
real-ti
me
adaptability
,
and
cooperati
v
e
consistenc
y
under
multi-constraint
en
vironments,
sho
wn
in
Figure
1.
Furthermore,
it
discusses
planning
challenges
under
multi-constraint
coupling,
including
scalability
,
real-time
performance,
and
hetero-
geneous
sw
arm
cooperation,
and
concludes
with
major
challenges,
future
research
directions,
and
application
prospects.
Figure
1.
Ov
erall
frame
w
ork
of
the
proposed
consistenc
y-a
w
are
path
planning
methodology
for
U
A
V
sw
arms
The
structures
of
the
paper
are
as
follo
ws.
Section
1
describes
models
spatial,
temporal,
and
co-
operati
v
e
constraints.
Section
2
re
vi
e
ws
classical
search,
intelligent
optimization,
and
DRL
methods.
Section
IAES
Int
J
Rob
&
Autom,
V
ol.
15,
No.
3,
September
2026:
621–638
Evaluation Warning : The document was created with Spire.PDF for Python.
IAES
Int
J
Rob
&
Autom
ISSN:
2722-2586
❒
623
3
highlights
scalability
,
real-time
consistenc
y
,
and
heterogeneous
cooperation
as
future
challenges.
Finally
,
the
conclusion
summarizes
the
paper
.
2.
SP
A
TIAL-TEMPORAL-COOPERA
TIVE
T
ASK
CONSENSUS
CONSTRAINTS
2.1.
Spatial,
temporal,
and
cooperati
v
e
task
constraints
in
path
planning
U
A
V
sw
arm
path
planning
must
satisfy
three
critical
constraint
types.
Spatial
constraints
ensure
that
the
sw
arm
operates
continuously
,
feasibly
,
and
without
collisions
within
geometric
and
en
vironmental
limits
[17].
T
emporal
constraints
maintain
stability
and
coordination
in
both
path
e
x
ecution
and
decision-making.
Cooperati
v
e
task
constraints
require
that
the
sw
arm’
s
beha
vi
o
r
s
do
not
conict
and
collecti
v
ely
achie
v
e
globally
optimal
performance
[18].
T
ogether
,
these
three
constraints
constitute
a
consensus
constraint
frame
w
ork
for
U
A
V
sw
arm
path
planning.
In
terms
of
spatial
constraints,
U
A
V
sw
arms
must
operate
within
a
predened
three-dimensi
onal
airspace,
a
v
oidi
ng
no-y
zones
and
pre
v
enting
collisions
with
obstacles
[19]
to
ensure
safe
and
ef
cient
mission
e
x
ecution.
Static
obstacles
require
that
planned
paths
ef
fecti
v
ely
a
v
oid
them
at
the
initial
planning
stage028,
whereas
dynamic
obstacles
dem
and
real-time
perception
[20]
and
adapti
v
e
traj
ectory
adjustment
[21].
Conse-
quently
,
path
planning
algorithms
must
possess
high-le
v
el
predicti
v
e
capabilities
and
ef
cient
update
mecha-
nisms
to
handle
such
challenges
ef
fecti
v
ely
.
Re
g
arding
temporal
constr
aints,
U
A
V
sw
arms
missions
must
be
e
x
ecuted
in
a
sequential
manner
to
pre
v
ent
congestion
or
collisions
caused
by
premature
or
delayed
actions
of
indi
vidual
U
A
V
sw
arm.
This
is
particularly
critical
in
scenarios
such
as
emer
genc
y
response,
disaster
monitoring,
and
rapid
inspection,
where
mission
e
x
ecution
must
be
temporally
continuous
and
coherent.
Such
requirements
place
high
demands
on
the
real-time
performance
and
computational
ef
cienc
y
of
path
planning
algorithms
[22].
T
emporal
constraints
additionally
require
rapid
replanning
capability
under
limited
computational
resources.
Cooperati
v
e
task
constraints
represent
a
distincti
v
e
feature
of
U
A
V
sw
arm
path
planning.
I
n
di
vidual
U
A
V
sw
arm
within
the
sw
arm
are
required
to
maintain
appropriate
spatial
distrib
ution
and
safety
separation
during
mission
e
x
ecution
to
a
v
oid
collisions
due
to
e
xcessi
v
e
proximity
.
Dif
ferent
U
A
V
sw
arms
must
achie
v
e
complementarity
and
coordination
in
task
allocation
[23].
Furthermore,
the
sw
arm
needs
to
maintain
netw
ork
connecti
vity
through
wireless
comm
unication
to
enable
real-time
information
sharing
and
decision-making
consistenc
y
.
Consequently
,
path
planning
must
consider
not
onl
y
geometric
and
temporal
feasibility
b
ut
also
communication
range
constraints
and
topological
rob
ustness,
ensuring
that
the
sw
arm
can
cooperate
stably
in
dynamic
en
vironments.
2.2.
Pr
oblem
classication
dimensions
The
cooperati
v
e
collision
a
v
oidance
problem
for
U
A
V
sw
arms
e
xhibits
di
v
erse
mission
object
i
v
es,
algorithmic
implementations,
and
operational
en
vironment
charact
eristics.
T
o
systematically
address
the
afore-
mentioned
multi-constraint
path
planning
problem,
e
xisting
research
typically
classies
the
approaches
along
the
follo
wing
dimensions.
In
Figure
2,
a
three-dimensional
classication
is
conducted
from
the
perspecti
v
es
of
task
types,
algorithmic
characteri
stics,
and
en
vironment
(En
v).
The
algorithmic
cate
gory
includes
determin-
istic
algorithms
(DEA),
randomized
algorithms
(RA),
centralized
algorithms
(CA),
and
distrib
uted
algorithms
(D
A).
Figure
2.
Cluster
problem
classication
dimension
Consensus-based
path
planning
for
U
A
V
swarms
under
multiple
constr
aints:
A
r
e
vie
w
(Y
ana
Lu)
Evaluation Warning : The document was created with Spire.PDF for Python.
624
❒
ISSN:
2722-2586
2.2.1.
T
ask
types
In
tracking
missions,
U
A
V
s
w
arms
are
required
to
collaborati
v
ely
track
mo
ving
tar
gets
[24].
These
tasks
are
commonly
encountered
in
securit
y
patrols
[25]
and
traf
c
o
w
monitoring
[26],
[27].
Collision
a
v
oidance
strate
gies
need
to
balance
tar
get-tracking
accurac
y
with
sw
arm
safety
separation,
particularly
when
both
tar
gets
and
obstacles
are
in
motion,
imposing
high
real-time
requirements
on
path
planning.
In
search
and
rescue
missions
[28],
such
as
disaster
relief
or
maritime
search
operations,
U
A
V
sw
arms
must
rapidly
locate
tar
gets
in
comple
x
and
dynamic
en
vironments
[29].
Collision
a
v
oidance
st
rate
gies
in
these
tasks
must
not
only
handle
static
obstacles
b
ut
also
adapt
to
dynamic
disturbances
such
as
debris,
smok
e,
and
mo
ving
rescue
equipment,
while
completing
mission
co
v
erage
in
minimal
time.
In
man
y
practical
applications,
missions
may
simultaneously
in
v
olv
e
co
v
erage,
tracking,
and
search-
and-rescue
objecti
v
es.
F
or
e
xample,
post-disaster
aerial
inspection
requires
area
co
v
erage,
real-time
tracking
of
k
e
y
tar
gets,
and
searching
for
potential
trapped
indi
viduals.
Collision
a
v
oidance
strate
gies
for
such
comple
x
missions
must
possess
high
e
xibility
and
the
capability
to
switch
task
priorities
dynamically
.
2.2.2.
Algorithmic
featur
es
From
the
perspecti
v
e
of
algorithmic
features,
U
A
V
sw
arm
path
planning
methods
can
be
primarily
classied
along
tw
o
dimensions:
deterministic
v
ersus
stochastic
and
centralized
v
ersus
distrib
uted.
Determinis-
tic
algorithms
[30]
produce
consistent
outputs
under
identical
initial
conditions
[31],
of
fering
high
predictability
and
stability
.
The
y
are
well
suited
for
scenarios
with
well-dened
en
vironmental
models
and
clear
mission
ob-
jecti
v
es
[32].
In
contrast,
stochastic
algorithms
[33]
introduce
probabilistic
elements
into
the
decision-making
process,
enabling
more
di
v
erse
solutions
in
highly
uncertain
en
vironments
or
lar
ge
search
spaces.
These
algo-
rithms
are
commonly
emplo
yed
for
global
search,
a
v
oidance
of
local
optima,
and
adaptation
to
dynamically
changing
mission
conditions.
Re
g
arding
decision-making
architectures,
centralized
algorithms
[34]
rely
on
one
or
a
fe
w
central
nodes
to
aggre
g
ate
i
nformation
and
mak
e
global
decisions.
The
y
can
maintain
full
a
w
areness
of
system
states
and
achie
v
e
globally
optimal
solutions;
ho
we
v
er
,
the
y
impose
high
requirements
on
communication
links
and
computational
resources,
and
f
ailure
of
central
nodes
may
result
in
mission
interruption.
Distrib
uted
algorithms
[35],
on
the
other
hand,
allo
w
indi
vidual
U
A
V
sw
arm
to
mak
e
autonomous
decisions
based
solely
on
local
information
e
xchange
[36],
of
fering
higher
rob
ustness
and
scalability
,
and
are
suitable
for
lar
ge-scale
sw
arms
and
communication-constrained
scenarios
[37].
Ne
v
ertheless,
distrib
uted
algorithms
may
e
xhibit
limitations
in
global
optimi
zation
and
con
v
er
gence
speed
[38],
which
can
be
mitig
ated
by
designing
appropriate
local
rules
and
information-sharing
mechanisms.
2.2.3.
En
vir
onmental
characteristics
Classication
based
on
en
vironmental
characteristics
primarily
depends
on
en
vironmental
stability
,
which
can
be
di
vided
into
static
and
dynamic
en
vironments.
In
static
en
vironments,
the
positions
of
tar
gets
and
obstacles
are
x
ed,
as
in
inspection
and
mapping
tasks,
allo
wing
for
pre-planned
trajectories
and
relati
v
ely
simple
algorithms.
In
dynamic
en
vironments,
tar
gets,
obstacles,
or
conditions
change
o
v
er
time,
as
in
maritime
search
and
rescue
or
traf
c
monitoring,
requiring
real-time
perception
and
trajectory
adjustment,
making
path
planning
more
challenging.
Additionally
,
en
vironments
can
be
cate
gorized
according
to
the
a
v
ailability
of
en
vironmental
infor
-
mation
into
kno
wn
and
unkno
wn
en
vironments
[39].
In
kno
wn
en
vironments,
information
such
as
maps
of
the
mission
area,
obstacle
locations,
and
tar
get
characteristics
is
a
v
ailable
before
mission
e
x
ecution,
f
acilitating
the
use
of
global
planning
strate
gies
[40].
In
unkno
wn
en
vironments,
such
information
is
not
fully
accessible
prior
to
task
e
x
ecution,
necessitating
the
use
of
U
A
V
sensors
[41]
and
cooperati
v
e
perception
capabilities
[42]
to
progressi
v
ely
b
uild
an
en
vironmental
model.
Online
decision-making
and
local
path
planning
techniques
are
then
emplo
yed
to
respond
to
une
xpected
changes
[43].
Dif
ferent
en
vironmental
char
acteristics
directly
af
fect
the
comple
xity
of
mission
planning,
perception
systems,
and
control
strate
gies,
making
this
classication
an
essential
dimension
in
U
A
V
sw
arms
task
design.
2.3.
Constraint
modeling
In
the
mathematical
modeling
of
path
planning,
the
proper
formulation
of
constraints
directly
deter
-
mines
the
feasibility
and
optimality
of
the
planned
trajectories.
Based
on
mission
e
x
ecution
characteristics
and
sw
arm
coordination
requirements,
spatial,
temporal,
and
cooperati
v
e
task
constraints
need
to
be
mathematically
IAES
Int
J
Rob
&
Autom,
V
ol.
15,
No.
3,
September
2026:
621–638
Evaluation Warning : The document was created with Spire.PDF for Python.
IAES
Int
J
Rob
&
Autom
ISSN:
2722-2586
❒
625
modeled
to
accurately
describe
the
states
of
U
A
V
sw
arms,
including
position,
v
elocity
,
and
heading.
P
ath
plan-
ning
solutions
are
then
obtained
through
objecti
v
e
optimization,
enabling
the
sw
arm
to
ef
ciently
coordinate
[44],
a
v
oid
obstacles
[45],
and
accomplish
task
allocation
and
e
x
ecution
[46]
during
mission
operations.
2.3.1.
Mathematical
modeling
The
mission
space
is
modeled
as
a
three-dimensional
Euclidean
space,
consisting
of
a
feasible
i
ght
re
gion
and
an
obstacle
set
that
do
not
o
v
erlap.
The
U
A
V
sw
arm
contains
a
total
of
N
unmanned
aerial
v
ehicles.
At
an
y
time,
the
state
of
each
U
A
V
includes
its
position
v
ector
,
its
v
elocity
,
and
its
heading
angle.
In
terms
of
spatial
modeling,
study
[47]
i
n
t
roduces
a
h
ybrid
approach
that
combines
discrete
and
con-
tinuous
states
to
describe
U
A
V
sw
arms.
By
linking
discrete
sw
arm
e
v
ents
with
continuous
state
v
ariables,
this
approach
captures
the
dynamic
motion
and
interaction
of
U
A
Vs
in
three-dimensional
space.
The
mission
objecti
v
e
is
commonl
y
dened
by
a
path
cost
function
J
,
which
considers
multiple
performance
metrics,
in-
cluding
ight
distance,
ener
gy
consumption,
mission
duration,
and
risk.
Accordingly
,
the
o
v
erall
objecti
v
e
can
be
formulated
as:
min
P
J
=
N
X
i
=1
(
α
1
L
i
+
α
2
E
i
+
α
3
T
i
+
α
4
R
i
)
(1)
Where
P
denotes
the
set
of
trajectories
of
all
U
A
V
sw
arms,
L
i
represents
the
total
path
length,
E
i
denotes
the
ener
gy
consumption,
T
i
is
the
mission
e
x
ecution
time,
and
R
i
corresponds
to
the
risk
cos
t.
a
k
denotes
the
weighting
coef
cients,
which
are
assigned
according
to
task
priorities.
The
weighting
coef
cients
α
k
in
(1)
allo
w
for
mission-specic
trade-of
fs,
e.g.,
prioritizing
ener
gy
(
α
2
)
o
v
er
path
length
(
α
1
)
for
long-
endurance
missions.
T
o
ensure
the
fea
sibility
of
planned
trajectories,
dynamic
constraints
are
introduced
at
the
modeling
stage
to
reect
the
kinematic
characteristics
and
platform
performance
limitations
of
U
A
V
sw
arm
[48].
T
yp-
ically
,
the
v
elocity
,
acceleration,
and
rate
of
change
of
heading
angle
of
each
U
A
V
are
required
to
satisfy:
ν
min
≤∥
ν
i
(
t
)
∥≤
ν
max
,
∥
˙
ν
i
(
t
)
∥≤
a
max
,
|
˙
ψ
i
(
t
)
|≤
ψ
max
(2)
In
(2),
v
min
,
v
max
,
a
max
,
and
ψ
max
represent
the
platform’
s
limits
on
speed,
acceleration,
and
turning
rate,
respecti
v
ely
.
These
dynamic
constraints
ensure
kinematic
feasibility
.
Spatial
constraints
are
introduced
to
ensure
the
safe
ight
of
U
A
V
sw
arm
within
the
mission
space.
Specically
,
spatial
safety
constraints
require
that
the
ight
trajectories
do
not
intersect
with
obstacle
sets
while
maintaining
a
prescribed
minimum
safety
distance
d
min
.
∥
p
i
(
t
)
−
p
j
(
t
)
∥≥
d
min
,
p
i
(
t
)
/
∈
Ω
o
(3)
This
constraint
ensures
that
U
A
Vs
f
o
l
lo
w
planned
trajectories
within
admissible
attitude
and
v
el
ocity
limits.
It
pre
v
ents
aggressi
v
e
maneuv
ers
that
may
cause
e
xcessi
v
e
ener
gy
consumption,
ight
instability
,
or
control
f
ailure,
thereby
guaranteeing
the
ph
ysical
feasibility
and
safety
of
the
trajectories.
T
emporal
constraints
describe
time
coordination
among
U
A
Vs
during
mission
e
x
ecution.
The
y
ensure
mission
completion
within
a
gi
v
en
time
windo
w
while
maintaining
synchronization
and
cooperati
v
e
beha
vior
.
In
U
A
V
sw
arm
planning,
time
af
fects
task
ef
cienc
y
,
coordination
accurac
y
,
and
system
stability
.
In
U
A
V
sw
arm
cooperati
v
e
missions
[49],
temporal
constraints
in
v
olv
e
more
than
indi
vidual
timing.
The
y
also
capture
task
sequencing
and
synchronization
within
the
sw
arm.
F
or
tasks
requiring
coordination,
such
as
collaborati
v
e
reconnaissance,
formation
ight,
distrib
uted
monitoring,
or
mult
i-point
strik
es,
the
fol-
lo
wing
constraints
must
be
satised:
t
ar
r
iv
e
i
≤
t
ar
r
iv
e
j
|
t
ar
r
iv
e
i
−
t
ar
r
iv
e
j
|≤
∆
t
sy
nc
(4)
Here,
∆
t
sy
nc
represents
the
allo
w
able
time
synchronization
de
viation.
The
former
applies
to
scenarios
with
sequential
task
requirements;
the
latter
applies
to
scenarios
requiri
ng
simultaneous
arri
v
al
or
synchronized
Consensus-based
path
planning
for
U
A
V
swarms
under
multiple
constr
aints:
A
r
e
vie
w
(Y
ana
Lu)
Evaluation Warning : The document was created with Spire.PDF for Python.
626
❒
ISSN:
2722-2586
e
x
ecution,
ensuring
that
all
U
A
V
sw
arms
perform
actions
within
the
permissible
time
de
viation,
thereby
main-
taining
o
v
erall
mission
coordination.
Cooperati
v
e
task
modeli
ng.
Cooperati
v
e
task
constraints
are
introduced
to
ensure
that
U
A
V
sw
arm
systems
achie
v
e
information
s
h
a
ring,
motion
coordination,
and
decision
consistenc
y
during
mission
e
x
ecution,
serving
as
k
e
y
elements
for
realizing
ef
cient
collecti
v
e
intelligence
beha
vior
[50].
F
ormation-k
eeping
constraints
are
introduced
to
ensure
that
U
A
V
sw
arm
maintain
a
prescribed
spatial
geometric
relationship
during
formation
or
cooperati
v
e
missions
[51].
Let
r
∗
i
denote
the
relati
v
e
position
of
U
A
V
i
with
respect
to
the
formation
center
in
the
ideal
formation.
During
actual
ight,
the
follo
wing
condition
must
be
satised:
∥
(
p
i
(
t
)
−
p
c
(
t
))
−
r
∗
i
∥≤
ε
f
(5)
where
p
c
(
t
)
represents
the
position
of
the
formation
center
and
ε
f
denotes
the
allo
w
able
formation
tolerance.
This
constraint
enables
the
U
A
V
sw
arm
to
maintain
o
v
erall
shape
during
maneuv
ers,
obstacle
a
v
oidance,
or
tar
get
tracking,
pre
v
enting
e
xcessi
v
e
dispersion
or
structural
disorder
of
the
formation.
Cooperati
v
e
task
constraints
capture
functional
di
vision
and
complementarity
among
U
A
Vs
at
the
task
le
v
el.
F
or
multi-objecti
v
e
missions,
task
allocati
on
must
be
both
unique
and
complete,
enabling
global
opti-
mization
under
sw
arm
cooperation.
These
constraints
help
the
sw
arm
maintain
structural
stability
,
coordinated
beha
vior
,
and
information
consistenc
y
in
comple
x,
dynamic
en
vironments,
thereby
impro
ving
operational
ef-
cienc
y
and
rob
ustness.
T
ogether
with
spatial
and
temporal
constraints,
the
y
form
the
foundational
modeling
frame
w
ork
for
U
A
V
sw
arm
path
planning.
The
planning
of
the
path
of
the
U
A
V
sw
arm
must
accommodate
v
arious
application
s
cenarios
and
strik
e
missions,
each
imposing
dif
f
erent
requirements
on
co
v
erage
ef
cienc
y
,
tar
get
prioritization,
formation
maintenance,
as
well
as
path
optimality
and
threat
a
v
oidance
[52].
F
or
heterogeneous
sw
arms,
dif
ferences
in
platform
e
nd
ur
ance,
payload
capacity
,
and
sensing
capabilities
must
also
be
considered
[53].
Consequently
,
this
problem
is
inherently
a
multi-objecti
v
e
constrained
optimization
task,
whose
c
o
m
ple
xity
arises
from
high-
dimensional
state
spaces,
dynamic
en
vironments,
and
sw
arm
coupling.
In
practice,
a
balance
must
be
struck
between
modeling
accurac
y
and
computational
feasibility
to
ensure
mission
ef
fecti
v
eness
while
maintaining
real-time
performance.
2.3.2.
Objecti
v
e
optimization
In
U
A
V
sw
arm
cooperati
v
e
path
planning
and
scheduling,
the
proper
formulation
of
optimization
objecti
v
es
directly
determines
system
performance
and
mission
ef
cienc
y
[54].
These
objecti
v
es
typically
encompass
multiple
dimensions,
including
path
length,
ener
gy
consumption,
time
ef
cienc
y
,
safety
,
and
task
co
v
erage.
Owing
to
v
arying
mission
requirements,
these
objecti
v
es
often
conict
with
one
another
,
necessitat-
ing
the
use
of
multi-objecti
v
e
optimization
methods
to
achie
v
e
ef
fecti
v
e
trade-of
fs
[55].
At
the
path
planning
l
e
v
el
,
minimizing
ight
distance
or
mission
completion
time
is
a
k
e
y
objecti
v
e
for
impro
ving
operational
ef
cienc
y
,
particularly
in
time-sensiti
v
e
tasks
such
as
emer
genc
y
response
and
in-
spection.
This
problem
is
typically
formulated
as
a
shortest-path
or
optimal
scheduling
problem
and
is
solv
ed
using
enhanced
methods
such
as
A*,
ant
colon
y
optimizati
on
,
or
particle
sw
arm
optimization,
balancing
com-
putational
ef
cienc
y
with
path
feasibility
.
At
the
mission
le
v
el,
sw
arm
planning
aims
to
maximize
task
completion
and
area
co
v
erage
ef
ciently
under
time
and
ener
gy
limits.
This
is
critical
for
area
search,
en
vironmental
monitoring,
and
agricultural
inspection,
typically
achie
v
ed
via
re
gion
partitioning
or
information-g
ain-based
co
v
erage
models.
From
the
perspecti
v
e
of
ener
gy
utilizati
on,
minimizing
ener
gy
consumption
is
a
critical
objecti
v
e
in
U
A
V
path
planning
[56],
[57].
Ener
gy
e
xpenditure
is
inuenced
by
f
actors
such
as
path
length,
i
g
ht
speed,
v
ehicle
dynamics,
and
en
vironmental
disturbances.
By
incorporating
ight
dynamics
models
and
optimizing
speed
proles
and
trajectory
design,
ener
gy
ef
cienc
y
can
be
signicantly
impro
v
ed,
which
is
particularly
important
in
scenarios
with
limited
endurance.
The
aforementioned
optimization
objecti
v
es
in
practical
missions
are
often
coupled
and
in
v
olv
e
trade-
of
fs
among
multiple
criteria.
Consequently
,
multi-objecti
v
e
optimizat
ion
frame
w
orks
are
typically
emplo
yed
in
modeling
and
solution
processes
[58],
using
weighting
coef
cients
to
achie
v
e
a
balanced
compromise
[59].
Such
optimization
strate
gies
enable
coordination
among
dif
ferent
performance
metrics,
allo
wing
the
U
A
V
sw
arm
to
achie
v
e
o
v
erall
mission
ef
fecti
v
eness
in
comple
x
en
vironments.
IAES
Int
J
Rob
&
Autom,
V
ol.
15,
No.
3,
September
2026:
621–638
Evaluation Warning : The document was created with Spire.PDF for Python.
IAES
Int
J
Rob
&
Autom
ISSN:
2722-2586
❒
627
3.
CONSISTENCY
-CONSTRAINED
PLANNING
METHODS
Research
on
U
A
V
sw
arm
path
planning
in
v
olv
es
v
arious
methods
that
dif
fer
in
computational
com
-
ple
xity
,
adaptability
,
global
optimal
ity
,
and
real-time
performance,
as
summarized
in
T
able
1.
These
methods
can
be
broadly
classied
into
classical
path
search,
intelligent
optimization,
machine
learning
and
deep
rein-
forcement
learning,
as
well
as
h
ybrid
and
multi-objecti
v
e
approaches
that
combine
multiple
techniques.
T
able
1.
Ov
ervie
w
of
typical
path
planning
techniques
and
their
characteristics
T
echnology
cate
gory
Main
methods
Adv
antages
Limitations
En
vironment
modeling
Grid-based
[60],
T
opological
graph
[61],
Continuous
model-
ing
[62],
Semantic
Map
[63]
Clear
information
representation,
adaptable
to
multiple
scenarios
T
rade-of
f
between
modeling
accu-
rac
y
and
real-time
performance
Classical
search
Dijkstra
[64],
A*
[65],
D*
[66],
Theta*
[67]
High
interpretability
,
guaranteed
optimality
under
certain
condi-
tions
High
computational
cost
in
high-
dimensional
scenarios
Sampling-based
meth-
ods
RR
T
[68],
RR
T*
[69],
PRM
[70]
Suitable
for
high-dimensional
continuous
spaces,
capable
of
handling
non-con
v
e
x
en
viron-
ments
Requires
post-optimization,
global
optimality
is
dif
cult
to
guarantee
Intelligent
optimiza-
tion
Genetic
algorithm
[71],
P
article
sw
arm
optimization
[72],
Ant
colon
y
optimization
[73]
Strong
global
search
capability
,
suitable
for
multi-objecti
v
e
opti-
mization
tasks
Slo
w
con
v
er
gence
and
sensiti
vity
to
parameter
settings
Deep
reinforcement
learning
DQN
[74],
PPO
[75],
SA
C
[76],
End-to-end
na
vig
ation
net-
w
orks
[77]
High
adaptability
and
scalability
,
suitable
for
dynamic
en
viron-
ments
High
training
cost
and
limited
generalization
capability
3.1.
Classical
path
sear
ch-based
methods
In
path
pl
anning,
graph
sear
ch
methods
ha
v
e
long
been
central
due
to
their
clear
structure,
i
nter
-
pretability
,
and
good
con
v
er
gence,
as
sho
wn
in
Figure
3.
W
ith
adv
ances
in
unmanned
systems
and
more
com-
ple
x
scenarios,
traditional
frame
w
orks
ha
v
e
been
e
xtended
to
impro
v
e
inte
grated
optimization
under
temporal,
spatial,
and
dynamic
consistenc
y
constraints.
These
enhancements
allo
w
planning
results
to
better
mat
ch
real
missions
and
en
vironmental
dynamics.
Figure
3.
Logic
diagram
of
classical
algorithm,
including
Dijkstra,
A*,
RR
T
,
and
PRM,
and
compares
their
respecti
v
e
characteristics
in
global
search,
heuristic
guidance,
and
sampling-based
planning
Consensus-based
path
planning
for
U
A
V
swarms
under
multiple
constr
aints:
A
r
e
vie
w
(Y
ana
Lu)
Evaluation Warning : The document was created with Spire.PDF for Python.
628
❒
ISSN:
2722-2586
The
de
v
elopment
of
classical
path
search
algorithms
under
consistenc
y
constraints
has
progressed
from
static
optimization
to
dynamic
consistenc
y
maintenance
and
then
to
multi-objecti
v
e
coordination.
Early
methods,
based
on
the
traditional
Dijkstra
algorithm
[78],
focused
on
nding
optimal
paths
in
static
graphs.
As
tasks
e
xpanded
to
comple
x
en
vironments
with
v
arying
attrib
utes,
In
2021
[79],
proposed
the
re
v
erse-label
Dijkstra
algorithm,
which
separates
road
tra
v
el
time
from
intersection
w
aiting
time
and
dynamically
updates
costs
using
a
re
v
erse-label
mechanism,
ensuring
path
timing
aligns
with
en
vironmental
attrib
utes.
In
2023,
consistenc
y
constraints
were
further
applied
to
state-space
search
and
sensor
calibration
path
planning
[80].
The
impro
v
ed
Dijkstra
method
uses
observ
ability
as
the
path
cost
and
emplo
ys
dynamic
error
estimation
with
multi-path
iteration
to
align
error
con
v
er
gence
with
observ
ability
e
v
aluation,
enabling
optimal
path
search
in
high-dimensional
error
space
and
impro
ving
the
stability
and
accurac
y
of
system-le
v
el
calibration.
As
application
demands
shift
to
w
ard
higher
dynamic
responsi
v
eness
and
continuity
,
path
smoothness
and
dynamic
feasibility
ha
v
e
become
k
e
y
e
xtensions
of
classical
search
algorithms.
T
o
addres
s
the
A*
algo-
rithm’
s
limitations
in
ef
cienc
y
and
smoothness,
study
[81]
proposed
the
self-adapti
v
e
neighborhood
search
A*
(SANSA)
algorithm.
SANSA
impro
v
es
ef
cienc
y
by
adjusting
the
se
arch
neighborhood
and
obstacle
han-
dling,
and
applies
path
smoothing
after
generation.
This
ensures
trajectories
maintain
geometric
and
kinematic
continuity
,
enabling
coordinated
preserv
ation
of
consistenc
y
and
feasibility
,
marking
a
shift
from
discrete
paths
to
continuous
trajectories.
W
ith
t
he
rise
of
U
A
V
sw
arm
planning,
classical
search
algorithms
ha
v
e
been
inte
grat
ed
with
coop-
erati
v
e
mechanisms
to
support
global
consistenc
y
and
multi-agent
coordination.
In
[82]
proposed
a
3D
jump
point
search
(3D-JPS)
cooperati
v
e
algorithm,
using
incremental
conict
resolution,
dual-objecti
v
e
cost
func-
tions,
and
consistenc
y
constraints
to
maintain
conict-free
trajectories.
In
[83]
introduced
a
h
ybrid
genetic
algorithm-D*
(GA-D*)
method,
combining
genetic
algorithm
s’
global
search
with
D*’
s
dynamic
replanning.
Probabilistic
constraints
and
incremental
consistenc
y
mechanisms
enable
coordinated
task
all
ocation
and
path
planning.
In
summary
,
while
classical
search
algorithms
lik
e
A*
and
Dijkstra
of
fer
theoretical
optimality
and
high
i
nterpretability
under
static
conditions
,
their
e
xtensions
to
dynamic
en
vironments
ofte
n
suf
fer
from
e
xpo-
nentially
increasing
computational
costs.
As
summarized
in
T
able
2,
their
primary
limitation
lies
in
maintaining
real-time
trajectory
continuity
and
scalability
for
lar
ge
sw
arms
,
a
g
ap
that
intelligent
optimization
algorithms
attempt
to
ll,
albeit
with
their
o
wn
con
v
er
gence
challenges.
T
able
2.
Classic
path
search
algorithms
and
their
consistenc
y
characteristics
Algorithm
Core
mechanism
Manifestation
of
consistenc
y
constraint
Dijkstra
[78]
Systematically
tra
v
erses
all
reachable
nodes
in
the
graph
to
guarantee
the
optimal
path
Static
optimality
consistenc
y
Re
v
erse-Label
Dijkstra
[79]
Separates
static
path
cost
from
dynamic
temporal
cost
to
impro
v
e
adaptability
in
time-v
arying
en
vironments
Cost
consistenc
y
Impro
v
ed
Dijkstra
[80]
Performs
optimal
path
search
in
high-dimensional
state
spaces,
such
as
error
or
augmented
state
spaces
State
consistenc
y
SANSA
[81]
Conducts
heuristic
search
follo
wed
by
path
smoothing
optimization
to
enhance
trajectory
feasibility
Motion
consistenc
y
3D-JPS
[82]
Resolv
es
path
conicts
among
multiple
agents
in
an
incremental
and
prioritized
manner
Multi-agent
cooperati
v
e
consistenc
y
GA-D*
[83]
Combines
the
global
search
capability
of
genetic
al-
gorithms
with
the
dynamic
replanning
of
heuristic
methods
Dynamic
cooperati
v
e
consistenc
y
3.2.
Intelligent
optimization-based
methods
Intelligent
optimization
algorithms,
grounded
in
sw
arm
intelligence
and
e
v
olutionary
m
echanisms,
e
xhibit
rob
ust
global
search
capabilities
and
adapt
ability
,
rendering
them
well-suited
for
multi-constraint
path
planning
in
U
A
V
sw
arms.
Early
genetic
algorithm-based
path
planning
focused
on
path
geometry
and
feasibility
consistenc
y
.
In
2024
[84],
proposed
a
h
ybrid
genetic
algorithm-rapidly
e
xploring
random
tree
(GA-RR
T)
method,
using
RR
T
to
b
uild
the
initial
population
and
applying
redundant
node
elimination
and
path
backtracking
to
impro
v
e
path
smoothness,
obstacle-a
v
oidance
continuity
,
and
e
x
ecutability
,
turning
consistenc
y
constraints
into
e
xplicit
operations.
As
U
A
V
sw
arm
applications
demand
partial
information,
high
dynamics,
and
strong
coupling,
IAES
Int
J
Rob
&
Autom,
V
ol.
15,
No.
3,
September
2026:
621–638
Evaluation Warning : The document was created with Spire.PDF for Python.
IAES
Int
J
Rob
&
Autom
ISSN:
2722-2586
❒
629
consistenc
y
e
xtends
to
strate
gy
and
task
coordination.
In
[85]
used
multi-agent
proximal
polic
y
optimization
(MAPPO)
with
a
centralized
v
alue
netw
ork
to
ensure
beha
vioral
and
strate
gy
consistenc
y
among
agents,
while
dynamic
cluster
particle
sw
arm
optimization
(DCPSO)
applied
dynamic
clustering,
potential
eld
constraints,
and
receding-horizon
replanning
to
align
search
direction,
v
elocity
,
and
path
con
v
er
gence.
These
approaches
enhance
system-le
v
el
consistenc
y
,
impro
v
e
stability
in
mul
ti-agent
s
cenarios,
and
reduce
issues
of
local
optima
and
beha
vioral
oscillations.
In
dynamic,
real-time
tasks,
intelligent
optimization
algorithms
e
xtend
consistenc
y
to
full-chain
syn-
chronization
across
information,
decisions,
and
planning.
Methods
lik
e
online
multi-iteration
ant
colon
y
with
receding-horizon
[86]
update
en
vironmental
perception,
resource
states,
and
local
costs
at
eac
h
step,
k
eeping
planned
paths
consistent
with
the
changing
en
vironment.
In
[87]
proposed
a
modied
ant
colon
y
optimization
(MA
CO)
within
a
tw
o-step
frame
w
ork.
ant
colon
y
optimization
(A
CO)
plans
feasible
paths
for
ground
v
ehicles,
and
a
genetic
algorithm
optimizes
U
A
V
ight
strate
gies.
By
enforcing
consistenc
y
constraints
on
tak
eof
f/landing
order
,
endurance,
and
path
coupling,
it
achie
v
es
global
coordination
for
heterogeneous
systems
under
road
netw
ork
limits.
This
reects
the
e
v
olution
of
intelligent
optimization
algorithms
from
single-step
to
full-process
consistenc
y
.
In
summary
,intelligent
optimization
algorithms
e
xcel
at
global
search
and
multi-objecti
v
e
optimiz
a-
tion,
making
them
suitable
for
com
ple
x
multi-constraint
path
planning.
As
summarized
in
T
able
3,
the
y
ha
v
e
e
v
olv
ed
from
ensuring
geometric
path
continuity
to
achie
ving
full-chain
information–decision–planning
con-
sistenc
y
.
Ho
we
v
er
,
their
slo
w
con
v
er
gence,
parameter
sensiti
vity
,
and
high
computational
o
v
erhead
limit
their
use
in
highly
dynamic
or
real-time
scenarios.
F
or
lar
ge-scale
sw
arms
or
rapidly
changing
en
vironments,
intel-
ligent
optimization
alone
is
insuf
cient,
often
requiring
h
ybrid
or
learning-based
inte
gration.
T
able
3.
Intelligent
optimization
algorithms
and
their
consistenc
y
characteristics
Algorithm
Core
mechanism
Manifestation
of
consistenc
y
constraints
GA-RR
T
[84]
Global
optimization
search
using
heuristic
methods
P
ath
geometric
consistenc
y
MAPPO
[85]
Centralized
learning
during
training
and
distrib
uted
decision-making
during
e
x
e-
cution
Multi-agent
strate
gy
consistenc
y
DCPSO
[85]
Adapti
v
e
optimization
of
particle
s
w
arm
in
dynamic
en
vironments
Search
process
consistenc
y
Online
Multi-iteration
ant
colon
y
strate
gy
[86]
Online
planning
strate
gy
based
on
a
receding-horizon
windo
w
Information-decision
consistenc
y
MA
CO
[87]
Cooperati
v
e
use
of
multiple
algorit
hms
to
solv
e
comple
x
planning
problems
Full-chain
consistenc
y
T
o
pro
vide
a
clear
and
intuiti
v
e
comparison
across
the
three
methodological
streams,
T
able
4
sum-
marizes
their
k
e
y
characterist
ics
in
terms
of
real-time
performance,
adaptabilit
y
to
dynamics,
multi-constraint
handling,
scalability
,
and
computational
cost.
T
able
4.
Comparati
v
e
summary
of
path
planning
methods
for
U
A
V
sw
arms
Algorithm
class
Real-time
perfor
-
mance
Adaptability
to
dy-
namics
Multi-constraint
han-
dling
Scalability
Computational
cost
Classical
search
High
Lo
w
to
medium
High
Lo
w
Medium
to
high
Intelligent
optimization
Lo
w
Medium
High
Medium
High
Deep
reinforcement
learning
High
High
Medium
High
Lo
w
Be
yond
the
qualitati
v
e
comparison
in
T
able
4,
e
xisting
studies
indicate
that
classical
search
algorithm
s
generally
achie
v
e
f
aster
deterministic
con
v
er
gence
in
lo
w-dimensional
static
en
vironments,
whereas
intelligent
optimization
methods
pro
vide
superior
global
e
xploration
capability
under
coupled
constraints.
Deep
reinforce-
ment
learning
approaches
demonstrate
stronger
adaptability
and
scalability
in
highly
dynamic
en
vironments,
although
their
training
comple
xity
and
simulation-to-reality
transfer
remain
signicant
challenges.
Therefore,
dif
ferent
planning
methods
e
xhibit
distinct
trade-of
fs
in
computational
ef
cienc
y
,
adaptability
,
and
cooperati
v
e
consistenc
y
maintenance.
Consensus-based
path
planning
for
U
A
V
swarms
under
multiple
constr
aints:
A
r
e
vie
w
(Y
ana
Lu)
Evaluation Warning : The document was created with Spire.PDF for Python.
630
❒
ISSN:
2722-2586
3.3.
Deep
r
einf
or
cement
lear
ning-based
methods
Deep
reinforcement
learning
(DRL)
has
become
a
k
e
y
approach
for
U
A
V
path
planning
in
unkno
wn
and
dynamic
en
vironments
due
to
its
adaptability
and
abi
lity
to
handle
high-dimensional
continuous
states.
T
o
address
dynamic,
multi-modal,
and
communication-constrained
scenarios,
researchers
enhance
DRL
with
im-
pro
v
ed
re
w
ard
functions,
e
xploration
strate
gies,
and
en
vironment-coupled
models,
impro
ving
path
continuity
,
decision
stability
,
and
en
vironmental
consistenc
y
.
Early
DRL
studies
introduced
consistenc
y
constraints
using
h
ybrid
approaches.
In
[88]
proposed
rey
algorithm-enhanced
deep
Q-netw
ork
(F
ADQN),
whi
ch
incorporated
rey-inspired
guidance
into
the
DRL
frame
w
ork.
By
simulating
rey
attraction,
it
pro
vided
directional
action
preferences
to
the
Q-netw
ork,
producing
smoother
and
more
continuous
paths
during
initial
training,
marking
a
shift
from
free
e
xploration
to
geometric
consistenc
y
.
As
research
mo
v
ed
from
geometric
features
to
dynamic
en
vironmental
consistenc
y
,
study
[89]
proposed
instructed
reinforcement
Q-l
earning
algorithm
(IR-QLA),
embedding
e
n
vi
ronmental
signal
strength
into
the
re
w
ard.
This
instructed
reinforcement
allo
ws
U
A
V
sw
arms
to
adjust
paths
in
real
time,
main-
taining
dynamic
compli
ance
with
temporal
and
en
vironmental
changes,
e
v
olving
from
stat
ic
state
consistenc
y
to
real-time
en
vironmental
consistenc
y
.
T
o
address
DRL
limitations
lik
e
slo
w
training
and
local
optima,
study
[90]
int
roduced
a
cumulati
v
e
re
w
ard
model
and
re
gion
se
gmentation
mechanism.
The
cumulati
v
e
re
w
ard
decomposes
path
progress
and
en-
vironmental
sparsity
into
continuous
re
w
ards,
impro
ving
temporal
consistenc
y
and
controllabil
ity
,
while
re
gion
se
gmentation
imposes
soft
global
constraints
to
pre
v
ent
local
loops
and
enhance
polic
y
stability
.
In
[91]
pro-
posed
the
constrained-interfered
uid
dynamical
system
(C-IFDS)
algorithm,
inte
grating
U
A
V
kinematics
and
constraints
into
a
DRL-based
reacti
v
e
disturbance
planning
frame
w
ork,
producing
high-quality
,
lo
w-conict
trajectories
with
maintained
trackability
and
consistenc
y
.
This
stage
reects
DRL
’
s
e
v
olution
from
ra
pid
con-
v
er
gence
to
global
consistenc
y
,
enhancing
rob
ustness
in
comple
x
en
vironments.
W
ith
the
inte
gration
of
unmanned
systems
and
communication
netw
orks,
consistenc
y
constraints
ha
v
e
e
v
olv
ed
to
w
ard
real-w
orld
modeling.
In
[92]
combined
dueling
double
DQN
(D3QN)
with
simultaneous
na
v-
ig
ation
and
radio
mapping
(SN
ARM)
to
b
uild
an
online
3D
radio
map,
using
urban
communication
channels
as
ph
ysical
consistenc
y
constraints.
This
allo
ws
U
A
V
sw
arms
to
coordinate
obstacle
a
v
oidance,
signal
quality
,
and
ight
time.
As
illustrated
in
Figure
4,
e
xisting
U
A
V
sw
arm
path
planning
approaches
can
be
broadly
c
lassied
into
three
cate
gories:
classical
path
search–based
methods,
intelligent
optimization–based
methods,
and
ma-
chine
learning
and
deep
reinforcement
learning
(DRL)–based
methods.
DRL
methods
of
fer
high
adaptability
and
scalability
for
dynamic
en
vironments,
handling
high-dimensional
state
spaces
and
decentralized
e
x
ecution
in
T
able
5.
Ho
we
v
er
,
the
y
suf
fer
from
high
training
costs,
poor
sample
ef
cienc
y
,
simulation-to-reality
g
aps,
and
lack
of
safety/consistenc
y
guarantees.
T
o
address
these
limitations,
h
ybrid
frame
w
orks
combining
classical
search,
intelligent
optimization,
and
DRL
ha
v
e
emer
ged
le
v
eraging
interpretabilit
y
,
global
search,
and
adapt-
ability
.
While
a
systematic
h
ybrid
solution
remains
an
open
challenge,
it
is
a
promising
direction
for
balancing
real-time
performance,
optimality
,
and
multi-constraint
consistenc
y
.
Figure
4.
Consistenc
y
constraint
method
IAES
Int
J
Rob
&
Autom,
V
ol.
15,
No.
3,
September
2026:
621–638
Evaluation Warning : The document was created with Spire.PDF for Python.