Inter
national
J
our
nal
of
Electrical
and
Computer
Engineering
(IJECE)
V
ol.
16,
No.
5,
October
2026,
pp.
2575
∼
2594
ISSN:
2088-8708,
DOI:
10.11591/ijece.v16i5.pp2575-2594
❒
2575
Rob
ust
r
esour
ce
allocation
in
multi-cell
UE-specic
RIS-assisted
D2D
r
elay
netw
orks
under
imperfect
CSI
Kay
ode
P
opoola
1
,
A
y
odeji
Ajani
2
,
Stuart
Nicholson
1
,
Muheeb
Ahmed
1
,
Srilatha
Narayangari
P
amuri
1
,
Ibrahim
Bala
Alhassam
3
1
Department
of
Computer
Science,
Dyson
Institute
of
Engineering
and
T
echnology
,
Malmesb
ury
,
W
iltshire,
United
Kingdom
2
Department
of
Computing,
Uni
v
ersity
of
Greater
Manchester
,
Manchester
,
United
Kingdom
3
Department
of
Electrical
Engineering,
Co
v
entry
Uni
v
ersity
,
Co
v
entry
,
United
Kingdom
Article
Inf
o
Article
history:
Recei
v
ed
May
18,
2026
Re
vised
Jun
15,
2026
Accepted
Aug
19,
2026
K
eyw
ords:
Channel
state
information
De
vice-to-de
vice
communication
Multi-cell
coordination
Recongurable
intelligent
surf
aces
Resource
allocation
Sum
spectral
ef
cienc
y
ABSTRA
CT
De
vice-to-de
vice
(D2D)
communication
enhances
spectral
ef
cienc
y
b
ut
re-
mains
constrained
by
l
imited
transmission
range,
underlay
interference,
and
the
half-duple
x
o
v
erhead
of
con
v
entional
relays.
User
equipment-specic
re-
congurable
intelligent
surf
aces
(UE-RIS)
of
fer
a
promising
alternati
v
e
by
en-
abling
passi
v
e
beamforming
to
strengthen
D2D
links
without
additional
spec-
trum
consumption.
Ho
we
v
er
,
e
xisting
studies
typically
assume
perfect
chan-
nel
state
informat
ion
(CSI)
and
single-cell
operation,
limiting
their
applicabil-
ity
to
practical
deplo
yments.
This
paper
proposes
a
rob
ust
multi-cell
resource
allocation
(RMRA)
frame
w
ork
for
UE-RIS-assisted
D2D
relay
netw
orks
un-
der
imperfect
CSI.
A
h
ybrid
uncertainty
model
is
adopted,
combining
statisti-
cal
Gauss-Mark
o
v
CSI
errors
for
intra-cell
links
with
bounded
norm-ball
un-
certainty
for
inter
-cell
links.
The
joint
optimisation
of
resource
reuse,
trans-
mit
po
wer
allocation,
and
RIS
phase
conguration
is
formulated
as
a
stochas-
tic
mix
ed-inte
ger
nonlinear
program
that
maximises
netw
ork
spectral
ef
cienc
y
while
satisfying
outage
and
quality-of-service
constraints.
T
o
ef
ciently
solv
e
the
problem,
a
three-stage
algorithm
is
proposed
comprising
distance-pruned
Hung
arian
assignment
,
rob
ust
po
wer
control
using
Bernstein-type
inequality
and
S-procedure
based
semidenite
programming,
and
soft
actor
-critic
(SA
C)
based
passi
v
e
beamforming.
Simulation
results
sho
w
that
RMRA
achie
v
es
a
94%
D2D
access
rate
at
light
load
and
o
v
er
75%
at
ful
l
load,
impro
v
es
sum
spectral
ef-
cienc
y
by
34.7%
and
70.2%
o
v
er
AF
relaying
and
direct
D2D,
respecti
v
ely
,
attains
118.5
bits/s/Hz/W
ener
gy
ef
cienc
y
,
and
maintains
30.2
bits/s/Hz
under
se
v
ere
CSI
uncertainty
.
This
is
an
open
access
article
under
the
CC
BY
-SA
license
.
Corresponding
A
uthor:
Kayode
Popoola
Department
of
Computer
Science,
Dyson
Institute
of
Engineering
and
T
echnology
Malmesb
ury
,
United
Kingdom
Email:
kayode.popoola@dysoninstitute.ac.uk
1.
INTR
ODUCTION
The
paradigm
shift
in
wireless
communication
from
fth-generation
(5G)
to
sixth-generation
(6G)
netw
orks
is
dri
v
en
by
the
need
for
h
yper
-connecti
vity
,
ultra-reliable
lo
w-latenc
y
communications
(URLLC),
and
e
xtreme
connection
densities.
While
5G
achie
v
es
signicant
g
ains
in
spectral
ef
cienc
y
through
mas-
si
v
e
MIMO
and
millimetre-w
a
v
e
(mmW
a
v
e)
technologies,
6G
requires
more
adv
anced
resource
management
strate
gies
to
handle
heterogeneous
traf
c
and
stringent
quality-of-service
(QoS)
requirements
[1],
[2].
De
vice-
J
ournal
homepage:
http://ijece
.iaescor
e
.com
Evaluation Warning : The document was created with Spire.PDF for Python.
2576
❒
ISSN:
2088-8708
to-de
vice
(D2D)
communication,
which
enables
direct
sidelinks
between
user
equipments
(UEs)
by
bypassing
the
base
station
(BS),
remains
a
cornerstone
of
this
e
v
olution
[3].
By
f
acilitating
localised
data
e
xchange,
D2D
can
enhance
area
spectral
ef
cienc
y
,
reduce
end-to-end
latenc
y
,
and
lo
wer
terminal
po
wer
consumption.
A
further
constraint
underlying
all
subsequent
design
choices
is
the
channel
coherence
time,
which
in
dense
6G
deplo
yments
at
mmW
a
v
e-adjacent
frequencies
can
be
on
the
order
of
a
fe
w
milliseconds
[4].
An
y
multi-cell
coordination
strat
e
gy
,
including
channel
state
information
(CSI)
e
xchange
o
v
er
backhaul,
combina-
torial
reuse
assignment,
rob
ust
po
wer
control,
and
RIS
phase
optimisation,
must
therefore
complete
within
this
windo
w
to
remain
v
alid
for
the
channel
realisat
ion
it
w
as
computed
for
.
Ex
ecuting
such
multi-cell
interference
coordination
within
this
constrained
windo
w
introduces
se
v
ere
latenc
y
bottlenecks,
necessitating
optimisation
algorithms
that
can
operate
with
minimal
online
computational
o
v
erhead.
Ho
we
v
er
,
the
lar
ge-scale
inte
gration
of
D2D
communications
into
cellular
netw
orks
f
aces
se
v
eral
sig-
nicant
technical
challenges.
First,
direct
sidelink
communication
is
highly
susceptible
to
se
v
ere
path
loss
and
en
vironmental
blockage,
thereby
limiting
the
ef
fecti
v
e
range
and
rel
iability
of
high-rate
transmissions
[5].
Although
con
v
entional
cooperati
v
e
relaying
techniques,
such
as
amplify-and-forw
ard
(AF)
and
decode-and-
forw
ard
(DF),
can
e
xtend
co
v
erage,
the
y
typically
operate
under
half-duple
x
constraints.
This
constraint
im-
poses
a
pre-log
spectral
ef
cienc
y
penalty
because
orthogonal
time
or
frequenc
y
resources
are
required
for
signal
forw
arding.
Furthermore,
when
D2D
groups
operate
in
an
underlay
mode
by
sharing
uplink
resources
with
cel
lular
users
(CUs),
the
y
introduce
comple
x
intra-cell
and
inter
-cell
interference
(ICI)
patterns
that
can
impact
the
stability
of
the
entire
multi-tier
netw
ork
[6],
[7].
Recongurable
intelligent
surf
aces
(RIS)
ha
v
e
recently
emer
ged
as
a
transformati
v
e
and
cost-
ef
fecti
v
e
technology
for
shaping
the
wireless
propag
ation
en
vironment
[8],
[9].
By
emplo
ying
an
array
of
nearly
passi
v
e
reecting
elements
to
manipulate
the
phase
of
incident
electromagnetic
w
a
v
es,
RIS
can
establish
virtual
line-of-
sight
(LoS)
links
and
mitig
ate
co-channel
interference.
A
promising
deplo
yment
strate
gy
in
v
olv
es
UE-specic
RIS,
where
compact
RIS
modules
are
inte
grated
with
relay-capable
user
equipments
(UEs)
to
enhance
D2D
sidelink
communications
without
i
ncurring
the
hardw
are
comple
xity
and
ener
gy
consumption
associated
with
acti
v
e
radio-frequenc
y
(RF)
chains
[10].
Although
RIS-assisted
transmission
does
not
pro
vide
t
rue
full-duple
x
relaying,
it
can
reduce
the
spectral-ef
cienc
y
loss
associated
with
con
v
ent
ional
half-duple
x
relay
operation
when
passi
v
e
reection
alone
is
suf
cient
to
maintain
link
connecti
vity
.
Despite
the
potential
of
RIS-assisted
D2D,
tw
o
critical
research
g
aps
persist.
First,
e
xisting
li
terature
focuses
on
single-cell
scenarios,
thereby
ne
glecting
the
inter
-cell
coordination
challenges
that
arise
in
dense
multi-cell
6G
deplo
yments.
In
such
en
vironments,
uncoordinated
resource
reuse
can
lead
to
se
v
ere
ICI
at
neighbouring
BSs.
Second,
perfect
CSI
is
often
assumed
for
analytical
con
v
enience,
although
it
is
practically
dif
cult
to
achie
v
e.
RIS
channels
are
particularly
challenging
to
estimate
because
the
y
in
v
olv
e
cascaded
f
ading
coef
cients
and
the
surf
aces
themselv
es
typically
lack
acti
v
e
sensing
hardw
are.
Consequently
,
CSI
acquired
via
pilot
signalling
or
backhaul
e
xchange
contains
non-ne
gligible
uncertainty
that
must
be
accounted
for
to
ensure
rob
ust
netw
ork
operation.
In
this
paper
,
we
address
these
challenges
by
proposing
a
rob
ust
multi-cell
resource
allocation
(RMRA)
frame
w
ork
for
UE-specic
RIS-assisted
D2D
relay
netw
orks.
Our
paper
jointly
addresses
three
interlink
ed
challenges:
i)
multi-cell
coordination
to
manage
ICI,
ii)
rob
ustness
ag
ainst
h
ybrid
statistical
and
bounded
CSI
imperfections,
and
iii)
mitig
ation
of
the
half-duple
x
penalty
through
passi
v
e
RIS-assisted
relaying.
Our
h
ybrid
uncertainty
model
treats
intra-cell
CSI,
acquired
through
uplink
pilots,
as
subject
to
Gaussian
estimation
er
-
rors
via
a
Gauss-Mark
o
v
model,
while
inter
-cell
CSI,
e
xchanged
o
v
er
limited-capacity
backhaul,
is
modelled
through
norm-ball
uncertainty
.
This
dual
treatment
captures
distinct
error
sources
in
a
ph
ysically
meaning-
ful
manner
.
Our
contrib
utions
are
fourfold:
we
model
a
realistic
h
ybrid
CSI
uncertainty
frame
w
ork,
formu-
late
a
rob
ust
sum
spectral
ef
cienc
y
maximisation
problem,
propose
an
algorithmic
decoupling
strate
gy
using
semidenite
programming,
and
implement
a
Soft
Actor
-Critic
reinforcement
learning
approach
for
passi
v
e
beamforming.
This
unied
frame
w
ork
aims
to
preserv
e
D2D
g
ains
without
compromising
cellular
reliability
under
stochastic
channel
uctuations
and
inter
-cell
interference.
The
optimisati
o
n
problem
is
a
stochastic
MINLP
and
is
NP-hard
in
general.
W
e
therefore
propose
a
three-stage
decomposition
that
separates
combinatorial
assignment,
continuous
po
wer
control,
and
non-
con
v
e
x
RIS
phase
optimisation.
The
rob
ust
po
wer
control
stage
is
cast
as
an
semidenite
program
(SDP)
using
the
Bernstein-type
inequality
for
probabilisti
c
CU
constraints
and
the
S-procedure
for
w
orst-case
D2D
constraints
[11],
[12].
The
RIS
phases
are
then
optimised
using
a
Soft
Actor
-Critic
deep
reinforcement
learning
agent,
which
is
well
suited
to
the
high-dimensional
continuous
action
space
and
unit-modulus
constraints
[13].
Int
J
Elec
&
Comp
Eng,
V
ol.
16,
No.
5,
October
2026:
2575-2594
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Elec
&
Comp
Eng
ISSN:
2088-8708
❒
2577
T
o
the
best
of
our
kno
wledge,
this
is
the
rst
w
ork
to
combine
such
a
h
ybrid
rob
ust
formulation
with
SA
C-based
passi
v
e
beamforming
in
a
multi-cell
D2D
underlay
scenario.
The
choice
of
SA
C
o
v
er
alternati
v
es
such
as
PPO,
DDPG,
or
con
v
entional
alternating
optimis
ation
(A
O)
is
moti
v
ated
by
three
f
actors
specic
to
the
RIS
phase-shift
design
problem.
First,
the
action
space
θ
∈
[0
,
2
π
)
N
K
is
high-dimensional
and
continuous;
A
O-based
phase
optimisation
typically
requires
per
-element
iterati
v
e
updates
with
comple
xity
that
scales
poorly
in
N
,
whereas
SA
C
produces
the
full
phase
v
ector
via
a
single
forw
ard
pass.
Second,
SA
C’
s
entrop
y-re
gularised
objecti
v
e
promotes
e
xploration
across
the
non-
con
v
e
x,
multi-modal
re
w
ard
landscape
induced
by
the
unit-modulus
constraints
and
multi-cell
interference
coupling,
reducing
the
risk
of
premature
con
v
er
gence
to
poor
local
optima
that
of
f-polic
y
deterministic
methods
such
as
DDPG
are
prone
to.
Third,
SA
C’
s
twin
Q-netw
ork
architecture
mitig
ates
the
o
v
erestimation
bias
that
destabilises
DDPG
training
under
the
highly
stochastic
re
w
ards
generated
by
Gauss-Mark
o
v
and
norm-ball
CSI
perturbations,
yielding
more
stable
con
v
er
gence
(section
5),
as
sho
wn
in
Figure
1.
T
r
a
i
n
i
n
g
E
p
i
s
o
d
e
s
0
100
200
300
400
500
600
700
800
900
1000
M
o
v
i
n
g
Av
e
r
a
g
e
R
e
w
a
r
d
(
S
u
m
S
E
)
0
5
10
15
20
25
30
35
40
45
RMRA (SAC)
Baseline (DDPG)
Figure
1.
DRL
con
v
er
gence
comparison:
mo
ving-a
v
erage
re
w
ard
(sum
SE)
v
ersus
training
episodes
for
the
proposed
SA
C
agent
(RMRA)
and
the
DDPG
baseline,
at
N
=
128
,
ρ
=
0
.
95
,
K
/
M
=
1
.
0
Recent
adv
ances
in
RIS-empo
wered
communications
ha
v
e
demonstrated
substantial
g
ains
in
ph
ysical-
layer
security
,
ener
gy
ef
cienc
y
,
and
non-orthogonal
multiple
access
[14],
[16].
Ho
we
v
er
,
much
of
this
litera-
ture
assumes
idealised
propag
ation
conditions
and
perfect
CSI,
which
limits
its
applicability
in
dense
multi-cell
deplo
yments.
In
parallel,
early
w
ork
on
D2D
underlay
systems
primarily
addressed
po
wer
control
and
mode
selection
under
perfect
CSI,
typically
within
a
single-cell
setting
[17],
[18].
F
or
RIS-assisted
D2D,
man
y
stud-
ies
ha
v
e
similarly
focused
on
simplied
single-cell
or
isolated-link
scenarios,
where
inter
-cell
interference
is
ignored
to
retain
analytical
tractability
[19].
A
gro
wing
body
of
w
ork
has
be
gun
to
e
xamine
interference
in
RIS-assisted
communications
more
carefully
.
These
studies
sho
w
that
performance
can
be
highly
sensiti
v
e
to
co-channel
interference,
b
ut
most
of
them
focus
on
the
interference
e
xperienced
at
the
user
side
rather
than
the
joint
ef
fect
of
interference
at
both
the
RIS
and
the
user
[20].
F
or
e
xample,
recent
analyses
ha
v
e
considered
interference-limited
RIS-aided
cellular
and
relaying
systems,
RIS-aided
mix
ed
optical/RF
links,
and
RIS-assisted
do
wnlink
scenarios
with
co-channel
interferers
[21],
[23].
Related
studies
on
RIS-aided
D2D
communications
ha
v
e
also
in
v
estig
ated
interference
between
cellular
and
D2D
transmissions,
as
well
as
interference
from
competing
D2D
links,
sometimes
with
learning-based
optimis
ation
of
po
wer
and
phase
shifts
[24],
[25].
Ne
v
ertheless,
these
contrib
uti
ons
generally
remain
limited
to
single-cell
or
weakly
coupled
deplo
yments,
and
the
y
do
not
fully
capture
the
spatially
selec-
ti
v
e
interference
coupling
introduced
by
RIS
beamforming
in
multi-cell
netw
orks.
Rob
ust
r
esour
ce
allocation
in
multi-cell
...
(Kayode
P
opoola)
Evaluation Warning : The document was created with Spire.PDF for Python.
2578
❒
ISSN:
2088-8708
While
recent
multi-cell
RIS-ass
isted
designs,
such
as
the
frame
w
ork
proposed
in
[26],
ef
fecti
v
ely
in-
corporate
inter
-cell
interference
into
the
optimisation
objecti
v
e,
the
y
critically
f
ail
to
account
for
the
se
v
ere
multi-cell
interference
coupling
that
occurs
when
backhaul
e
xchanged
CSI
is
corrupted
by
bounded
quantisa-
tion
errors.
Our
frame
w
ork
address
es
this
e
xact
limitation
by
embedding
norm-ball
uncertainty
models
directly
into
the
multi-cell
resource
allocation
constraints.
Moti
v
ated
by
these
limitations,
our
w
ork
considers
a
more
realistic
h
ybrid
uncertainty
model
that
treats
intra-cell
links
using
statistical
Gauss-Mark
o
v
uncertainty
and
inter
-cell
links
using
bounded
norm-ball
uncertainty
.
Unlik
e
prior
single-cell
RIS-D2D
designs
or
multi-cell
schemes
with
perfect
CSI,
we
jointly
address
resource
reuse,
rob
ust
po
wer
allocation,
and
RIS
phase
design
in
a
coordinated
multi-cell
underlay
netw
ork.
This
allo
ws
us
to
capture
both
the
ef
fect
of
interference
at
the
user
and
the
reected
coupling
through
the
RIS,
while
ensuring
tractable
optimisation
through
BTI-,
S-procedure-,
and
SA
C-based
decomposition.
The
principal
contrib
utions
are
summarised
as
follo
ws:
−
Hybrid
CSI
uncert
ainty
modelling:
W
e
de
v
elop
a
multi-cell
uplink
underlay
system
model
that
simultane-
ously
captures
intra-cell
and
inter
-cell
interference.
Distinguishing
our
w
ork
from
single-ti
er
models,
we
adopt
a
h
ybrid
uncertainty
frame
w
ork
where
intra-cell
pilot-based
CSI
is
modelled
using
statistical
Gauss-
Mark
o
v
uncertainty
,
while
inter
-cell
backhaul-based
CSI
is
characterised
by
bounded
norm-ball
uncertainty
.
−
Rob
ust
joint
optim
isation
frame
w
ork:
W
e
formulate
a
stochastic
MINLP
and
propose
a
tractable
solution
using
an
alternating
optimiSation
frame
w
ork.
W
e
handle
w
orst-case
inter
-cell
CSI
errors
using
the
S-
Procedure
and
probabilistic
intra-cell
errors
using
the
Bernstein-T
ype
Inequality
,
transforming
them
into
linear
matrix
inequalities.
−
Three-stage
algorithmic
decoupling:
T
o
address
the
computational
intractability
of
the
MINLP
,
we
propose
a
three-stage
RMRA
algorithm
consisting
of:
i)
a
distance-pruned
Hung
arian-based
assignment
strate
gy
,
ii)
a
rob
ust
po
wer
-control
stage
using
the
BTI
for
Gaussian
uncertainty
and
the
S-procedure
for
bounded
un-
certainty
,
both
reformulated
as
tractable
semidenite
programming
(SDPs),
and
iii)
a
passi
v
e
beamforming
stage
optimised
through
a
SA
C
deep
reinforcement
learning
agent.
−
Numerical
v
alidation:
The
proposed
frame
w
ork
achie
v
es
a
34.7%
impro
v
ed
sum-rate
than
non-rob
ust
AF
baselines
and
maintains
a
94%
D2D
access
rate
under
light
loads.
Furthermore,
we
demonstrate
that
RMRA
preserv
es
stri
ct
rob
ustness,
e
xceeding
the
performance
of
perfect-CSI
acti
v
e
relaying
e
v
en
under
se
v
ere
channel
estimation
errors.
The
combination
of
underlay
D2D
netw
orks,
multi-cell
frequenc
y
reuse,
and
UE-specic
RIS
creates
a
uniquely
se
v
ere
interfe
rence
en
vironment.
While
UE-specic
RIS
can
theoretically
reco
v
er
the
half-duple
x
penalty
through
passi
v
e
beamforming,
optimising
these
highly
directional
reecti
v
e
beams
can
cause
se
v
ere,
uncontrolled
interference
leakage
to
neighbouring
cells.
Addressing
this
specic
challenge
is
the
primary
moti
v
ation
for
our
rob
ust
multi-cell
frame
w
ork.
The
remainder
of
this
paper
is
structured
as
follo
ws.
Section
2
details
the
multi-cell
system
archi-
tecture,
the
channel
propag
ation
models,
and
the
h
ybrid
CSI
uncertainty
frame
w
ork.
Section
3
presents
the
mathematical
formulation
of
the
rob
ust
joint
optimisati
on
problem.
Section
4
de
v
elops
the
proposed
three-
stage
RMRA
algorithm,
including
the
SDP
reformulations
and
the
RL
agent
architecture.
Section
5
presents
the
numerical
results
and
performance
comparisons.
Section
6
dis
cusses
practical
limitations
and
practical
implementation
considerations,
and
section
7
concludes
the
paper
.
Boldf
ace
lo
wercase
and
uppercase
letters
represent
v
ectors
and
matrices,
respecti
v
ely
.
The
operators
(
·
)
T
,
(
·
)
∗
,
and
(
·
)
H
denote
the
transpose,
conjug
ate,
and
Hermitian
transpose,
respecti
v
ely
.
The
set
of
comple
x
m
×
n
matrices
is
denoted
by
C
m
×
n
.
W
e
use
∥
·
∥
for
the
Euclidean
norm
and
|
·
|
for
the
modulus
of
a
scalar
.
A
comple
x
circularly
symmetric
Gaussian
distrib
ution
with
mean
µ
and
co
v
ariance
Σ
is
denoted
as
C
N
(
µ
,
Σ
)
.
The
notation
A
⪰
0
indicates
that
A
is
Hermitian
positi
v
e
semidenite,
and
diag
(
·
)
denotes
a
diagonal
matrix.
The
e
xpectation
operator
is
represented
by
E
[
·
]
.
2.
SYSTEM
AND
CHANNEL
MODEL
W
e
consider
a
coordinated
multi-cell
uplink
netw
ork
as
sho
wn
in
Figure
2,
comprising
J
base
stations
(BSs),
denoted
by
the
set
B
=
{
B
1
,
.
.
.
,
B
J
}
,
arranged
in
a
re
gular
he
xagonal
layout.
Each
BS
B
j
serv
es
M
j
cellular
users
(CUs),
C
j
=
{
C
j
,
1
,
.
.
.
,
C
j
,M
j
}
,
which
are
distrib
uted
uniformly
within
the
corresponding
V
oronoi
cell.
The
netw
ork
supports
K
de
vice-to-de
vice
(D2D)
groups,
D
=
{
D
1
,
.
.
.
,
D
K
}
,
which
operate
Int
J
Elec
&
Comp
Eng,
V
ol.
16,
No.
5,
October
2026:
2575-2594
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Elec
&
Comp
Eng
ISSN:
2088-8708
❒
2579
in
an
underlay
mode.
Each
D2D
group
k
∈
D
consists
of
a
source
S
k
,
a
destination
R
k
,
and
a
potential
relay
node
Π
k
.
The
relay
node
is
equipped
with
a
UE-specic
RIS
consisting
of
N
nearly
passi
v
e
reecting
elements.
BSs
are
interconnected
via
a
high-capacity
backhaul
link
to
f
acilitate
the
e
xchange
of
quantised
CSI
and
interference
coordination
data
within
the
channel
coherence
interv
al.
Figure
2.
Conceptual
system
model
of
a
multi-cell
UE-specic
RIS-assisted
D2D
system:
Source
S
k
transmits
to
Destination
R
k
via
a
direct
D2D
link
(red
dashed)
and
an
RIS-reected
path
through
the
UE-specic
RIS
with
N
passi
v
e
elements
(blue
solid).
The
uplink
cellular
user
C
j
,m
sends
to
its
serving
BS
B
j
(blue
arro
w).
Orange
dashed
arro
ws
denote
co-channel
interference
from
cellular
and
other
D2D
transmitters
to
w
ard
the
BS
and
the
D2D
recei
v
er
The
netw
ork
utilises
a
uni
v
ersal
frequenc
y
reuse
f
actor
of
1.
A
D2D
group
k
in
cell
j
that
reuses
the
resource
block
(RB)
of
CU
C
j
,m
e
xperiences
co-channel
interference
from
i)
the
intra-cell
CU
C
j
,m
,
ii)
co-channel
CUs
in
neighbouring
cells
j
′
̸
=
j
,
and
iii)
other
co-channel
D2D
groups.
Correspondingly
,
BS
B
j
e
xperiences
interference
from
all
co-channel
D2D
sources
and
inter
-cell
CUs
when
decoding
the
uplink
signal
from
its
associated
CU.
2.1.
D2D
operational
modes
T
o
maximise
spectral
ef
cienc
y
and
range,
each
D2D
group
k
dynamically
selects
one
of
three
trans-
mission
modes
based
on
the
instantaneous
link
quality
and
the
a
v
ailable
CSI:
−
Mode
0
(Direct
Sidelink):
The
source
S
k
transmits
directly
to
R
k
.
This
mode
is
selected
if
the
l
ink
distance
is
small
and
the
SINR
satises
the
required
QoS
threshold
without
assistance.
−
Mode
1
(single-slot
RIS-assisted):
The
relay
Π
k
utilises
its
UE-specic
RIS
to
reect
the
signal
from
S
k
to
R
k
.
This
occurs
in
a
single
time
slot,
thereby
a
v
oiding
the
half-duple
x
throughput
penalty
.
This
mode
is
preferred
for
its
high
ener
gy
ef
cienc
y
and
passi
v
e
nature.
−
Mode
2
(T
w
o-slot
AF
relay):
If
modes
0
and
1
are
insuf
cient,
Π
k
functions
as
an
acti
v
e
half-duple
x
amplify-and-forw
ard
relay
.
T
ransmission
is
completed
in
tw
o
slots,
with
reception
in
slot
1
and
forw
arding
in
slot
2.
The
achie
v
able
rate
in
this
mode
is
scaled
by
a
f
actor
of
1
/
2
to
account
for
the
half-duple
x
constraint.
Rob
ust
r
esour
ce
allocation
in
multi-cell
...
(Kayode
P
opoola)
Evaluation Warning : The document was created with Spire.PDF for Python.
2580
❒
ISSN:
2088-8708
2.2.
Lar
ge-scale
and
small-scale
fading
models
The
comple
x
baseband
channel
coef
ci
ent
between
an
y
tw
o
nodes
(
u,
v
)
is
modelled
as
h
u,v
=
p
β
u,v
˜
h
u,v
.
The
lar
ge-scale
g
ain
β
u,v
in
linear
po
wer
is
dened
as
[27]
β
u,v
=
β
0
L
−
α
u,v
10
ξ
u,v
10
,
(1)
where
L
u,v
is
the
distance,
α
is
the
path-loss
e
xponent,
and
β
0
is
the
path-loss
at
the
ref
erence
distance.
The
term
10
ξ
u,v
10
represents
log-normal
shado
wing
with
ξ
u,v
∼
N
(0
,
σ
2
sh
)
.
Note
that
β
u,v
is
a
po
wer
g
ain;
the
amplitude
scaling
is
p
β
u,v
,
ensuring
ph
ysical
consistenc
y
in
the
signal-le
v
el
model.
F
or
links
e
xhibiting
a
line-of-sight
(LoS)
component,
such
as
those
in
v
olving
the
RIS,
we
emplo
y
the
Rician
f
ading
model:
˜
h
u,v
=
r
K
R
1
+
K
R
˜
h
LoS
u,v
+
r
1
1
+
K
R
˜
h
NLoS
u,v
,
(2)
where
K
R
is
the
Rician
K
-f
actor
,
˜
h
LoS
u,v
is
the
deterministic
LoS
component,
and
˜
h
NLoS
u,v
∼
C
N
(0
,
1)
represents
the
Rayleigh
f
ading
component.
2.3.
Effecti
v
e
RIS-assisted
channel
Let
h
S
k
,
Π
k
∈
C
N
×
1
and
h
Π
k
,R
k
∈
C
N
×
1
represent
the
channels
from
the
source
to
the
RIS
and
from
the
RIS
to
the
destination,
respecti
v
ely
.
The
RIS
phase-shift
matrix
is
Φ
k
=
diag
(
e
j
θ
k
,
1
,
.
.
.
,
e
j
θ
k
,N
)
,
where
θ
k
,n
∈
[0
,
2
π
)
.
In
Mode
1,
the
equi
v
alent
end-to-end
channel
for
group
k
is
gi
v
en
by:
h
eq
k
=
h
S
k
,R
k
+
h
H
Π
k
,R
k
Φ
k
h
S
k
,
Π
k
,
(3)
where
h
S
k
,R
k
is
the
direct
link.
The
ef
fecti
v
e
channel
po
wer
g
ain
is
g
(1)
k
=
|
h
eq
k
|
2
.
The
passi
v
e
nature
of
the
RIS
implies
that
the
noise
at
the
RIS
is
ne
gligible
compared
to
the
recei
v
er
noise
at
R
k
.
2.4.
Hybrid
CSI
uncertainty
framew
ork
T
o
accurately
reect
ph
ysical
netw
ork
constraints,
we
e
xplicitly
dene
a
h
ybrid
error
model:
a
sta-
tistical
Gauss-Mark
o
v
model
for
intra-cell
pilot
estimation
errors,
and
a
bounded
norm-ball
model
to
handle
w
orst-case
quantisation
errors
o
v
er
inter
-cell
backhaul
links.
−
Statistical
Gaussian
uncertainty
(Intra-Cell):
Intra-cell
links
are
estimated
via
pilot
signalling,
where
esti-
mation
noise
is
the
dominant
error
source.
The
actual
channel
h
u,v
is
related
to
the
estimate
ˆ
h
u,v
via
the
Gauss-Mark
o
v
model
h
u,v
=
ρ
ˆ
h
u,v
+
p
1
−
ρ
2
e
u,v
,
e
u,v
∼
C
N
(0
,
σ
2
h
)
,
(4)
where
ρ
∈
[0
,
1]
is
the
correlation
coef
cient.
−
Bounded
uncertainty
(Inter
-Cell):
Inter
-cell
links
are
acquired
through
backhaul
e
xchange,
where
quantisa-
tion
and
signalling
delays
result
in
a
bounded
error
.
W
e
model
this
as
h
∈
H
in
ter
≜
{
ˆ
h
+
∆
h
:
∥
∆
h
∥
≤
ϵ
}
,
(5)
where
ϵ
is
the
uncertainty
radius.
This
allo
ws
for
a
rob
ust
w
orst-case
design
for
inter
-cell
interference
mitig
ation.
The
estimation
error
v
ariance
is
further
analysed
via
the
Cram
´
er
-Rao
bound
in
Appendix
C.
In
practical
multi-cell
deplo
yments,
the
uncertainty
radius
ϵ
is
not
treated
as
an
arbitrary
constant
b
ut
is
determined
by
the
dominant
sources
of
inter
-cell
CSI
de
gradation:
quantisation
of
the
backhaul-e
xchanged
channel
estimate
at
b
bits
per
coef
cient,
and
the
backhaul
e
xchange
latenc
y
τ
d
relati
v
e
to
the
cohere
n
c
e
time.
Concretely
,
we
set
this
as:
ϵ
2
=
2
−
b
+
κ
d
τ
d
∥
ˆ
h
∥
2
,
(6)
where
the
rst
term
captures
the
quantisation
error
oor
and
the
second
captures
channel
ageing
o
v
er
the
backhaul
delay
,
with
κ
d
a
Doppler
-dependent
scaling
constant.
This
formulation
ties
ϵ
directly
to
deplo
yable
hardw
are
parameters
(quantisation
resolution,
backhaul
latenc
y).
Int
J
Elec
&
Comp
Eng,
V
ol.
16,
No.
5,
October
2026:
2575-2594
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Elec
&
Comp
Eng
ISSN:
2088-8708
❒
2581
2.5.
K
ey
system
parameters
Representati
v
e
system
parameters
for
a
multi-cell
urban
6G
deplo
yment
are
summarised
in
T
able
1.
These
v
alues
are
chosen
in
accordance
wi
th
commonly
adopted
3GPP
channel
and
deplo
yment
models
for
urban
macro/micro
scenarios
[28],
reecting
realistic
propag
ation
conditions
and
hardw
are
constraints.
The
noise
po
wer
is
computed
as
σ
2
N
=
N
0
B
,
where
N
0
=
−
174
dBm/Hz
is
the
noise
po
wer
spectral
density
and
B
=
180
kHz
is
the
per
-resource
block
bandwidth.
T
able
1.
System
parameters
P
arameter
Symbol
V
alue/Range
Cell
radius
R
500
m
Maximum
direct
D2D
range
r
50
m
Carrier
frequenc
y
f
c
3.5
GHz
Per
-RB
bandwidth
B
180
kHz
Noise
po
wer
spectral
density
N
0
−
174
dBm/Hz
Shado
wing
standard
de
viation
σ
sh
8
dB
P
ath-loss
e
xponent
α
3.5
RIS
elements
per
UE-specic
RIS
N
25–250
Phase-shift
quantisation
b
3
bits
Maximum
CU/D2D
po
wer
P
C
max
,
P
D
max
23
dBm,
20
dBm
2.6.
Pilot
signalling
o
v
erhead
Acquisition
of
the
cascaded
RIS-assisted
channel
h
S
k
,
Π
k
,
h
Π
k
,R
k
∈
C
N
×
1
requires
orthogonal
pilot
transmissions
whose
duration
scales
linearly
with
the
number
of
RIS
elements
N
.
W
ithin
a
coherence
block
of
length
T
c
symbols,
the
pilot
o
v
erhead
consumes
τ
p
=
κN
symbols,
where
κ
is
the
per
-element
pilot
duration.
The
fraction
of
the
block
a
v
ailable
for
data
transmission
is
therefore
η
(
N
)
=
max
0
,
1
−
κN
T
c
.
(7)
This
pre-log
f
actor
is
applied
multiplicati
v
ely
to
all
RIS-assisted
(Mode
1)
achie
v
able
rates.
3.
PR
OBLEM
FORMULA
TION
3.1.
Decision
v
ariables
and
objecti
v
e
Let
X
=
[
x
(
j
)
k
,m
]
∈
{
0
,
1
}
K
×
M
j
×
J
denote
the
binary
reuse-assignment
tensor
,
where
x
(
j
)
k
,m
=
1
indicates
that
D2D
group
k
in
cell
j
reuses
the
RB
of
cellular
user
m
in
cell
j
.
W
e
dene
P
C
∈
R
P
j
M
j
+
and
P
D
∈
R
K
+
as
the
collections
of
CU
and
D2D
transmit
po
wers,
respecti
v
ely
,
and
Φ
=
{
Φ
k
}
K
k
=1
as
the
set
of
RIS
phase
congurations.
The
objecti
v
e
is
to
maximise
the
netw
ork
sum
spectral
ef
cienc
y
(SE),
formulated
as:
P
0
:
max
X
,
P
C
,
P
D
,
Φ
J
X
j
=1
M
j
X
m
=1
R
C
j
,m
+
K
X
k
=1
x
(
j
)
k
,m
R
D
k
!
(8)
s.t.
(C1)
to
(C5)
.
where
R
C
j
,m
and
R
D
k
represent
the
achie
v
able
rates
of
CU
C
j
,m
and
D2D
group
k
,
respecti
v
ely
.
3.2.
Rate
expr
essions
and
interfer
ence
modelling
The
achie
v
able
rate
for
CU
C
j
,m
at
its
associated
BS
B
j
is
gi
v
en
by
R
C
j
,m
=
log
2
(1
+
γ
j
,m,B
)
.
The
SINR
is:
γ
j
,m,B
=
P
C
j
,m
|
g
j
,m,B
|
2
P
K
k
=1
x
(
j
)
k
,m
P
D
k
|
h
k
,B
j
|
2
+
I
C
in
ter
,j
,m
+
σ
2
N
,
(9)
where
g
j
,m,B
denotes
the
desired
channel
g
ain
from
C
j
,m
to
B
j
.
The
aggre
g
ate
ICI
observ
ed
at
B
j
on
the
shared
RB
is:
I
C
in
ter
,j
,m
=
X
j
′
̸
=
j
P
C
j
′
,m
|
g
j
′
,m,B
j
|
2
+
X
k
′
x
(
j
′
)
k
′
,m
P
D
k
′
|
h
k
′
,B
j
|
2
!
,
(10)
Rob
ust
r
esour
ce
allocation
in
multi-cell
...
(Kayode
P
opoola)
Evaluation Warning : The document was created with Spire.PDF for Python.
2582
❒
ISSN:
2088-8708
which
captures
the
contrib
ution
of
co-channel
CUs
and
D2D
sources
from
neighbouring
cells.
F
or
Mode
1
D2D
groups,
the
achie
v
able
rate
is
R
D
,
(1)
k
=
log
2
(1
+
γ
(1)
k
)
,
where
γ
(1)
k
=
P
D
k
|
h
eq
k
|
2
I
in
tra
k
+
I
in
ter
k
+
σ
2
N
.
(11)
The
intra-cell
interference
is
I
in
tra
k
=
P
C
j
,m
|
h
j
,m,R
k
|
2
,
while
I
in
ter
k
is
t
he
sum
of
co-channel
interference
from
adjacent
cells.
T
o
account
for
the
pilot
signalling
o
v
erhead
required
for
CSI
acquisition,
let
T
c
denote
the
channel
coherence
block
length
and
τ
p
=
κN
denote
the
pilot
training
duration,
where
κ
is
a
constant
scaling
f
actor
.
The
ef
fecti
v
e
achie
v
able
rate
for
mode
1
D2D
groups
is
e
xpressed
as:
R
D
,
(1)
k
=
1
−
κN
T
c
log
2
1
+
γ
(1)
k
,
(12)
3.3.
Rob
ust
constraints
The
follo
wing
constraints
ensure
netw
ork
stability
and
QoS
under
h
ybrid
CSI
uncertainty:
(C1)
Probabilistic
CU
QoS:
Subject
to
Gauss-Mark
o
v
uncertainty
,
we
require
Pr
{
γ
j
,m,B
≥
γ
C
thr
}
≥
1
−
δ
C
for
all
j
,
m
.
(C2)
W
orst-case
D2D
QoS:
Under
bounded
uncertainty
H
in
ter
,
the
D2D
SINR
must
sat
isfy
min
∆
h
∈H
in
ter
γ
k
≥
γ
D
thr
for
all
k
.
(C3)
Po
wer
b
udgets:
0
≤
P
C
j
,m
≤
P
C
max
and
0
≤
P
D
k
≤
P
D
max
.
(C4)
RIS
hardw
are
constraints:
Phase
shifts
are
restricted
to
the
unit
circle
and
quantised
as
θ
k
,n
∈
{
2
π
ℓ/
2
b
}
2
b
−
1
ℓ
=0
.
(C5)
RB
uniquenes
s:
Each
CU
RB
can
be
reused
by
at
most
one
D2D
group,
and
each
group
reuses
at
most
one
RB.
Remark.
Problem
P
0
is
a
stochastic
non-con
v
e
x
MINLP
.
Its
NP-hardness
arises
from
i)
the
combinat
orial
nature
of
reuse
assignment,
ii)
the
multiplicati
v
e
coupling
between
transmit
po
wers
and
RIS
phase
v
ariables,
and
iii)
the
non-tractable
nature
of
probabilistic
and
semi-innite
w
orst-case
constraints.
4.
PR
OPOSED
R
OB
UST
MUL
TI-CELL
RESOURCE
ALLOCA
TION
ALGORITHM
T
o
address
the
stochastic
mix
ed-inte
ger
nonlinear
programme
in
P
0
ef
ciently
,
we
propose
a
three-
stage
decoupling
frame
w
ork.
This
approach
separates
the
discrete
reuse
assignment
from
the
continuous
po
wer
allocation
and
the
non-con
v
e
x
RIS
phase-shift
design.
The
h
i
gh-le
v
el
e
x
ecution
o
w
is
s
u
m
marised
in
Algo-
rithm
1.
Algorithm
1
.
Rob
ust
multi-cell
resource
allocation
(RMRA)
1:
Input:
Estimated
CSI
ˆ
h
,
parameters
ρ
,
ϵ
,
b
udgets
P
C
/D
max
.
2:
Initialise:
SA
C
actor/critic
netw
orks,
P
C
,
P
D
,
Φ
.
3:
f
or
each
coherence
block
T
c
do
4:
Obtain
intra-cell
ˆ
h
via
pilots
and
inter
-cell
ˆ
h
via
backhaul.
5:
//
Stage
1:
Distance-Pruned
Assignment
6:
Prune
pairs
(
k
,
m
)
where
lar
ge-scale
g
ain
β
R
k
,C
j
,m
>
β
max
.
7:
Compute
cost
matrix
[∆
χ
k
,m
]
and
solv
e
Hung
arian
matching
for
X
∗
.
8:
//
Stage
2:
Rob
ust
P
o
wer
Optimisation
(SDP)
9:
F
ormulate
LMIs
using
BTI
(Eq.
16)
and
S-Procedure
(Eq.
18).
10:
Solv
e
SDP
via
Block
Coordinate
Descent
to
yield
P
C
∗
,
P
D
∗
.
11:
//
Stage
3:
P
assi
v
e
Beamf
orming
(SA
C
Infer
ence)
12:
Observ
e
state
s
t
(current
channels,
po
wers,
ICI).
13:
Ex
ecute
SA
C
forw
ard
pass
to
obtain
continuous
phase
shifts
θ
.
14:
Quantise
θ
to
b
-bit
resolution
yielding
Φ
∗
.
15:
end
f
or
16:
Output:
X
∗
,
P
C
∗
,
P
D
∗
,
Φ
∗
.
Int
J
Elec
&
Comp
Eng,
V
ol.
16,
No.
5,
October
2026:
2575-2594
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Elec
&
Comp
Eng
ISSN:
2088-8708
❒
2583
4.1.
Stage
1:
Distance-pruned
Hungarian
assignment
T
o
manage
the
computational
comple
xity
of
the
multi-cell
assignment,
we
rst
prune
the
set
of
po-
tential
reuse
partners.
Rather
than
relying
on
a
geometric
distance
threshold
L
min
,
which
does
not
account
for
shado
wing
v
ariability
in
urban
en
vironments,
the
pruning
criterion
is
e
xpressed
in
terms
of
the
lar
ge-scale
channel
g
ain
β
u,v
dened
in
(1),
which
already
incorporates
both
pathloss
and
log-normal
shado
wing.
A
pair
(
k
,
m
)
is
considered
feasible
only
if
the
lar
ge-scale
isolation
satises
β
R
k
,C
j
,m
≤
β
max
,
(13)
where
β
max
is
a
x
ed
isolation
threshold
on
the
cross-link
g
ain.
Because
β
u,v
depends
on
L
u,v
,
α
,
and
the
realised
shado
wi
ng
term
ξ
u,v
,
this
criterion
adapts
to
local
shado
wing
conditions
rather
than
enforcing
a
x
ed
geometric
separation,
and
tw
o
D2D-CU
pairs
at
the
same
ph
ysical
distance
b
ut
wi
th
dif
ferent
shado
wing
reali-
sations
may
be
pruned
dif
ferently
.
F
or
all
feasible
pairs,
we
dene
the
assignment
cost
∆
χ
k
,m
as
the
estimated
system
sum-rate
increment:
∆
χ
k
,m
=
E
[
R
C
j
,m
+
R
D
k
]
shared
−
E
[
R
C
j
,m
]
unshared
,
(14)
where
the
e
xpected
rates
are
calculated
using
a
x
ed-po
wer
proxy
and
the
statistical
CSI
means.
This
formula-
tion
ensures
that
the
Hung
arian
algorithm
select
s
partners
that
maximise
the
mar
ginal
spectral
ef
cienc
y
.
The
comple
xity
of
this
stage
is
O
(
n
3
)
,
where
n
=
max(
K
,
M
j
)
,
which
is
manageable
for
real-time
6G
deplo
y-
ments.
4.2.
Stage
2:
Rob
ust
po
wer
contr
ol
via
BTI
and
S-pr
ocedur
e
W
ith
x
ed
reuse
part
ners,
Sta
g
e
2
optimises
transmit
po
wers
to
satisfy
the
h
ybrid
CSI
unce
rtainty
requirements.
4.2.1.
BTI
r
ef
ormulation
f
or
CU
r
eliability
The
probabilistic
constrai
n
t
(C1)
under
Gaussian
uncertainty
e
∼
C
N
(
0
,
σ
2
h
I
)
can
be
e
xpressed
in
the
quadratic
form
Pr
{
e
H
Qe
+
2
ℜ{
r
H
e
}
+
s
≤
0
}
≤
δ
C
.
(15)
T
o
render
this
tractable,
we
apply
the
one-sided
Bernstein-type
ine
q
ua
lity
,
yielding
the
follo
wing
deterministic
LMIs:
T
r(
Q
)
−
p
−
2
ln
δ
C
ν
+
ln
δ
C
µ
+
s
≥
0
,
(16)
v
ec(
Q
)
√
2
r
≤
ν
,
µ
I
+
Q
⪰
0
,
µ
≥
0
,
(17)
where
Q
,
r
,
and
s
are
linear
functions
of
P
C
j
,m
and
P
D
k
deri
v
ed
from
the
SINR
denominator
and
the
threshold
γ
C
thr
.
A
detailed
deri
v
ation
of
the
BTI-based
conic
reformulation
is
pro
vided
in
Appendix
A.
4.2.2.
S-Pr
ocedur
e
f
or
D2D
r
ob
ustness
F
or
the
bounded
inter
-cell
uncertainty
(C2),
we
require
the
D2D
SINR
to
hold
for
all
∥
∆
h
∥
≤
ϵ
.
This
semi-innite
constraint
is
transformed
using
the
S-lemma
into
a
single
LMI.
Specically
,
(C2)
is
satised
if
there
e
xists
a
scalar
λ
≥
0
such
that
A
+
λ
I
b
b
H
c
−
λϵ
2
⪰
0
,
(18)
where
A
,
b
,
and
c
represent
the
quadratic,
linear
,
and
constant
coef
cients
of
the
D2D
SINR
mar
gin,
respec-
ti
v
ely
.
The
resulting
SDP
is
solv
ed
globally
across
cells
via
BCD,
which
ensures
a
non-decreasing
objecti
v
e
sequence.
The
equi
v
alence
to
the
LMI
is
pro
v
ed
in
Appendix
B.
Rob
ust
r
esour
ce
allocation
in
multi-cell
...
(Kayode
P
opoola)
Evaluation Warning : The document was created with Spire.PDF for Python.
2584
❒
ISSN:
2088-8708
4.3.
Stage
3:
P
assi
v
e
beamf
orming
via
soft
actor
-critic
The
non-con
v
e
xity
of
the
RIS
phase-shift
optimisation,
compounded
by
the
multi-cell
interference
and
unit-modulus
constraints,
mak
es
con
v
entional
optimisation
techniques
dif
cult
t
o
apply
.
W
e
utilise
the
SA
C
algorithm,
which
emplo
ys
an
entrop
y-re
gularised
frame
w
ork
to
maintain
an
appropriate
balance
between
e
xploration
and
e
xploitation
in
continuous
action
spaces.
−
State
space
(
s
t
):
The
state
includes
the
estimated
channel
g
ains
ˆ
h
S
k
,
Π
k
,
ˆ
h
Π
k
,R
k
,
ˆ
h
S
k
,R
k
,
the
current
po
wer
le
v
els,
and
the
aggre
g
ate
ICI
po
wer
measured
at
the
BS.
−
Action
space
(
a
t
):
The
action
is
the
v
ector
of
continuous
phase
shifts
θ
∈
[0
,
2
π
)
N
K
,
which
are
mapped
to
the
b
-bit
resolution
grid
before
application.
−
Re
w
ard
function
(
r
t
):
The
re
w
ard
is
the
total
system
SE
penalised
by
a
weighted
term
Λ
for
an
y
violation
of
the
QoS
thresholds
γ
C
thr
or
γ
D
thr
:
r
t
=
X
j
,m
R
C
j
,m
+
X
k
R
D
k
−
Λ
⊮
{
QoS
violation
}
.
(19)
The
state
v
ector
has
dimension
|
s
t
|
=
2
N
K
+
K
+
J
,
comprising
the
real
and
imaginary
parts
of
the
three
estimated
channel
v
ectors
per
D2D
group
(stack
ed),
current
po
wer
allocations
{
P
D
k
}
,
and
the
per
-BS
aggre
g
ate
ICI
measurements.
The
action
v
ector
θ
∈
[0
,
2
π
)
N
K
has
dimension
N
K
.
The
penalty
weight
Λ
in
(16)
is
not
x
ed
a
priori;
it
is
annealed
during
training
according
to
Λ
(
t
)
=
Λ
0
1
+
t
T
anneal
,
(20)
starting
from
Λ
0
and
increasing
linearly
o
v
er
training
step
t
up
to
a
cap,
so
that
early
e
xploration
is
not
o
v
erly
constrained
by
QoS
penalties
while
later
training
increasingly
enforces
feasibility
.
Λ
0
and
T
anneal
are
tuned
via
grid
search
to
the
v
alues
reported
in
T
able
2
(
Λ
0
=
5
,
T
anneal
=
5
×
10
4
steps),
selected
as
the
conguration
yielding
the
lo
west
QoS-violation
rate
without
de
grading
the
con
v
er
ged
re
w
ard.
The
SA
C
implementation
utilises
twin
Q-netw
orks
to
mitig
ate
o
v
erestimation
bias
and
a
stochast
ic
actor
for
rob
ust
polic
y
learning.
Hyperparameters
used
for
the
e
xperimental
v
alidation
are
summarised
in
T
able
2.
T
able
2.
SA
C
agent
h
yperparameters
Hyperparameter
V
alue
Learning
rate
(actor
and
critic)
3
×
10
−
4
Discount
f
actor
(
ζ
)
0.99
Entrop
y
coef
cient
(
α
entrop
y
)
0.2
T
ar
get
smoothing
coef
cient
(
τ
)
0.005
Replay
b
uf
fer
size
10
6
Batch
size
256
Hidden
layers
2
layers,
256
nodes
each
Acti
v
ation
function
ReLU
4.4.
Complexity
considerations
T
o
strictly
adhere
to
millisecond-le
v
el
coherence
time
constraints,
the
iterati
v
e
e
x
ecution
of
Stage
2
and
Stage
3
is
restricted
to
the
of
ine
trai
n
i
ng
and
lar
ge-scale
f
ading
tracking
phases.
During
real-time
online
deplo
yment,
the
SA
C
agent
e
x
ecutes
a
single
forw
ard
inference
pass
O
(1)
,
and
the
transmit
po
wers
are
optimised
via
a
single
BCD
iteration
w
arm-started
by
the
pre
vious
coherence
block,
thereby
bypassing
latenc
y-
intensi
v
e
con
v
er
gence
loops.
The
computational
o
v
erhead
of
the
proposed
RMRA
frame
w
ork
is
anal
ysed
for
each
of
the
three
stages.
The
comple
xity
of
Stage
1
is
dominated
by
the
Hun
g
ari
an
algorithm,
which
scales
as
O
(max(
K
,
M
j
)
3
)
per
cell.
Stage
2
in
v
olv
es
solving
an
SDP
reformulated
via
BTI
and
the
S-procedure.
Using
primal-dual
interior
-point
methods,
the
w
orst-case
comple
xity
for
Stage
2
is
O
(
L
(
N
K
)
3
.
5
log
(1
/ϵ
tol
))
,
where
L
denotes
the
number
of
iterations
required
for
BCD
to
reach
a
stationary
point.
The
f
actor
(
N
K
)
3
.
5
reects
the
size
of
the
LMIs
in
the
SDP
.
F
or
Stage
3,
the
comple
xity
of
SA
C
is
concentrated
in
the
of
ine
training
phase,
which
scales
with
the
depth
and
width
of
the
neural
netw
orks
and
the
size
of
the
replay
b
uf
fer
.
During
online
deplo
yment,
the
RIS
phase
conguration
is
determined
via
a
single
forw
ard
pass
through
the
actor
netw
ork,
resulting
in
constant-time
inference
that
is
suitable
for
real-time
operation
within
the
channel
coherence
time.
Int
J
Elec
&
Comp
Eng,
V
ol.
16,
No.
5,
October
2026:
2575-2594
Evaluation Warning : The document was created with Spire.PDF for Python.