Inter national J our nal of Electrical and Computer Engineering (IJECE) V ol. 16, No. 5, October 2026, pp. 2575 ∼ 2594 ISSN: 2088-8708, DOI: 10.11591/ijece.v16i5.pp2575-2594 ❒ 2575 Rob ust r esour ce allocation in multi-cell UE-specic RIS-assisted D2D r elay netw orks under imperfect CSI Kay ode P opoola 1 , A y odeji Ajani 2 , Stuart Nicholson 1 , Muheeb Ahmed 1 , Srilatha Narayangari P amuri 1 , Ibrahim Bala Alhassam 3 1 Department of Computer Science, Dyson Institute of Engineering and T echnology , Malmesb ury , W iltshire, United Kingdom 2 Department of Computing, Uni v ersity of Greater Manchester , Manchester , United Kingdom 3 Department of Electrical Engineering, Co v entry Uni v ersity , Co v entry , United Kingdom Article Inf o Article history: Recei v ed May 18, 2026 Re vised Jun 15, 2026 Accepted Aug 19, 2026 K eyw ords: Channel state information De vice-to-de vice communication Multi-cell coordination Recongurable intelligent surf aces Resource allocation Sum spectral ef cienc y ABSTRA CT De vice-to-de vice (D2D) communication enhances spectral ef cienc y b ut re- mains constrained by l imited transmission range, underlay interference, and the half-duple x o v erhead of con v entional relays. User equipment-specic re- congurable intelligent surf aces (UE-RIS) of fer a promising alternati v e by en- abling passi v e beamforming to strengthen D2D links without additional spec- trum consumption. Ho we v er , e xisting studies typically assume perfect chan- nel state informat ion (CSI) and single-cell operation, limiting their applicabil- ity to practical deplo yments. This paper proposes a rob ust multi-cell resource allocation (RMRA) frame w ork for UE-RIS-assisted D2D relay netw orks un- der imperfect CSI. A h ybrid uncertainty model is adopted, combining statisti- cal Gauss-Mark o v CSI errors for intra-cell links with bounded norm-ball un- certainty for inter -cell links. The joint optimisation of resource reuse, trans- mit po wer allocation, and RIS phase conguration is formulated as a stochas- tic mix ed-inte ger nonlinear program that maximises netw ork spectral ef cienc y while satisfying outage and quality-of-service constraints. T o ef ciently solv e the problem, a three-stage algorithm is proposed comprising distance-pruned Hung arian assignment , rob ust po wer control using Bernstein-type inequality and S-procedure based semidenite programming, and soft actor -critic (SA C) based passi v e beamforming. Simulation results sho w that RMRA achie v es a 94% D2D access rate at light load and o v er 75% at ful l load, impro v es sum spectral ef- cienc y by 34.7% and 70.2% o v er AF relaying and direct D2D, respecti v ely , attains 118.5 bits/s/Hz/W ener gy ef cienc y , and maintains 30.2 bits/s/Hz under se v ere CSI uncertainty . This is an open access article under the CC BY -SA license . Corresponding A uthor: Kayode Popoola Department of Computer Science, Dyson Institute of Engineering and T echnology Malmesb ury , United Kingdom Email: kayode.popoola@dysoninstitute.ac.uk 1. INTR ODUCTION The paradigm shift in wireless communication from fth-generation (5G) to sixth-generation (6G) netw orks is dri v en by the need for h yper -connecti vity , ultra-reliable lo w-latenc y communications (URLLC), and e xtreme connection densities. While 5G achie v es signicant g ains in spectral ef cienc y through mas- si v e MIMO and millimetre-w a v e (mmW a v e) technologies, 6G requires more adv anced resource management strate gies to handle heterogeneous traf c and stringent quality-of-service (QoS) requirements [1], [2]. De vice- J ournal homepage: http://ijece .iaescor e .com Evaluation Warning : The document was created with Spire.PDF for Python.
2576 ❒ ISSN: 2088-8708 to-de vice (D2D) communication, which enables direct sidelinks between user equipments (UEs) by bypassing the base station (BS), remains a cornerstone of this e v olution [3]. By f acilitating localised data e xchange, D2D can enhance area spectral ef cienc y , reduce end-to-end latenc y , and lo wer terminal po wer consumption. A further constraint underlying all subsequent design choices is the channel coherence time, which in dense 6G deplo yments at mmW a v e-adjacent frequencies can be on the order of a fe w milliseconds [4]. An y multi-cell coordination strat e gy , including channel state information (CSI) e xchange o v er backhaul, combina- torial reuse assignment, rob ust po wer control, and RIS phase optimisation, must therefore complete within this windo w to remain v alid for the channel realisat ion it w as computed for . Ex ecuting such multi-cell interference coordination within this constrained windo w introduces se v ere latenc y bottlenecks, necessitating optimisation algorithms that can operate with minimal online computational o v erhead. Ho we v er , the lar ge-scale inte gration of D2D communications into cellular netw orks f aces se v eral sig- nicant technical challenges. First, direct sidelink communication is highly susceptible to se v ere path loss and en vironmental blockage, thereby limiting the ef fecti v e range and rel iability of high-rate transmissions [5]. Although con v entional cooperati v e relaying techniques, such as amplify-and-forw ard (AF) and decode-and- forw ard (DF), can e xtend co v erage, the y typically operate under half-duple x constraints. This constraint im- poses a pre-log spectral ef cienc y penalty because orthogonal time or frequenc y resources are required for signal forw arding. Furthermore, when D2D groups operate in an underlay mode by sharing uplink resources with cel lular users (CUs), the y introduce comple x intra-cell and inter -cell interference (ICI) patterns that can impact the stability of the entire multi-tier netw ork [6], [7]. Recongurable intelligent surf aces (RIS) ha v e recently emer ged as a transformati v e and cost- ef fecti v e technology for shaping the wireless propag ation en vironment [8], [9]. By emplo ying an array of nearly passi v e reecting elements to manipulate the phase of incident electromagnetic w a v es, RIS can establish virtual line-of- sight (LoS) links and mitig ate co-channel interference. A promising deplo yment strate gy in v olv es UE-specic RIS, where compact RIS modules are inte grated with relay-capable user equipments (UEs) to enhance D2D sidelink communications without i ncurring the hardw are comple xity and ener gy consumption associated with acti v e radio-frequenc y (RF) chains [10]. Although RIS-assisted transmission does not pro vide t rue full-duple x relaying, it can reduce the spectral-ef cienc y loss associated with con v ent ional half-duple x relay operation when passi v e reection alone is suf cient to maintain link connecti vity . Despite the potential of RIS-assisted D2D, tw o critical research g aps persist. First, e xisting li terature focuses on single-cell scenarios, thereby ne glecting the inter -cell coordination challenges that arise in dense multi-cell 6G deplo yments. In such en vironments, uncoordinated resource reuse can lead to se v ere ICI at neighbouring BSs. Second, perfect CSI is often assumed for analytical con v enience, although it is practically dif cult to achie v e. RIS channels are particularly challenging to estimate because the y in v olv e cascaded f ading coef cients and the surf aces themselv es typically lack acti v e sensing hardw are. Consequently , CSI acquired via pilot signalling or backhaul e xchange contains non-ne gligible uncertainty that must be accounted for to ensure rob ust netw ork operation. In this paper , we address these challenges by proposing a rob ust multi-cell resource allocation (RMRA) frame w ork for UE-specic RIS-assisted D2D relay netw orks. Our paper jointly addresses three interlink ed challenges: i) multi-cell coordination to manage ICI, ii) rob ustness ag ainst h ybrid statistical and bounded CSI imperfections, and iii) mitig ation of the half-duple x penalty through passi v e RIS-assisted relaying. Our h ybrid uncertainty model treats intra-cell CSI, acquired through uplink pilots, as subject to Gaussian estimation er - rors via a Gauss-Mark o v model, while inter -cell CSI, e xchanged o v er limited-capacity backhaul, is modelled through norm-ball uncertainty . This dual treatment captures distinct error sources in a ph ysically meaning- ful manner . Our contrib utions are fourfold: we model a realistic h ybrid CSI uncertainty frame w ork, formu- late a rob ust sum spectral ef cienc y maximisation problem, propose an algorithmic decoupling strate gy using semidenite programming, and implement a Soft Actor -Critic reinforcement learning approach for passi v e beamforming. This unied frame w ork aims to preserv e D2D g ains without compromising cellular reliability under stochastic channel uctuations and inter -cell interference. The optimisati o n problem is a stochastic MINLP and is NP-hard in general. W e therefore propose a three-stage decomposition that separates combinatorial assignment, continuous po wer control, and non- con v e x RIS phase optimisation. The rob ust po wer control stage is cast as an semidenite program (SDP) using the Bernstein-type inequality for probabilisti c CU constraints and the S-procedure for w orst-case D2D constraints [11], [12]. The RIS phases are then optimised using a Soft Actor -Critic deep reinforcement learning agent, which is well suited to the high-dimensional continuous action space and unit-modulus constraints [13]. Int J Elec & Comp Eng, V ol. 16, No. 5, October 2026: 2575-2594 Evaluation Warning : The document was created with Spire.PDF for Python.
Int J Elec & Comp Eng ISSN: 2088-8708 ❒ 2577 T o the best of our kno wledge, this is the rst w ork to combine such a h ybrid rob ust formulation with SA C-based passi v e beamforming in a multi-cell D2D underlay scenario. The choice of SA C o v er alternati v es such as PPO, DDPG, or con v entional alternating optimis ation (A O) is moti v ated by three f actors specic to the RIS phase-shift design problem. First, the action space θ ∈ [0 , 2 π ) N K is high-dimensional and continuous; A O-based phase optimisation typically requires per -element iterati v e updates with comple xity that scales poorly in N , whereas SA C produces the full phase v ector via a single forw ard pass. Second, SA C’ s entrop y-re gularised objecti v e promotes e xploration across the non- con v e x, multi-modal re w ard landscape induced by the unit-modulus constraints and multi-cell interference coupling, reducing the risk of premature con v er gence to poor local optima that of f-polic y deterministic methods such as DDPG are prone to. Third, SA C’ s twin Q-netw ork architecture mitig ates the o v erestimation bias that destabilises DDPG training under the highly stochastic re w ards generated by Gauss-Mark o v and norm-ball CSI perturbations, yielding more stable con v er gence (section 5), as sho wn in Figure 1. T r a i n i n g E p i s o d e s 0 100 200 300 400 500 600 700 800 900 1000 M o v i n g Av e r a g e R e w a r d ( S u m S E ) 0 5 10 15 20 25 30 35 40 45 RMRA (SAC) Baseline (DDPG) Figure 1. DRL con v er gence comparison: mo ving-a v erage re w ard (sum SE) v ersus training episodes for the proposed SA C agent (RMRA) and the DDPG baseline, at N = 128 , ρ = 0 . 95 , K / M = 1 . 0 Recent adv ances in RIS-empo wered communications ha v e demonstrated substantial g ains in ph ysical- layer security , ener gy ef cienc y , and non-orthogonal multiple access [14], [16]. Ho we v er , much of this litera- ture assumes idealised propag ation conditions and perfect CSI, which limits its applicability in dense multi-cell deplo yments. In parallel, early w ork on D2D underlay systems primarily addressed po wer control and mode selection under perfect CSI, typically within a single-cell setting [17], [18]. F or RIS-assisted D2D, man y stud- ies ha v e similarly focused on simplied single-cell or isolated-link scenarios, where inter -cell interference is ignored to retain analytical tractability [19]. A gro wing body of w ork has be gun to e xamine interference in RIS-assisted communications more carefully . These studies sho w that performance can be highly sensiti v e to co-channel interference, b ut most of them focus on the interference e xperienced at the user side rather than the joint ef fect of interference at both the RIS and the user [20]. F or e xample, recent analyses ha v e considered interference-limited RIS-aided cellular and relaying systems, RIS-aided mix ed optical/RF links, and RIS-assisted do wnlink scenarios with co-channel interferers [21], [23]. Related studies on RIS-aided D2D communications ha v e also in v estig ated interference between cellular and D2D transmissions, as well as interference from competing D2D links, sometimes with learning-based optimis ation of po wer and phase shifts [24], [25]. Ne v ertheless, these contrib uti ons generally remain limited to single-cell or weakly coupled deplo yments, and the y do not fully capture the spatially selec- ti v e interference coupling introduced by RIS beamforming in multi-cell netw orks. Rob ust r esour ce allocation in multi-cell ... (Kayode P opoola) Evaluation Warning : The document was created with Spire.PDF for Python.
2578 ❒ ISSN: 2088-8708 While recent multi-cell RIS-ass isted designs, such as the frame w ork proposed in [26], ef fecti v ely in- corporate inter -cell interference into the optimisation objecti v e, the y critically f ail to account for the se v ere multi-cell interference coupling that occurs when backhaul e xchanged CSI is corrupted by bounded quantisa- tion errors. Our frame w ork address es this e xact limitation by embedding norm-ball uncertainty models directly into the multi-cell resource allocation constraints. Moti v ated by these limitations, our w ork considers a more realistic h ybrid uncertainty model that treats intra-cell links using statistical Gauss-Mark o v uncertainty and inter -cell links using bounded norm-ball uncertainty . Unlik e prior single-cell RIS-D2D designs or multi-cell schemes with perfect CSI, we jointly address resource reuse, rob ust po wer allocation, and RIS phase design in a coordinated multi-cell underlay netw ork. This allo ws us to capture both the ef fect of interference at the user and the reected coupling through the RIS, while ensuring tractable optimisation through BTI-, S-procedure-, and SA C-based decomposition. The principal contrib utions are summarised as follo ws: − Hybrid CSI uncert ainty modelling: W e de v elop a multi-cell uplink underlay system model that simultane- ously captures intra-cell and inter -cell interference. Distinguishing our w ork from single-ti er models, we adopt a h ybrid uncertainty frame w ork where intra-cell pilot-based CSI is modelled using statistical Gauss- Mark o v uncertainty , while inter -cell backhaul-based CSI is characterised by bounded norm-ball uncertainty . − Rob ust joint optim isation frame w ork: W e formulate a stochastic MINLP and propose a tractable solution using an alternating optimiSation frame w ork. W e handle w orst-case inter -cell CSI errors using the S- Procedure and probabilistic intra-cell errors using the Bernstein-T ype Inequality , transforming them into linear matrix inequalities. − Three-stage algorithmic decoupling: T o address the computational intractability of the MINLP , we propose a three-stage RMRA algorithm consisting of: i) a distance-pruned Hung arian-based assignment strate gy , ii) a rob ust po wer -control stage using the BTI for Gaussian uncertainty and the S-procedure for bounded un- certainty , both reformulated as tractable semidenite programming (SDPs), and iii) a passi v e beamforming stage optimised through a SA C deep reinforcement learning agent. − Numerical v alidation: The proposed frame w ork achie v es a 34.7% impro v ed sum-rate than non-rob ust AF baselines and maintains a 94% D2D access rate under light loads. Furthermore, we demonstrate that RMRA preserv es stri ct rob ustness, e xceeding the performance of perfect-CSI acti v e relaying e v en under se v ere channel estimation errors. The combination of underlay D2D netw orks, multi-cell frequenc y reuse, and UE-specic RIS creates a uniquely se v ere interfe rence en vironment. While UE-specic RIS can theoretically reco v er the half-duple x penalty through passi v e beamforming, optimising these highly directional reecti v e beams can cause se v ere, uncontrolled interference leakage to neighbouring cells. Addressing this specic challenge is the primary moti v ation for our rob ust multi-cell frame w ork. The remainder of this paper is structured as follo ws. Section 2 details the multi-cell system archi- tecture, the channel propag ation models, and the h ybrid CSI uncertainty frame w ork. Section 3 presents the mathematical formulation of the rob ust joint optimisati on problem. Section 4 de v elops the proposed three- stage RMRA algorithm, including the SDP reformulations and the RL agent architecture. Section 5 presents the numerical results and performance comparisons. Section 6 dis cusses practical limitations and practical implementation considerations, and section 7 concludes the paper . Boldf ace lo wercase and uppercase letters represent v ectors and matrices, respecti v ely . The operators ( · ) T , ( · ) ∗ , and ( · ) H denote the transpose, conjug ate, and Hermitian transpose, respecti v ely . The set of comple x m × n matrices is denoted by C m × n . W e use ∥ · ∥ for the Euclidean norm and | · | for the modulus of a scalar . A comple x circularly symmetric Gaussian distrib ution with mean µ and co v ariance Σ is denoted as C N ( µ , Σ ) . The notation A ⪰ 0 indicates that A is Hermitian positi v e semidenite, and diag ( · ) denotes a diagonal matrix. The e xpectation operator is represented by E [ · ] . 2. SYSTEM AND CHANNEL MODEL W e consider a coordinated multi-cell uplink netw ork as sho wn in Figure 2, comprising J base stations (BSs), denoted by the set B = { B 1 , . . . , B J } , arranged in a re gular he xagonal layout. Each BS B j serv es M j cellular users (CUs), C j = { C j , 1 , . . . , C j ,M j } , which are distrib uted uniformly within the corresponding V oronoi cell. The netw ork supports K de vice-to-de vice (D2D) groups, D = { D 1 , . . . , D K } , which operate Int J Elec & Comp Eng, V ol. 16, No. 5, October 2026: 2575-2594 Evaluation Warning : The document was created with Spire.PDF for Python.
Int J Elec & Comp Eng ISSN: 2088-8708 ❒ 2579 in an underlay mode. Each D2D group k ∈ D consists of a source S k , a destination R k , and a potential relay node Π k . The relay node is equipped with a UE-specic RIS consisting of N nearly passi v e reecting elements. BSs are interconnected via a high-capacity backhaul link to f acilitate the e xchange of quantised CSI and interference coordination data within the channel coherence interv al. Figure 2. Conceptual system model of a multi-cell UE-specic RIS-assisted D2D system: Source S k transmits to Destination R k via a direct D2D link (red dashed) and an RIS-reected path through the UE-specic RIS with N passi v e elements (blue solid). The uplink cellular user C j ,m sends to its serving BS B j (blue arro w). Orange dashed arro ws denote co-channel interference from cellular and other D2D transmitters to w ard the BS and the D2D recei v er The netw ork utilises a uni v ersal frequenc y reuse f actor of 1. A D2D group k in cell j that reuses the resource block (RB) of CU C j ,m e xperiences co-channel interference from i) the intra-cell CU C j ,m , ii) co-channel CUs in neighbouring cells j ′ ̸ = j , and iii) other co-channel D2D groups. Correspondingly , BS B j e xperiences interference from all co-channel D2D sources and inter -cell CUs when decoding the uplink signal from its associated CU. 2.1. D2D operational modes T o maximise spectral ef cienc y and range, each D2D group k dynamically selects one of three trans- mission modes based on the instantaneous link quality and the a v ailable CSI: − Mode 0 (Direct Sidelink): The source S k transmits directly to R k . This mode is selected if the l ink distance is small and the SINR satises the required QoS threshold without assistance. − Mode 1 (single-slot RIS-assisted): The relay Π k utilises its UE-specic RIS to reect the signal from S k to R k . This occurs in a single time slot, thereby a v oiding the half-duple x throughput penalty . This mode is preferred for its high ener gy ef cienc y and passi v e nature. − Mode 2 (T w o-slot AF relay): If modes 0 and 1 are insuf cient, Π k functions as an acti v e half-duple x amplify-and-forw ard relay . T ransmission is completed in tw o slots, with reception in slot 1 and forw arding in slot 2. The achie v able rate in this mode is scaled by a f actor of 1 / 2 to account for the half-duple x constraint. Rob ust r esour ce allocation in multi-cell ... (Kayode P opoola) Evaluation Warning : The document was created with Spire.PDF for Python.
2580 ❒ ISSN: 2088-8708 2.2. Lar ge-scale and small-scale fading models The comple x baseband channel coef ci ent between an y tw o nodes ( u, v ) is modelled as h u,v = p β u,v ˜ h u,v . The lar ge-scale g ain β u,v in linear po wer is dened as [27] β u,v = β 0 L − α u,v 10 ξ u,v 10 , (1) where L u,v is the distance, α is the path-loss e xponent, and β 0 is the path-loss at the ref erence distance. The term 10 ξ u,v 10 represents log-normal shado wing with ξ u,v ∼ N (0 , σ 2 sh ) . Note that β u,v is a po wer g ain; the amplitude scaling is p β u,v , ensuring ph ysical consistenc y in the signal-le v el model. F or links e xhibiting a line-of-sight (LoS) component, such as those in v olving the RIS, we emplo y the Rician f ading model: ˜ h u,v = r K R 1 + K R ˜ h LoS u,v + r 1 1 + K R ˜ h NLoS u,v , (2) where K R is the Rician K -f actor , ˜ h LoS u,v is the deterministic LoS component, and ˜ h NLoS u,v ∼ C N (0 , 1) represents the Rayleigh f ading component. 2.3. Effecti v e RIS-assisted channel Let h S k , Π k ∈ C N × 1 and h Π k ,R k ∈ C N × 1 represent the channels from the source to the RIS and from the RIS to the destination, respecti v ely . The RIS phase-shift matrix is Φ k = diag ( e j θ k , 1 , . . . , e j θ k ,N ) , where θ k ,n ∈ [0 , 2 π ) . In Mode 1, the equi v alent end-to-end channel for group k is gi v en by: h eq k = h S k ,R k + h H Π k ,R k Φ k h S k , Π k , (3) where h S k ,R k is the direct link. The ef fecti v e channel po wer g ain is g (1) k = | h eq k | 2 . The passi v e nature of the RIS implies that the noise at the RIS is ne gligible compared to the recei v er noise at R k . 2.4. Hybrid CSI uncertainty framew ork T o accurately reect ph ysical netw ork constraints, we e xplicitly dene a h ybrid error model: a sta- tistical Gauss-Mark o v model for intra-cell pilot estimation errors, and a bounded norm-ball model to handle w orst-case quantisation errors o v er inter -cell backhaul links. − Statistical Gaussian uncertainty (Intra-Cell): Intra-cell links are estimated via pilot signalling, where esti- mation noise is the dominant error source. The actual channel h u,v is related to the estimate ˆ h u,v via the Gauss-Mark o v model h u,v = ρ ˆ h u,v + p 1 − ρ 2 e u,v , e u,v ∼ C N (0 , σ 2 h ) , (4) where ρ ∈ [0 , 1] is the correlation coef cient. − Bounded uncertainty (Inter -Cell): Inter -cell links are acquired through backhaul e xchange, where quantisa- tion and signalling delays result in a bounded error . W e model this as h ∈ H in ter ≜ { ˆ h + ∆ h : ∥ ∆ h ∥ ≤ ϵ } , (5) where ϵ is the uncertainty radius. This allo ws for a rob ust w orst-case design for inter -cell interference mitig ation. The estimation error v ariance is further analysed via the Cram ´ er -Rao bound in Appendix C. In practical multi-cell deplo yments, the uncertainty radius ϵ is not treated as an arbitrary constant b ut is determined by the dominant sources of inter -cell CSI de gradation: quantisation of the backhaul-e xchanged channel estimate at b bits per coef cient, and the backhaul e xchange latenc y τ d relati v e to the cohere n c e time. Concretely , we set this as: ϵ 2 = 2 − b + κ d τ d ∥ ˆ h ∥ 2 , (6) where the rst term captures the quantisation error oor and the second captures channel ageing o v er the backhaul delay , with κ d a Doppler -dependent scaling constant. This formulation ties ϵ directly to deplo yable hardw are parameters (quantisation resolution, backhaul latenc y). Int J Elec & Comp Eng, V ol. 16, No. 5, October 2026: 2575-2594 Evaluation Warning : The document was created with Spire.PDF for Python.
Int J Elec & Comp Eng ISSN: 2088-8708 ❒ 2581 2.5. K ey system parameters Representati v e system parameters for a multi-cell urban 6G deplo yment are summarised in T able 1. These v alues are chosen in accordance wi th commonly adopted 3GPP channel and deplo yment models for urban macro/micro scenarios [28], reecting realistic propag ation conditions and hardw are constraints. The noise po wer is computed as σ 2 N = N 0 B , where N 0 = − 174 dBm/Hz is the noise po wer spectral density and B = 180 kHz is the per -resource block bandwidth. T able 1. System parameters P arameter Symbol V alue/Range Cell radius R 500 m Maximum direct D2D range r 50 m Carrier frequenc y f c 3.5 GHz Per -RB bandwidth B 180 kHz Noise po wer spectral density N 0 − 174 dBm/Hz Shado wing standard de viation σ sh 8 dB P ath-loss e xponent α 3.5 RIS elements per UE-specic RIS N 25–250 Phase-shift quantisation b 3 bits Maximum CU/D2D po wer P C max , P D max 23 dBm, 20 dBm 2.6. Pilot signalling o v erhead Acquisition of the cascaded RIS-assisted channel h S k , Π k , h Π k ,R k ∈ C N × 1 requires orthogonal pilot transmissions whose duration scales linearly with the number of RIS elements N . W ithin a coherence block of length T c symbols, the pilot o v erhead consumes τ p = κN symbols, where κ is the per -element pilot duration. The fraction of the block a v ailable for data transmission is therefore η ( N ) = max 0 , 1 − κN T c . (7) This pre-log f actor is applied multiplicati v ely to all RIS-assisted (Mode 1) achie v able rates. 3. PR OBLEM FORMULA TION 3.1. Decision v ariables and objecti v e Let X = [ x ( j ) k ,m ] ∈ { 0 , 1 } K × M j × J denote the binary reuse-assignment tensor , where x ( j ) k ,m = 1 indicates that D2D group k in cell j reuses the RB of cellular user m in cell j . W e dene P C ∈ R P j M j + and P D ∈ R K + as the collections of CU and D2D transmit po wers, respecti v ely , and Φ = { Φ k } K k =1 as the set of RIS phase congurations. The objecti v e is to maximise the netw ork sum spectral ef cienc y (SE), formulated as: P 0 : max X , P C , P D , Φ J X j =1 M j X m =1   R C j ,m + K X k =1 x ( j ) k ,m R D k ! (8) s.t. (C1) to (C5) . where R C j ,m and R D k represent the achie v able rates of CU C j ,m and D2D group k , respecti v ely . 3.2. Rate expr essions and interfer ence modelling The achie v able rate for CU C j ,m at its associated BS B j is gi v en by R C j ,m = log 2 (1 + γ j ,m,B ) . The SINR is: γ j ,m,B = P C j ,m | g j ,m,B | 2 P K k =1 x ( j ) k ,m P D k | h k ,B j | 2 + I C in ter ,j ,m + σ 2 N , (9) where g j ,m,B denotes the desired channel g ain from C j ,m to B j . The aggre g ate ICI observ ed at B j on the shared RB is: I C in ter ,j ,m = X j ′ ̸ = j   P C j ′ ,m | g j ′ ,m,B j | 2 + X k ′ x ( j ′ ) k ′ ,m P D k ′ | h k ′ ,B j | 2 ! , (10) Rob ust r esour ce allocation in multi-cell ... (Kayode P opoola) Evaluation Warning : The document was created with Spire.PDF for Python.
2582 ❒ ISSN: 2088-8708 which captures the contrib ution of co-channel CUs and D2D sources from neighbouring cells. F or Mode 1 D2D groups, the achie v able rate is R D , (1) k = log 2 (1 + γ (1) k ) , where γ (1) k = P D k | h eq k | 2 I in tra k + I in ter k + σ 2 N . (11) The intra-cell interference is I in tra k = P C j ,m | h j ,m,R k | 2 , while I in ter k is t he sum of co-channel interference from adjacent cells. T o account for the pilot signalling o v erhead required for CSI acquisition, let T c denote the channel coherence block length and τ p = κN denote the pilot training duration, where κ is a constant scaling f actor . The ef fecti v e achie v able rate for mode 1 D2D groups is e xpressed as: R D , (1) k = 1 − κN T c log 2 1 + γ (1) k , (12) 3.3. Rob ust constraints The follo wing constraints ensure netw ork stability and QoS under h ybrid CSI uncertainty: (C1) Probabilistic CU QoS: Subject to Gauss-Mark o v uncertainty , we require Pr { γ j ,m,B ≥ γ C thr } ≥ 1 − δ C for all j , m . (C2) W orst-case D2D QoS: Under bounded uncertainty H in ter , the D2D SINR must sat isfy min ∆ h ∈H in ter γ k ≥ γ D thr for all k . (C3) Po wer b udgets: 0 ≤ P C j ,m ≤ P C max and 0 ≤ P D k ≤ P D max . (C4) RIS hardw are constraints: Phase shifts are restricted to the unit circle and quantised as θ k ,n ∈ { 2 π ℓ/ 2 b } 2 b − 1 ℓ =0 . (C5) RB uniquenes s: Each CU RB can be reused by at most one D2D group, and each group reuses at most one RB. Remark. Problem P 0 is a stochastic non-con v e x MINLP . Its NP-hardness arises from i) the combinat orial nature of reuse assignment, ii) the multiplicati v e coupling between transmit po wers and RIS phase v ariables, and iii) the non-tractable nature of probabilistic and semi-innite w orst-case constraints. 4. PR OPOSED R OB UST MUL TI-CELL RESOURCE ALLOCA TION ALGORITHM T o address the stochastic mix ed-inte ger nonlinear programme in P 0 ef ciently , we propose a three- stage decoupling frame w ork. This approach separates the discrete reuse assignment from the continuous po wer allocation and the non-con v e x RIS phase-shift design. The h i gh-le v el e x ecution o w is s u m marised in Algo- rithm 1. Algorithm 1 . Rob ust multi-cell resource allocation (RMRA) 1: Input: Estimated CSI ˆ h , parameters ρ , ϵ , b udgets P C /D max . 2: Initialise: SA C actor/critic netw orks, P C , P D , Φ . 3: f or each coherence block T c do 4: Obtain intra-cell ˆ h via pilots and inter -cell ˆ h via backhaul. 5: // Stage 1: Distance-Pruned Assignment 6: Prune pairs ( k , m ) where lar ge-scale g ain β R k ,C j ,m > β max . 7: Compute cost matrix [∆ χ k ,m ] and solv e Hung arian matching for X ∗ . 8: // Stage 2: Rob ust P o wer Optimisation (SDP) 9: F ormulate LMIs using BTI (Eq. 16) and S-Procedure (Eq. 18). 10: Solv e SDP via Block Coordinate Descent to yield P C ∗ , P D ∗ . 11: // Stage 3: P assi v e Beamf orming (SA C Infer ence) 12: Observ e state s t (current channels, po wers, ICI). 13: Ex ecute SA C forw ard pass to obtain continuous phase shifts θ . 14: Quantise θ to b -bit resolution yielding Φ ∗ . 15: end f or 16: Output: X ∗ , P C ∗ , P D ∗ , Φ ∗ . Int J Elec & Comp Eng, V ol. 16, No. 5, October 2026: 2575-2594 Evaluation Warning : The document was created with Spire.PDF for Python.
Int J Elec & Comp Eng ISSN: 2088-8708 ❒ 2583 4.1. Stage 1: Distance-pruned Hungarian assignment T o manage the computational comple xity of the multi-cell assignment, we rst prune the set of po- tential reuse partners. Rather than relying on a geometric distance threshold L min , which does not account for shado wing v ariability in urban en vironments, the pruning criterion is e xpressed in terms of the lar ge-scale channel g ain β u,v dened in (1), which already incorporates both pathloss and log-normal shado wing. A pair ( k , m ) is considered feasible only if the lar ge-scale isolation satises β R k ,C j ,m ≤ β max , (13) where β max is a x ed isolation threshold on the cross-link g ain. Because β u,v depends on L u,v , α , and the realised shado wi ng term ξ u,v , this criterion adapts to local shado wing conditions rather than enforcing a x ed geometric separation, and tw o D2D-CU pairs at the same ph ysical distance b ut wi th dif ferent shado wing reali- sations may be pruned dif ferently . F or all feasible pairs, we dene the assignment cost ∆ χ k ,m as the estimated system sum-rate increment: ∆ χ k ,m = E [ R C j ,m + R D k ] shared − E [ R C j ,m ] unshared , (14) where the e xpected rates are calculated using a x ed-po wer proxy and the statistical CSI means. This formula- tion ensures that the Hung arian algorithm select s partners that maximise the mar ginal spectral ef cienc y . The comple xity of this stage is O ( n 3 ) , where n = max( K , M j ) , which is manageable for real-time 6G deplo y- ments. 4.2. Stage 2: Rob ust po wer contr ol via BTI and S-pr ocedur e W ith x ed reuse part ners, Sta g e 2 optimises transmit po wers to satisfy the h ybrid CSI unce rtainty requirements. 4.2.1. BTI r ef ormulation f or CU r eliability The probabilistic constrai n t (C1) under Gaussian uncertainty e ∼ C N ( 0 , σ 2 h I ) can be e xpressed in the quadratic form Pr { e H Qe + 2 ℜ{ r H e } + s ≤ 0 } ≤ δ C . (15) T o render this tractable, we apply the one-sided Bernstein-type ine q ua lity , yielding the follo wing deterministic LMIs: T r( Q ) − p − 2 ln δ C ν + ln δ C µ + s ≥ 0 , (16) v ec( Q ) √ 2 r ≤ ν , µ I + Q ⪰ 0 , µ ≥ 0 , (17) where Q , r , and s are linear functions of P C j ,m and P D k deri v ed from the SINR denominator and the threshold γ C thr . A detailed deri v ation of the BTI-based conic reformulation is pro vided in Appendix A. 4.2.2. S-Pr ocedur e f or D2D r ob ustness F or the bounded inter -cell uncertainty (C2), we require the D2D SINR to hold for all ∥ ∆ h ∥ ≤ ϵ . This semi-innite constraint is transformed using the S-lemma into a single LMI. Specically , (C2) is satised if there e xists a scalar λ ≥ 0 such that A + λ I b b H c − λϵ 2 ⪰ 0 , (18) where A , b , and c represent the quadratic, linear , and constant coef cients of the D2D SINR mar gin, respec- ti v ely . The resulting SDP is solv ed globally across cells via BCD, which ensures a non-decreasing objecti v e sequence. The equi v alence to the LMI is pro v ed in Appendix B. Rob ust r esour ce allocation in multi-cell ... (Kayode P opoola) Evaluation Warning : The document was created with Spire.PDF for Python.
2584 ❒ ISSN: 2088-8708 4.3. Stage 3: P assi v e beamf orming via soft actor -critic The non-con v e xity of the RIS phase-shift optimisation, compounded by the multi-cell interference and unit-modulus constraints, mak es con v entional optimisation techniques dif cult t o apply . W e utilise the SA C algorithm, which emplo ys an entrop y-re gularised frame w ork to maintain an appropriate balance between e xploration and e xploitation in continuous action spaces. − State space ( s t ): The state includes the estimated channel g ains ˆ h S k , Π k , ˆ h Π k ,R k , ˆ h S k ,R k , the current po wer le v els, and the aggre g ate ICI po wer measured at the BS. − Action space ( a t ): The action is the v ector of continuous phase shifts θ ∈ [0 , 2 π ) N K , which are mapped to the b -bit resolution grid before application. − Re w ard function ( r t ): The re w ard is the total system SE penalised by a weighted term Λ for an y violation of the QoS thresholds γ C thr or γ D thr : r t = X j ,m R C j ,m + X k R D k − Λ ⊮ { QoS violation } . (19) The state v ector has dimension | s t | = 2 N K + K + J , comprising the real and imaginary parts of the three estimated channel v ectors per D2D group (stack ed), current po wer allocations { P D k } , and the per -BS aggre g ate ICI measurements. The action v ector θ ∈ [0 , 2 π ) N K has dimension N K . The penalty weight Λ in (16) is not x ed a priori; it is annealed during training according to Λ ( t ) = Λ 0 1 + t T anneal , (20) starting from Λ 0 and increasing linearly o v er training step t up to a cap, so that early e xploration is not o v erly constrained by QoS penalties while later training increasingly enforces feasibility . Λ 0 and T anneal are tuned via grid search to the v alues reported in T able 2 ( Λ 0 = 5 , T anneal = 5 × 10 4 steps), selected as the conguration yielding the lo west QoS-violation rate without de grading the con v er ged re w ard. The SA C implementation utilises twin Q-netw orks to mitig ate o v erestimation bias and a stochast ic actor for rob ust polic y learning. Hyperparameters used for the e xperimental v alidation are summarised in T able 2. T able 2. SA C agent h yperparameters Hyperparameter V alue Learning rate (actor and critic) 3 × 10 − 4 Discount f actor ( ζ ) 0.99 Entrop y coef cient ( α entrop y ) 0.2 T ar get smoothing coef cient ( τ ) 0.005 Replay b uf fer size 10 6 Batch size 256 Hidden layers 2 layers, 256 nodes each Acti v ation function ReLU 4.4. Complexity considerations T o strictly adhere to millisecond-le v el coherence time constraints, the iterati v e e x ecution of Stage 2 and Stage 3 is restricted to the of ine trai n i ng and lar ge-scale f ading tracking phases. During real-time online deplo yment, the SA C agent e x ecutes a single forw ard inference pass O (1) , and the transmit po wers are optimised via a single BCD iteration w arm-started by the pre vious coherence block, thereby bypassing latenc y- intensi v e con v er gence loops. The computational o v erhead of the proposed RMRA frame w ork is anal ysed for each of the three stages. The comple xity of Stage 1 is dominated by the Hun g ari an algorithm, which scales as O (max( K , M j ) 3 ) per cell. Stage 2 in v olv es solving an SDP reformulated via BTI and the S-procedure. Using primal-dual interior -point methods, the w orst-case comple xity for Stage 2 is O ( L ( N K ) 3 . 5 log (1 /ϵ tol )) , where L denotes the number of iterations required for BCD to reach a stationary point. The f actor ( N K ) 3 . 5 reects the size of the LMIs in the SDP . F or Stage 3, the comple xity of SA C is concentrated in the of ine training phase, which scales with the depth and width of the neural netw orks and the size of the replay b uf fer . During online deplo yment, the RIS phase conguration is determined via a single forw ard pass through the actor netw ork, resulting in constant-time inference that is suitable for real-time operation within the channel coherence time. Int J Elec & Comp Eng, V ol. 16, No. 5, October 2026: 2575-2594 Evaluation Warning : The document was created with Spire.PDF for Python.