I nd o ne s ia n J o urna l o f   E lect rica l En g ineering   a nd   Co m pu t er   Science   Vo l.   4 3 ,   No .   1 Ju ly   2 0 2 6 ,   p p .   299 - 31 3   I SS N:  2 5 0 2 - 4 7 5 2 ,   DOI : 1 0 . 1 1 5 9 1 /ijeecs.v 4 3 .i 1 . p p 299 - 31 3           299       J o ur na l ho m ep a g e h ttp : //ij ee cs.ia esco r e. co m   Deep  Q   lea rn ing   a lg o rithm f o r de t ecting D Do S attac ks o   Io T dev ices       L a na   K a m la   Ahm ed 1 K a y h a n Z ra G ha f o o r 2   1 D e p a r t me n t   o f   C o m p u t e r   S c i e n c e   a n d   I n f o r ma t i o n   T e c h n o l o g y ,   C o l l e g e   o f   S c i e n c e ,   S a l a h a d d i n   U n i v e r si t y - Er b i l ,   Er b i l ,   I r a q   2 D e p a r t me n t   o f   I n f o r mat i o n   a n d   C o m mu n i c a t i o n   Te c h n o l o g y   En g i n e e r i n g ,   Er b i l   P o l y t e c h n i c   U n i v e r si t y ,   Er b i l ,   I r a q       Art icle  I nfo     AB S T RAC T   A r ticle  his to r y:   R ec eiv ed   Feb   2 1 ,   2 0 2 6   R ev is ed   J u n   1 0 ,   2 0 2 6   Acc ep ted   J u n   2 7 ,   2 0 2 6       Th e   ra p id   e x p a n sio n   o i n tern e t   o th i n g (Io T)  n e two r k h a h e ig h ten e d   se c u rit y   risk s ,   p a rt icu larly   re g a r d in g   d istri b u ted   d e n ial  o f   se rv ic e   (DD o S a tt a c k a g a in st  d e v ice with   li m it e d   c o m p u t in g   c a p a c it y .   Hi g h   d e tec ti o n   a c c u ra c y   is  c ru c ial  fo th e se   re so u rc e - c o n stra in e d   e n v iro n m e n ts,  wh e re   fa lse   p o siti v e c a n   d isr u p t   leg it ima te  t ra ffic  a n d   fa lse   n e g a ti v e a ll o a tt a c k to   p e rsist.   Ho we v e r,   m o d e rn   re i n fo r c e m e n lea rn in g   (RL)   a n d   m a c h i n e   lea rn in g   (M L)  in tru si o n   d e tec ti o n   so l u ti o n o ften   e x h ib it   p o o g e n e ra li z a ti o n   d u e   t o   sta ti c   sta te  re p re se n tatio n s.   To   a d d re ss   th is,   th is  p a p e p r o p o se a   d e e p     Q - l e a rn in g   (DQ L)  fra m e wo rk   t h a in teg ra tes   K - m e a n c l u ste rin g   d irec tl y   in to   th e   R a c ti o n   sp a c e .   U n li k e   p ri o RL - b a se d   IDS  m o d e ls,  o u r   a p p r o a c h   d y n a m ica ll y   in teg ra tes   c lu ste ri n g   in t o   th e   lea rn in g   p ro c e ss ,   e n a b li n g   a d a p ti v e   sta te  re p re se n tati o n   a n d   imp ro v e d   g e n e ra li z a ti o n   to   u n se e n   traffic   p a tt e rn s.  T h e   sy ste m   is  fo rm u late d   a a   M a rk o v   d e c isio n   p ro c e ss   wh e re   th e   a g e n o p ti m ize a   c o m p o site   re wa rd   fu n c ti o n   b a se d   o n   a c c u ra c y ,   p re c isio n ,   re c a ll ,   a n d   F 1 - sc o re .   Ev a l u a ted   o n   th e   N - Ba Io d a tas e u sin g   1 0 - fo l d     c ro ss - v a li d a ti o n ,   t h e   p r o p o se d   m e th o d   a c h iev e s   a   c las sifica ti o n   a c c u ra c y   o f   9 8 . 9 5 %   a n d   a   we ig h ted   F 1 - sc o r e   o 9 8 . 7 3 % ,   si g n if ica n tl y   o u tp e rfo rm in g   trad it io n a M a n d   RL   b a s e li n e s.  Th e se   re su lt s   d e m o n stra te  th e     fra m e wo rk ' e ffe c ti v e n e ss   a a   sc a lab le,  a d a p t iv e   s o lu ti o n   f o r   in tel li g e n t   Io T   DD o S   d e tec ti o n .   K ey w o r d s :   Dee p   Q - lear n in g   Dis tr ib u ted   d en ial  o f   s er v ice   I n ter n et  o f   th in g s   s ec u r ity   lear n in g   R ein f o r ce m en t le ar n i n g   T h is i a n   o p e n   a c c e ss   a rticle   u n d e r th e   CC B Y - SA   li c e n se .     C o r r e s p o nd ing   A uth o r :   L an Kam la  Ah m ed   Dep ar tm en t o f   C o m p u ter   Scin ec an d   I n f o r m atio n   T ec h n o lo g y ,   C o lleg o f   Scien ce   Salah ad d in   Un iv er s ity - E r b il   Ker k u k   R o ad ,   E r b il,  I r aq   E m ail:  k lan am an atik @ g m ail. c o m       1.   I NT RO D UCT I O N   T h r ap id   ev o lu tio n   o f   th i n ter n et  o f   th in g s   ( I o T )   h as  tr an s f o r m ed   h o d ev ices  in te r ac wh ile  ex p o s in g   s y s tem s   to   ad v a n ce d   s ec u r ity   th r ea ts .   Su ch   attac k s ,   in clu d in g   SYN  f lo o d s   an d   UDP  am p lific atio n ,   ar k n o wn   as  d is tr ib u te d   d e n i al - of - s er v ice  ( DDo S)  attac k s   th at  ex p lo it  t h lim ited   d ef en s iv ca p ab ilit ies  o f   I o T   h ar d war e,   g iv en   its   r eso u r ce - co n s tr ain ed   n atu r ( 2 0 2 4   s tatis t ics  s h o th at  o v er   2 1 . 3   m illi o n   attac k s   ar r eg is ter ed )   [ 1 ] .   Hig h   d etec tio n   r ates  m ea n   ev er y th in g   to   s u ch   r eso u r ce - co n s tr ain e d   d ev i ce s f alse  p o s itiv es   ca n   d is r u p leg itima te  tr a f f ic ,   b u f alse  n e g ativ es  ca n   en ab le  ad v er s ar ial  s u r v i v al.   C o n v en tio n al  s ec u r ity   f r am ewo r k s   th at  ce n ter   o n   id e n tity   co n tr o l,  ac ce s s   co n tr o l,  a n d   au th o r izatio n   o f ten   ca n n o h an d le  lar g e - s ca le,   d y n am ic  c y b er attac k s   th at  h av b ec o m d o m in an t   f ea t u r o f   th I o T   [ 2 ] ,   [ 3 ] .   B esid es,  h ig h   r ates  o f   m alwa r e,   p ac k et - s n if f in g ,   an d   s p o o f in g   r eq u ir d etec tio n   m o d els  th at  ca n   ad ap in   r ea tim to   ch an g in g   tr af f ic  p atter n s   [ 4 ] ,   [ 5 ] .   Evaluation Warning : The document was created with Spire.PDF for Python.
                      I SS N :   2 5 0 2 - 4 7 5 2   I n d o n esian   J   E lec  E n g   &   C o m p   Sci Vo l.  4 3 ,   No .   1 ,   Ju ly   20 2 6 :   299 - 31 3   300   Alth o u g h   p r im itiv m ac h in e - lear n in g   ( ML )   a n d   d ee p - lear n in g   ( DL )   m o d els,  lik co n v o lu tio n al   n eu r al  n etwo r k s   ( C NNs),   r ec u r r en n eu r al  n etwo r k s   ( R NNs),   an d   a u to en co d er s ,   r o u tin el y   m o n ito r   co m m o n   m etr ics  o f   th n etwo r k ,   s u ch   as  th s o u r c I a d d r ess es,  th ey   ten d   n o t   to   a d ap to   v ar y in g   tr af f ic   p atter n s     [ 6 ] - [ 8 ] .   Key b o ar d s   th er m al  a n d   en er g y   an aly s is   with   ex ter io r   m esh   s cr ee n s   o f   o f f ice  s p ac es  u n d er   d iv e r s e   clim atic  co n d itio n s .   B u ex is t in g   r ein f o r ce m en t - b ased   I DS  r ely   o n   f ix ed   s tate  m o d el lin g   an d   lack   ad ap tiv e   f ea tu r ab s tr ac tio n   [ 9 ] [ 1 0 ] .   E v en   th o u g h   th liter atu r in d icate s   th at  p r ed ictiv ac cu r ac y   h as  im p r o v ed   o n ly   m ar g in ally ,   s ig n if ican ch alle n g es  r em ain .   T h class ical  ML   p ip elin es  ( h y b r id   K - m e an s   co m b in ed   with   ar tific ial  n eu r al  n etw o r k s   ( A NNs)  [ 1 1 ] ,   s u p p o r t - v ec to r   m a ch in es  [ 1 2 ] ,   o r   r an d o m - f o r est  en s em b les  [ 1 3 ] )   ar e   o f ten   lin k ed   to   eith er   h ig h   co m p u tatio n al  co s t o r   p o o r   g en e r aliza tio n .   Similar ly ,   th ac cu r ac y   o f   d ee p - lea r n in g   m o d els  s u ch   as  DE E Sh ield   [ 1 4 ] ,   [ 1 5 ] ,   XGBo o s [ 1 6 ] ,   an d   B i L STM - C NN  [ 1 7 ]   is   also   q u ite  r ea s o n ab le,   b u t   th ey   p o s s ig n if ican im p lem en tatio n   ch allen g es  in   r eso u r ce - lim ited   co n tex ts   [ 1 8 ] ,   [ 1 9 ] .   At  th s am tim e,   RL - b ased   m o d els  lik MA DDPG  [ 2 0 ]   an d   d ee p   Q - lear n in g   ( DQL ) - I DS  [ 2 1 ]   ar e   ad ap ti v b u ar e   g en e r ally   test ed   in   s y n th etic  s im u latio n s   an d   h a v n o t b ee n   p r o p er ly   as s ess ed   o n   u n s ee n   tr af f ic  s am p l es [ 2 2 ] - [ 2 5 ] .   T h is   Stu d y   f ills   th is   g ap   b y   m ak in g   ad a p tiv clu s ter in g   a n   in teg r al  p ar o f   t h lear n in g   p r o ce s s .   Un lik p r e v io u s   m o d els,  we   a ls o   in tr o d u ce   clu s ter in g   in to   t h ag e n t' s   d ec is io n   s p ac e,   th e r eb y   d ev elo p in g   an   ad ap tiv s tate  r e p r esen tatio n   th at  g en e r alize s   to   n ew  tr af f ic  to p o lo g ies.  I n   p ar ticu lar ,   th e   K - m ea n s   alg o r it h m   is   ad d ed   to   th DQL   ag en t' s   ac tio n   r ep er to ir t o   en ab le  it to   i d en tify   th o p tim al  n u m b er   o f   clu s ter s   ( K)   in   r ea tim e.   As  an   ex ten s io n   o f   th n o n - a d ap tiv p r e p r o ce s s in g ,   we  u s co m p o s ite  r ewa r d   s c h em th at  in teg r ates  ac cu r ac y ,   p r ec is io n ,   r ec all,   a n d   F1 - s co r t o   jo in tly   o p tim ize  clu s ter in g   d ec is io n s   an d   in tr u s io n - d etec tio n   p o licies.  T h is   p ar ad ig m   e n ab les  th ag en to   o b tain   h ig h er - lev el  s tate  ab s tr ac tio n   an d   s tab ilize  Q - v alu co n v er g en ce ,   th e r eb y   im p r o v i n g   th ef f ec tiv en ess   o f   DDo d etec tio n .   T h m ain   co n tr i b u tio n s   o f   t h e   wo r k   ar s u m m ar ize d   as  f o llo ws:   th cr ea tio n   o f   DQL   ar ch itectu r e   th at  allo ws  cr ea tin g   ad ap tiv r ep r esen tatio n s   b ased   o n   th ac tio n   s p ac e,   th d esig n   o f   n ew  co m p o s ite  m ec h an is m   o f   r ewa r d s   t o   o p t im ize  th em   s im u ltan eo u s ly ,   a n d   an   em p ir ical  v alid atio n   is   co n d u cted   with   s ig n if ican am o u n o f   r ig o r   u s in g   ten - f o ld   cr o s s - v alid atio n   o n   th N - B aI o T   d ataset  to   p r o v th e   s u p er io r ity   o v er   tr ad itio n al  b aselin es.  T h r est  o f   th m an u s cr ip will  b s tr u ctu r ed   an d   p r esen ted   as  f o llo ws:   s ec tio n   2   will  d escr ib th m ater ials   an d   m eth o d o lo g y ,   s ec tio n   3   will  p r esen th em p ir ical  f in d in g s ,   an d   s ec tio n   4   will  p r esen t th co n clu s io n   an d   r ec o m m en d atio n s   f o r   f u tu r e   r esear ch .       2.   M E T H O D   T o   m ain tain   tr an s p a r en cy ,   r ep r o d u cib ilit y ,   an d   s cien tific   r i g o r ,   th e   s ec tio n   d escr ib es  th e   in s tr u m en ts ,   m ater ials ,   an d   ex p er im en tal  m eth o d s   ap p lied   wh e n   ca r r y i n g   o u th cu r r en s tu d y .   Fig u r 1   illu s tr ates  th ex p er im en tal   p ip elin e   u s ed   i n   th e   p r esen r esear ch .   I s ta r ts   with   th a b s o r p tio n   o f   th N - B aI o T   d ataset,   co n tin u es  with   d ata  p r ep r o ce s s in g   an d   f ea tu r en g in ee r i n g ,   an d   en d s   with   th cr ea tio n   an d   test in g   o f   two   r ein f o r ce m e n lear n in g   m o d els,  i.e . ,   Q - lear n in g   a n d   DQL .   T h two   m o d els  co m b in th K - m ea n s   clu s ter in g   m eth o d   a n d   ar e   tr ain ed   an d   ev alu ated   u s in g   K - f o ld   cr o s s - v alid atio n .   T h e   p ip elin e   co n clu d es  with   a   co m p ar at iv p er f o r m an ce   a n al y s is   u s in g   p r ed ef in e d   ev alu ati o n   m etr ics.           Fig u r 1 .   E x p er im e n tal  p ip elin in teg r atin g   a d ap tiv clu s ter i n g   with   DQL       2 . 1 .     Da t a s et   d escript io n   T h d ataset  u s ed   to   tr ain   th p r o p o s ed   m o d el  is   th NB aI o T   d ataset,   p u b licly   av ailab le  o n   Kag g le   [ 2 6 ] .   T h is   d ataset  co m p r is es  n etwo r k   tr af f ic  f r o m   n in r ea l - wo r ld   I o T   d ev ices,  s u ch   as  we b ca m s ,   th er m o s tats ,   Evaluation Warning : The document was created with Spire.PDF for Python.
I n d o n esian   J   E lec  E n g   &   C o m p   Sci     I SS N:   2502 - 4 7 5 2       Dee p   lea r n in g   a l g o r ith fo r   d etec tin g   DDo S   a tta ck s   o n   I o T d ev ices  ( La n a   K a mla   A h m ed )   301   an d   b ab y   m o n ito r s ,   ca p tu r ed   d u r in g   r eg u lar   o p er ati o n   a n d   u n d er   v a r iety   o f   attac k   co n d itio n s ,   in clu d in g   SYN  f lo o d ,   UDP  f lo o d ,   an d   d ata  ex f iltra tio n .   Alth o u g h   th d ataset  is   r ea li s tic  in   ter m s   o f   th d y n am ics  o f   a   ty p ical  I o T   d ev ice,   it  h as  d is p r o p o r tio n ate  class   d is tr ib u t io n ,   with   th attac k   tr af f ic  s a m p le  s ig n if ican tly   lar g er   th an   t h at  o f   n o r m al   tr a f f ic.   T h is   im b alan ce   m ay   h a v an   im p ac o n   th e   m o d el   tr a in in g   a n d   o u tp u ts .   T ab le  1   will  b th s u m m a r y   o f   th k ey   p o in ts   r elate d   to   t h d ataset,   s u ch   as  s am p le  s izes  an d   d ataset  an d   tr af f ic  f ea tu r s izes [ 2 6 ] .         T ab le  1 .   C h ar ac ter is tics   an d   s am p le  s tr u ctu r ( N - B aI o T   d ataset)   A t t r i b u t e   D e scri p t i o n   N u mb e r   o f   sam p l e s   7 , 0 0 0 , 0 0 0 +   N u mb e r   o f   f e a t u r e s   1 1 5   Tr a f f i c   t y p e s   B e n i g n ,   M i r a i - b a se d ,   B A S H LI TE - b a s e d ,   T C P ,   U D P   f l o o d s ,   s c a n s,  c o mb o s   Ty p e   o f   b o t n e t   M i r a i   a n d   B A S H LI TE  ( a . k . a .   G a f g y t )   C l a s si f i c a t i o n   f r a mi n g   B i n a r y   ( b e n i g n   v s   a t t a c k )   a n d   mu l t i c l a ss (1 1   t o t a l   c l a s ses)   K e y   f e a t u r e   d o ma i n s   P r o t o c o l   s t a t s,   p a c k e t   r a t e ,   I C M P   f r e q u e n c y ,   st r e a m a n o ma l i e s   C l a s d i s t r i b u t i o n   I mb a l a n c e d   -   b e n i g n   d o mi n a t e s   I o d e v i c e t y p e s   C o l l e c t e d   f r o m   d i f f e r e n t   t y p e o f   I o r e a l   d e v i c e   i n c l u d i n g   Th e r mo s t a t ,   P h i l i p s   B a b y   mo n i t o r ,   p r o v i s i o n   sec u r i t y   c a m e r a ,   s a ms u n g   sm a r t   TV ,   S i mp l e   H o m e   C a mer a s,  a n d   W e M o   S mart   P l u g ,   D a mi n i   D o o r b e l l ,   Ec o   B e e   T o .       2 . 2 .     Da t a   p re pro ce s s ing   T h m o s t im p o r tan t p a r t in   th e   p r ep ar atio n   o f   th d atasets   was p r ep r o ce s s in g .   T h in itial p h ase  o f   th p r o ce s s   was  to   f ill  g ap s   in   v alu es  with   th h elp   o f   th n p .   n an   to   n u m   f u n ctio n ,   wh ic h   s u b s titu tes  m is s in g   v alu es  with   n u m er ical   esti m at es.  T h a d d itio n al   d im e n s io n a lity   r ed u ctio n   was  d o n e   th r o u g h   th e   r e m o v al   o f   r ed u n d an lo w - v ar ian ce   an d   D Do S - d etec tio n - ir r elev a n f ea tu r es,  wh ich   h el p ed   t o   f ilter   t h n u m b er   o f   f ac to r s   th at  m ay   n o co n tr ib u te   to   th e   tr ain in g   o f   t h m o d els.  Su b s eq u en n o r m al izatio n   o f   co n ti n u o u s   f ea t u r es  was  p er f o r m ed   u s in g   Z - s co r s ta n d ar d izatio n   with   Stan d ar d S ca ler ,   wh ich   s tab ilizes  th d ata  d is tr ib u tio n   an d   im p r o v es  lear n i n g   [ 2 7 ] .   A n o th er   ch an g is   th at  th s cik it - lear n   KB in s Dis cr etize r   was   u s ed   to   g en e r ate  s tatis t ically   e q u alize d   b in   b o u n d ar ies,  th er e b y   p r o v i d in g   a   d ee p er   s tate  r ep r esen tatio n   f o r   th DQL   ag en t.   Sp ec if ically ,   th N - B aI o T   d a taset  is   n o ca teg o r ical;  h o w ev er ,   s in ce   t h r aw  i n p u t   was  d is cr etize d ,   th e   r esu ltin g   b in s   wer ca teg o r ica an d   co n v e r ted   in to   a n   en u m er ated ,   o n e - h o f o r m at  to   m a k th em   co m p atib le   with   th u n d e r ly in g   lea r n in g   m o d el.     2 . 3 .     M o del  a rc hite ct ure   T h cu r r en r esear ch   o u tlin es  th f r am ewo r k   f o r   a   DQL   s y s tem   im p lem en ted   as  f u lly   co n n ec te d   f ee d - f o r war d   n eu r al  n etwo r k ,   as  s h o wn   in   Fig u r 2 .   T h n etwo r k   h as  n etwo r k   in p u t   lay er ,   two   h id d en   lay er s ,   an d   an   o u tp u lay er .   No n lin ea r   p atter n s   in   th in p u d ata  ar ca p tu r ed   b y   th h i d d en   lay er s   u s in g   th e   r ec tifie d   lin ea r   u n it  ( R eL U)   ac tiv atio n   f u n ctio n ,   wh ich   h as  b ee n   p r o v en   e f f ec tiv e   in   d ee p   lear n i n g   a r ch itectu r es  [ 2 8 ] .   Fu r th er m o r e,   to   a d d r ess   class   im b alan ce   an d   en s u r e   h ig h   d etec tio n   q u ality   [ 2 9 ] ,   we  d esig n ed   a   n o v el  c o m p o s ite  r e war d   f u n ctio n .   Ou r   r ewa r d   s ig n al,   u n lik s tan d a r d   m o d els,  c o m b in es  Acc u r ac y ,   Pre cisi o n ,   R ec all,   an d   F1 - s co r ( R   0 . 2 5   *   Acc   0 . 2 5   *   Pre 0 . 2 5   *   R ec   0 . 2 5   *   F1 ) .   Alth o u g h   th e   r ewa r d   f u n ctio n   is   co m p u ted   f r o m   th g r o u n d - tr u th   la b els  o f   th N - B aI o T   d ataset,   th is   is   tactica d esig n   d ec is io n   f o r   th o f f lin tr ain i n g   s tag e.   T h is   en ab les  th ag en to   lear n   an   o p tim al  class if ica t io n   p o licy   th at  ca n   u ltima tely   b ap p lied   i n   an   au to n o m o u s ,   lab el - f r ee   en v ir o n m en t.  m u ltid im en s io n al  f ee d b ac k   allo ws   r ed u cin g   th r ate  o f   f alse  p o s itiv es  an d   f alse  n eg ativ es  at  th s am tim e,   wh ich   is   n ec es s a r y   to   g u ar a n tee  th e   av ailab ilit y   o f   s er v ices  i n   r eso u r ce - lim ited   I o T   s ettin g s .   T h e   o u tp u lay e r   u s es  lin ea r   ac ti v atio n   t o   esti m ated   Q - v alu es  o f   ev er y   p o s s ib le  ac tio n   av ailab le,   th er ef o r allo win g   th ag en to   ca teg o r ize  th n etwo r k   tr af f ic  as  b en ig n   o r   m alicio u s   ac c o r d in g   to   d y n am ically   esti m ated   cl u s ter   ass ig n m en ts .   I n   o r d e r   to   m a k th r es p o n s e   to   c h a n g i n g   t r a f f ic   c o n f ig u r a t io n s   o f   t h e   I o T   n e tw o r k s   m o r f le x i b l e,     K - m ea n s   cl u s t er in g   is   u s e d   to   co m p le m e n t   t h e   DQ L   f r a m e w o r k .   T h e   p r o c ess   o f   in te g r ati o n   al lo ws  s t at es o f   t h n et wo r k s   t o   b e   ag g r e g a te d   i n   a   s tr u ct u r ed   f o r m   to   g i v e   s ta te   i n f o r m at io n   o f   t h e   h i g h - d i m e n s i o n al   s tat es   o f   t h e   s y s te m .   T u p l es o f   ( s t ate ,   ac ti o n ,   r ew ar d ,   n e x t   s t ate )   tr a n s iti o n s   o f   s t at ar f e d   i n t o   r e p l ay   b u f f er   i n   tr ai n i n g   t o   r e g u la r i ze   th tr ai n i n g   p r o c e s s   a n d   t o   r e d u c c o r r el at i o n s   am o n g s r e lat ed   s a m p les .   I ap p lies   ε - g r e e d y   alg o r it h m   i n   o r d e r   to   m ak t r ad e - o f f   wi th   e x p l o r ati o n   an d   e x p l o ita ti o n   wi th   t h e   ai m   o f   m a x i m iz in g   lo n g - te r m   cu m u lat iv r e tu r n s .   O v e r al l,   t h is   ar c h it ec t u r will   al lo D QL   a g en t o   b t r ai n ed   t o   a d o p u s ef u l   p o l ici es  t o   d et ec t   DD o S   i n   I o T   e n v ir o n m en ts .   T h e   a r c h it ec t u r e   p r o v i d e s   a   s tat e   c h a n g e   s eq u e n ce ,   a n d   als o   u s es   ad a p ti v e   clu s te r ,   a n d   th is   m a k es   i p o s s i b le   t o   le ar n   a n d   d e te ct  p atte r n s   i n   h i g h - d i m e n s i o n a n e tw o r k   t r a f f ic .         Evaluation Warning : The document was created with Spire.PDF for Python.
                      I SS N :   2 5 0 2 - 4 7 5 2   I n d o n esian   J   E lec  E n g   &   C o m p   Sci Vo l.  4 3 ,   No .   1 ,   Ju ly   20 2 6 :   299 - 31 3   302       Fig u r 2 .   Pro p o s ed   DQL   ar c h i tectu r f o r   d y n a m ic  Q - v alu a p p r o x im atio n       2 . 4 .     P r o po s ed  enha nced  DQ L   det ec t io n m o del   T o   s elec an d   ev alu ate  r esea r ch   m o d el  o f   DQL ,   th p r esen p ap er   f o llo wed   s y s tem atic  ap p r o ac h   an d   b eg a n   b y   estab lis h in g   t h m ain   h y p e r p ar am ete r s ,   wh ich   d ef in ed   t h b ac k g r o u n d   o f   th ex p er im e n p r o to co l,   i.e . ,   th e   d ataset,   lea r n in g   r ate,   d is co u n f ac to r ,   m em o r y   s ize  an d   th n u m b er   o f   K - f o l d s .   T h N - B aI o T   was  in   th e   f ir s s tep s   p r ep r o ce s s ed   with   th e   n o r m aliza tio n   an d   f ea tu r e n co d in g   to   e n h an ce   th q u ality   o f   d ata  an d   co r r ec th p ar tia m is s in g n ess .   Fo llo win g   th is ,   th K - f o ld   cr o s s - v alid atio n   d iv id ed   th d ata  to   f o r m   tr ain i n g   an d   cr o s s - v ali d atio n   f o l d s ,   th u s   allo win g   i n ten s iv ass ess m en o f   s ev er al  f o ld s   [ 3 0 ] T h is   wo r k f lo m ain tain s   p ar am ete r   s ettin g s   in   p r o ce s s   a s   s h o wn   in   Fig u r 3 .   I h elp s   in   th u ltima te  ev alu atio n   o f   ag g r eg ate  p e r f o r m an ce   m ea s u r es a n d ,   h e n ce ,   g o o d   m o n ito r in g   o f   DDo S a ttack s .   On   to p   o f   th is   wo r k f lo w,   th s y s tem   d ir ec tly   in clu d es  K - m ea n s   clu s ter in g   m o d u le,   th e   DQL   ag en t,   m eta - lear n er   [ 3 1 ] .   T h ag e n t,  as  o p p o s ed   to   u s in g   th s tatic  m eth o d   o f   clu s ter in g ,   f ig u r es  o u th o p tim al   n u m b er   o f   clu s ter s   ( K)   o n   th e   f ly ,   an d   th is   is   co n s id er ed   to   b th ac tio n   p er f o r m ed   b y   th K - m ea n s   m o d u le.   T h en a b lin g   a d ap tiv e   ap p r o ac h   e n ab les  n o is y   f ea tu r es   to   b clu s ter ed   au to m atica lly   an d   b o o s ts   th s ep ar atio n   o f   class es  in   c o m p l icate d   tr af f ic.   T h e   r e p lay   m e m o r y   ( D)   is   em p lo y ed   to   s tab ilize  th n etwo r k   b y   k ee p in g   p r ev io u s   ex p er ien ce s ,   an d   an   e p s ilo n - g r ee d y   p o licy   is   em p lo y ed   to   m ak th p r o g r am   b ala n ce d   in   ex p lo r in g   n ew  attac k   p atter n s   in   ad d itio n   to   th p r ev i o u s   o n es.  T h e   wh o le  r atio n ale  o f   th wo r k in g   m ec h an is m   o f   th e   s y s tem ,   t h p a r ticu lar   p ar am eter   s ettin g ,   a n d   th d esig n   elem e n ts   ar s u m m ar ized   in     T ab le  1   [ 3 2 ]   a n d   f o r m alize d   in   Alg o r ith m   1 .           Fig u r 3 .   Sy s tem   wo r k f lo f o r   ag en t - b ased   o p tim al  clu s ter   s elec tio n       T ab le  2 .   DQL   fr am ewo r k   p ar a m eter s   an d   d esig n   elem e n ts   s u m m ar y   P a r a me t e r   D e f i n i t i o n   S t a t e ( S )   D i scret e   n e t w o r k   t r a f f i c   c h a r a c t e r i st i c s o f   t h e   d a t a se t   A c t i o n ( A )   C h o o si n g   t h e   b e st   n u m b e r   o f   c l u st e r s   ( K )   o f   t h e   K - m e a n s m o d u l e .   R e w a r d   ( R )     A   c o m p o si t e   si g n a l :   0 . 2 5   * ( A c c u r a c y   +   P r e c i si o n   +   R e c a l l   +   F 1 - sc o r e )   P o l i c y     A n     e p s i l o n - g r e e d y   s t r a t e g y   ( st a r t i n g   a t   1 . 0 ,   d e c a y i n g   t o   0 . 0 1 )   Q   ( st ,   a t )   C u r r e n t   Q - v a l u e   s t a t e   a n d   a c t i o n   α (Lea r n i n g   r a t e )   R e g u l a t i n g   t h e   w e i g h t   o f   n e w   i n f o r m a t i o n   γ   ( D i sc o u n t   f a c t o r )   D e t e r m i n i n g   t h e   i m p o r t a n c e   o f   f u t u r e   r e w a r d s   Evaluation Warning : The document was created with Spire.PDF for Python.
I n d o n esian   J   E lec  E n g   &   C o m p   Sci     I SS N:   2502 - 4 7 5 2       Dee p   lea r n in g   a l g o r ith fo r   d etec tin g   DDo S   a tta ck s   o n   I o T d ev ices  ( La n a   K a mla   A h m ed )   303   Alg o r ith m   1 .   E n h a n ce d   DQL     f o r   I o T   DDo d etec tio n   in pu t:   da ta se t_ fo ld er al ph ),   ga mm ),   ϵ_ st ar t,   ϵ_ de ca y,   ϵ_ m in ,n um _e pi so de s,   replaymemory d, clustersize (possible - ks),k_fold    output:  performance metrics, confusion matrix        1. Initialize replay memory        2. initialize q - network q(s, a; θ) with random weights        3. set ϵ = ϵ_start        4. for each file in dataset_folder do       4.1 Preprocess x (normalize features, handle missing values       4.2 Perform k - fold cross - validation:           for each fold in k_fold do                 Split data into xtrain, ytrain, xtest, ytest              for episode = 1 to num_episodes do         4.2.1 set state = get_features(xtrain)         4.2.2 while training_state  exists do                            if (random(0,1) < ϵ) then                                    set action = randomaction()                              else                                     set action = argmax(q(state))                           end if          4.2.2. 1 Perform k - means clustering using k = action          4.2.2.2 Map clusters to labels using majority voting          4.2.2.3 set acc = calc_accuracy()                                    set prec = calc_precision()                                   set rec = calc_recal l()                                    set f1 = calc_f1_score()          4.2.2.4 set reward = (0.25 * acc) + (0.25 * prec) + (0.25 * rec) + (0.25 * f1)          4.2.2.5 store (state, action, reward, next_state) in d          4.2.2.6 update q(state, action) using bellm an:                                       set q = q + α * (reward + γ * max_q  -   q)          4.2.2.7 set ϵ = max(ϵ * ϵ_decay, ϵ_min)                         end while                    end for        4.3 set optimal_action = argmax(q(test_state))        4.4 Evaluate xtest using k - means with k = optimal_action        4.5 Calculate final metrics and confusion matrix              end for   END FOR         5. AGGREGATE metrics_across_folds.         6. RETURN overall metrics,  Confusion Matrix.     T h DQL   m o d el  h y p er p a r am e ter s   wer ch o s en   to   g u a r an tee   s tab le  co n v er g en ce   a n d   a   s tr o n g   ab ilit y   to   d etec d if f e r en p atter n s   o f   attac k s .   T h lear n in g   r ate  (   0 . 0 0 1 )   was  s elec ted   to   h av b alan ce   b etwe en   th e   s p ee d   at  wh ich   th weig h ts   ar u p d ated th h ig h er   t h v alu es,  t h f aster   th n etwo r k   wo u ld   g et  awa y ,   an d   t h lo wer   th v alu es,  t h s o o n e r   t h n etwo r k   wo u ld   lear n .   T h d is co u n f ac to r   (   γ   )   was  d eter m in ed   to   b 0 . 9 5   s o   th at  th ag en t   co u ld   m o d el   th f u tu r n etwo r k   s tates,  an d   th is   is   n ec ess ar y   in   o r d er   to   m o d el  t h tim d ep en d e n cies  o f   m u lti - s tag DDo attac k .   T h ep s ilo n - d e ca y   s ch ed u le  b eg in s   with   1 . 0   an d   d ec ay s   to   0 . 0 1   th r o u g h o u t   th f ir s 1 0 0   ep is o d es,  wh ich   allo ws  th a g en t   to   ex p lo r e   th h ig h - d im en s i o n al  f ea tu r es  o f   N - B aI o T   b r o a d ly   at   th s tar b ef o r ch a n g in g   o n   to   s tab le  e x p lo itatio n   p o licy .     T h e   s elec ted   v alu es  c o r r esp o n d   to   th b est  v alid atio n   r esu lt s   o b tain ed   d u r in g   th e x p er i m en tal  ev alu atio n .   Sp ec if ical ly ,   α   c o n tr o ls   th e   m ag n itu d o f   weig h u p d ates  in   Q - n etwo r k   o p tim izatio n ,   t h er eb y   en s u r in g   p r u n e d   lear n i n g   o f   k n o wled g e .   γ ,   o n   th e   o th er   h an d ,   h ig h lig h ts   th f u tu r b en ef its ,   allo win g   th ag e n to   ta k in to   co n s i d er atio n   l o n g - te r m   b en ef its   in   th co m b in atio n   o f   cu m u lativ r ewa r d s .   T h ex p l o r atio n   co e f f icien ( ϵ )   is   th ten s io n   b etwe en   th e   d esire s   to   ex p lo r an d   to   ex p lo it  in   f av o r   o f   an   ac tio n   s el ec tio n   p r o ce s s   th at  n av ig ates   th ag en b etwe en   p o licy   r ef in em en an d   s u cc ess iv iter atio n .   T h is   is   s to r ed   in   r ep lay   b u f f er   ( D)   to   r ed u ce   t r ain in g   in s tab ilit y .   E p is o d ic  ex p e r ien ce s ,   in clu d in g   ac tio n s ,   s tates,  an d   r e war d s ,   ar tem p o r ar ily   s to r ed   h er e,   allo win g   m in ib atch es  to   b e   s to ch asti ca lly   ex tr ac ted   a n d   r e d u cin g   co r r elatio n s   in h er en in   tem p o r ally   o r d e r ed   d ata.   T h e   s ize  o f   th ch o s en   b atc h   d eter m in es  h o m an y   e x p er ie n ce s   ar ad d ed   t o   th b u f f er   ea c h   tim an   o p tim i za tio n   s tep   is   p er f o r m ed ,   wh ich ,   in   tu r n ,   in f lu e n ce s   th co n v er g en ce   r ate  an d   o v er all  lear n in g   s tab ilit y .   B esid es,  th ε  - d ec ay   s ch e d u le  a n d   m in im al   ex p lo r atio n   t h r esh o ld   ( θ)   d e f in th e   tem p o r al  d ec a d en ce   o f   ε   an d   p lace   a n   ab s o lu te  lo wer   lim it,  e n s u r in g   th at  o cc asio n al  e x p lo r ato r y   b eh av io r   is   m ai n tain ed   d u r in g   t h tr ain in g   h is to r y .   I u s es  th Q - n etwo r k ,   wh ich   i s   r ef er r ed   to   as  ( s ,   a,   θ) ,   to   ap p r o x im ate  Q - v alu es  o f   ac tio n - s tate  p air s ,   an d   θ   is   th tr ain ab le  p ar am eter   o f   t h n eu r al  n etwo r k .   T h n u m b e r   o f   t r ain in g   e p is o d es (   Nu m   ep   ep is o d es)  tells   th e   am o u n t o f   tr ai n in g .   T h p h ase   o f   esti m atio n   o f   th K - m ea n s   clu s ter in g   is   co n tin g en t o n   th e   v ar iety   o f   p o s s ib le   ca n d id ate  clu s ter   v al u es  ( KS) .   I g iv es  th a g en g u id elin e s   o n   th way   to   d eter m in t h b est  n u m b e r   o f   clu s ter s .   I n   o r d er   to   g iv s tr o n g   an d   o b jectiv esti m ates  o f   th p e r f o r m an ce ,   K - f o l d   cr o s s - v alid atio n   is   u s e d   o n   s er ies  o f   d ata  s p lits .   T h s et  o f   f ea tu r es  ( X)   an d   lab els  o f   th s am ( y )   ar d iv i d ed   in t o   tr ain in g   ( Xtr ain ,   y tr ain )   an d   test   ( Xtest,  y test )   s et.   Evaluation Warning : The document was created with Spire.PDF for Python.
                      I SS N :   2 5 0 2 - 4 7 5 2   I n d o n esian   J   E lec  E n g   &   C o m p   Sci Vo l.  4 3 ,   No .   1 ,   Ju ly   20 2 6 :   299 - 31 3   304   E p s ilo n - g r ee d y   s tr ateg y   is   tak en   to   p ick   th b est  ac tio n   an d   tar g et  Q - v alu es  ar ca lc u lated   in   B ellm an   eq u atio n .   T h e   ev alu a tio n   m etr ic   co n s is ts   o f   ac c u r ac y   an d   F1 - s co r e,   a n d ,   tem p o r ar ily ,   is   im p lem en te d   b y   g en er atin g   th e   f ee d b ac k   s ig n al  r   o f   t h a g en u s in g   g r o u n d - tr u th   lab els.  T h leg alit y   o f   th e   ap p r o ac h   is   th lack   o f   c o n tr o l   o v e r   th e   d e cisi o n s   m ad b y   th e   ag en t   in   h is   o p er atio n ,   wh er e   lear n in g   is   p r o m o ted   o n ly   b y   lab elled   in f o r m atio n .   On ce   t h tr ain in g   e n d s ,   a   co n f u s io n   tab le  will  b c r ea t ed   t o   d eter m in th e   lev el   o f   ef f icien cy   o f   th e   s y s tem   in   s ep ar atin g   n o r m al  a n d   m alicio u s   n etwo r k   tr a f f ic.   T h ad a p tiv clu s ter in g   is   in   tu r n   in teg r ated   s o   th at  it  ca n   d ea with   s u ch   h ar s h   class   im b alan ce   o f   th N - B aI o T   d ata  s et  as  th d y n am ic  K   wo u ld   all o th e   ag en t o   d et ec m in o r   p atter n s   o f   attac k   t h at  co u ld   b o v er lo o k ed   b y   g lo b al  class if icatio n .   T h b atch   n o r m aliza tio n   a n d   d r o p o u e x p er im e n tal  test s   wer d o n e,   b u n eith er   o f   th two   f ac ilit ated   co n v er g en ce   s p ee d   o r   s tab il ity   s ig n if ican tly .   T h ey   wer e   o m itted   to   m ak e   th e   ar ch i tectu r s im p le  an d   co m p u tatio n ally   ef f icien t to   t h ed g d e v ices o f   th I o T   with o u t sacr if icin g   th d etec tio n   p er f o r m a n ce .   T h p r o p o s ed   alg o r ith m   was  co d ed   in   Py th o n   with   th e   h el p   o f   th e   T en s o r Flo f r am ewo r k ,   a n d   t h e   k ey   co m p o n e n ts   o f   th s im p le   r ein f o r ce m en lear n in g   o f   ex p er ien ce   r ep lay ,   r ewa r d   s elec ti o n ,   ac tio n   s elec tio n   p o licy ,   an d   s u cc ess iv u p d ate s   o f   Q - n etwo r k s   wer u tili ze d .   T h in p u f ile  was  N - B aI o T   d ata,   an d   th K - m ea n s   p r o v is io n in g   m o d el  th at  was  u tili ze d   wa s   em p lo y ed   to   d r aw  r elatio n   b etwe en   th ac tio n s   o f   th ag en t.  T r ain in g   was  ca r r ied   o u o n   1 0 0   ep is o d es,   an d   p er f o r m an ce   was  e v alu ated   b y   u s in g   th e   1 0 - f o ld   cr o s s - v alid atio n   in   o r d er   t o   ad d   s tr en g th   a n d   s tab ilit y   to   it.  Fig u r 4   also   d em o n s tr ates  th o v er all  wo r k f l o a n d   lo o p   o f   in ter ac tio n   b etwe en   t h ag en a n d   en v ir o n m en t,  in clu d in g   th e   ch o ice  o f   ac tio n ,   o b s er v atio n   o f   th e   s itu atio n ,   an d   Q - v alu u p d ate  th r o u g h   th n eu r al  n etwo r k .   T h m o d el  ca n   co m m u n icate   with   th I n ter n et - of - t h in g s   e n v ir o n m en b y   m o n ito r in g   th cu r r en s itu atio n   o f   th n etwo r k   tr af f ic  ( S)  a n d   m ak in g   a n   ac tio n   ch o ice  ( A) ,   wh ich ,   ac c o r d i n g   to   an   e - g r ee d y   ex p lo r atio n   p o licy ,   e v alu ates  clu s ter   s ize  o f   KK.   T o   r ew ar d   th e   ag e n t,  a   r ew ar d   ( R )   is   o b tain e d   wh e n   th e   ch o s en   ac tio n   is   im p lem en ted   b ased   o n   th m etr ics o f   class if icatio n   p er f o r m an ce ,   e. g . ,   ac c u r ac y   an d   F1 - s co r e.   E ac h   ex p er ien ce   is   ad d e d   to   a   r ep lay   b u f f er ,   wh ich   is   s u b s e q u en tly   u tili ze d   to   u p d ate  t h e   weig h ts   o f   th e   Q - n etwo r k   an d   en h an ce   th lea r n in g   p r o ce s s   b y   m ak in g   it  m o r s tab le  an d   co n v e r g en t.  T h way   th ag en t   p r o v id es  f ee d b ac k   o n   th r e war d ,   r e - u p d ates  Q - v alu es,  b eh av es  in   th en v i r o n m e n t,  a n d   s elec ts   th n ex ac tio n   is   f ir s d escr ib ed   i n   d e tail  in   Fig u r e   4 ,   an d   illu s tr at e s   th DQL   tr ain i n g   lo o p .   T h is   tr ain in g   p r o ce s s   in v o lv es  ex p er ien ce   r e p lay   a n d   tr ies  to   b alan ce   b etwe en   ex p lo r atio n   an d   ex p lo itatio n ,   s o   th at  it  ca n   allo w   o p tim al  lear n in g   o f   t h attac k   p atter n s   an d   r esp o n s iv en ess   to   n o v el  attac k   p atter n s .           Fig u r 4 .   I te r ativ tr ain in g   lo o p   u s in g   co m p o s ite  r ewa r d   f ee d b ac k       2 . 5 .     Det ec t io n a nd   f lo o f   I o T   tr a f f ic   T h p r o p o s ed   DDo attac k - d etec tin g   s y s tem   is   d ep lo y ed   a s   DQL - b ased   tr a f f ic  class if ier   th at  h as  th p o ten tial to   s o r t t h n o r m a l a n d   m alicio u s   tr af f ic.   T h o v er all  p r o ce s s   o f   th m o d el  is   d ep icted   in   Fig u r 5 .           Fig u r e   5 .   DQL - b ased   m o d el  f o r   DDo de tectio n   in   I o T   t r af f ic   Evaluation Warning : The document was created with Spire.PDF for Python.
I n d o n esian   J   E lec  E n g   &   C o m p   Sci     I SS N:   2502 - 4 7 5 2       Dee p   lea r n in g   a l g o r ith fo r   d etec tin g   DDo S   a tta ck s   o n   I o T d ev ices  ( La n a   K a mla   A h m ed )   305   2 . 6 .     E x perim ent a e qu ipm e nt    Py th o n   3 . 1 2 . 3   was  u s ed   as  th im p lem en tatio n   lan g u ag f o r   th p r o p o s ed   m o d el,   en a b lin g   s cien tific   co m p u tin g   an d   m ac h i n e - lear n in g - b ased   r esear ch   in   m o d e r n   en v ir o n m en t.  T e n s o r Flo 2 . 1 7 . 0   was  u s ed   to   d ev elo p   an d   tr ain   th e   DQL   m o d el,   as  it   is   m atu r e   an d   co m p atib le  with   T en s o r Flo w   L ite,   e n ab lin g   d ep lo y m en t   o n   r eso u r ce - lim ited   I o T   ed g e   d ev ices.  Scik it - lear n   was  u s ed   to   p r ep r o ce s s   an d   clu s ter   d ata  with   K - m ea n s ,   wh ile  Nu m Py   an d   Pan d as  ass is ted   with   d ata  m an ip u latio n ,   an d   Ma tp l o tlib   was  u s ed   f o r   p e r f o r m an ce   v is u aliza tio n .   All  ca lcu latio n s   wer p er f o r m e d   o n   m u lti - c o r I n tel  C o r i7 - 1 0 7 5 0 C PU  at  2 . 6 0   GHz   with   1 6 GB   o f   R AM   an d   5 1 2   G B   SS D,   s im u latin g   r ea lis tic  h ig h - e n d   I o T   ed g e   g atew ay .   T h ex p e r im en tal  p r o ce s s   in v o lv e d   p r o ce s s in g   9 1   d is tin ct  d ata  f iles   f r o m   th e   N - B aI o T   d ataset.   T h m ea n   c r o s s - v alid atio n   tim e   o f   6 5 . 6   s ec o n d s ,   o r   6 5 . 6   s ec o n d s   p er   f o ld   o f   cr o s s - v alid atio n ,   im p lied   to tal  tr ain in g   a n d   cr o s s - v alid atio n   tim o f   ab o u 3 . 9   h o u r s   in   t h e   f u ll  c o u r s o f   th 1 0 - f o ld   cr o s s   v alid atio n .   T h h i g h est  m e m o r y   co n s u m p tio n   was m ain tain ed   at  5 0 0   MB,  an d   th av er ag i n f er en ce   tim was 0 . 0 5   m s   p er   in s tan ce .   T h e s v alu es v er if y   th at   th m o d el  is   co m p u tatio n ally   ef f icien t a n d   ca n   b u s ed   in   lo w - r eso u r ce   an d   lo w - m em o r y   I o T   ap p licatio n s .     2 . 7 .     P er f o r m a nce  m e t rics   Fo u r   p o p u lar   m etr ics,  n am el y ,   ac cu r ac y ,   p r ec is io n ,   r ec all,   an d   F1 - s co r e,   wer e   u s ed   to   ass es s   th class if icatio n   p er f o r m an ce   o f   th p r o p o s ed   DQL   m o d el  [ 3 3 ] .   Su ch   m ea s u r es  ca n   b e   u s ed   to   d eter m in th e   ef f ec tiv en ess   o f   th m o d el  i n   d ec id in g   th p r esen ce   o f   m alig n an an d   leg itima te  n et wo r k   tr af f ic  in   th d etec tio n   o f   I o T - b ased   DDo S.   Acc u r ac y   is   m ea s u r o f   th f r ac tio n   o f   tr u s am p les in   s et   o f   s am p les:     A c c ura c y =  +   +  +  +    ( 1 )     T h ac cu r ac y   is   m ea s u r e   o f   th n u m b er   o f   p o s itiv es  th a ar p r e d icted   co r r ec tly   o u o f   all  th p r ed icted   p o s itiv es:     Pr e c ision =   +                                                      ( 2 )     R ec all:  R ec all  i s   r atio   th at   is   u s ed   to   m ea s u r e   h o w   m an y   p o s itiv ca s es  ar co r r e ctly   id e n tifie d   o u t   o f   all  th tr u e   p o s itiv ca s es:     R e c a l l =   +                                                                                        ( 3 )     F1 - s co r is   th h ar m o n ic  m ea n   o f   p r ec is io n   an d   r ec all:     1 = 2 × Pre c isi o n × Re c a l l Pre c isi o n + Re c a l l                                                               ( 4 )     All  o f   th ese  m etr ics  r ep r esen s tr o n g   s y s tem   th at  ca n   b u s ed   to   ev alu ate   th d etec tio n   p o ten tial  o f   th m o d el   in   p a r ticu lar   ca s es  o f   im b alan ce d   d ata,   wh en   t h p r esen ce   o f   attac k   tr af f ic  s am p l es  m ay   o v er wh elm   th n o r m a l   d ata  tr af f ic.   T h ese   m etr ics  ar ap p lied   in   o r d er   to   co m p ar th p r o p o s ed   DQL - K - m ea n s   m o d el   with   m u ltip le  b aselin es,  s u ch   as  R an d o m   Fo r est,  C NN,   an d   Q - L ea r n in g ,   in   o r d e r   to   g u a r a n tee  m eth o d o lo g ical  r ig o r .   Als o ,   an   ab latio n   s tu d y   is   ca r r ied   o u with   th u s o f   t h s am m etr ics  to   ch ec k   th e   p er f o r m an ce   o f   th m o d el  wh e n   th e   ad ap tiv clu s t er in g   a n d   t h elem en ts   o f   c o m p o s ite  r ewa r d   ar a b s en t.  All  r esu lts   ar g iv en   as  th m ea n   o f   th 1 0 - f o ld   cr o s s - v alid atio n   f o ld s   in   o r d er   to   co n s id er   s tatis tical  v ar i ab ilit y   a n d   o b tain   s ig n if ican ce ,   alo n g   with   th s tan d ar d   d ev iatio n   ( ± σ ) .       3.   RE SU L T S AN D I SCU SS I O N   I n   th is   s ec tio n ,   th r esu lts   o f   th ex p e r im en o f   th e   p r o p o s ed   DQL - b ased   f r am ewo r k   a r e   p r o v id ed .   T h is   ev alu atio n   h as  b ee n   p er f o r m ed   b ased   o n   th - B aI o T   d ata  s et  an d   th co m m o n   p e r f o r m an ce   m ea s u r es  o f   th N - B aI o T ac cu r ac y ,   F1 - s co r e,   tr u p o s itiv es  ( T P),   tr u n e g ativ es  ( T N) ,   f alse  p o s it iv es  ( FP ) ,   an d   f alse   n eg ativ es  ( FN) .   T h co n v er g en ce   v is u aliza tio n s   an d   th t em p o r al  p r o g r ess io n   o f   th m o d el  d escr ib e   th e   lear n in g   p r o ce s s   o f   th e   m o d el  in   a   s er ies  o f   ep is o d es  an d   d is p lay   s tab ilit y   an d   s tab le  co n v er g en ce   b eh av i o u r Als o ,   m ea s u r es  o f   co m p u tati o n al  ef f icien cy   we r m ad i n   ter m s   o f   tim ( in   s ec o n d s )   tak en   to   tr ain   an   ex am p le,   p er - s am p le  in f er en c laten cy ,   an d   m em o r y   co n s u m p tio n .   All  th ese  f in d in g s   ar co m p r eh e n s iv r ev iew  o f   th e   ef f ec tiv en ess   an d   ef f icien cy   o f   th e   s u g g ested   DQL   m o d el.       Evaluation Warning : The document was created with Spire.PDF for Python.
                      I SS N :   2 5 0 2 - 4 7 5 2   I n d o n esian   J   E lec  E n g   &   C o m p   Sci Vo l.  4 3 ,   No .   1 ,   Ju ly   20 2 6 :   299 - 31 3   306   3 . 1 .   Cla s s if ica t io perf o r m a nce  re s ults   Acc u r ac y ,   p r ec is io n ,   r ec all,   a n d   F1 - s co r e   ar f o u r   s tan d a r d   m ea s u r es  th at  wer em p lo y ed   to   test   th e   p r o p o s ed   DQL   m o d el.   T h v alu es  u n d er   th ese  m etr ics  h av b ee n   ca lcu lated   th r o u g h   t h elem en ts   o f   th e   co n f u s io n   m atr ices  b ec a u s t h elem en ts   h av b ee n   s u m m ar ized   in   T ab le   3   [ 3 3 ] .   T h p r o p o s ed   m o d el  h as  d em o n s tr ated   h ig h   lev el  o f   d etec tio n   an d   an   ac ce p tab le  p r o p o r ti o n   o f   p r ec is io n   an d   r ec all,   as  p r esen ted   in   th r esu lts   o f   t h ex p er im en t.   T h ese  v a lu es  o f   F1 - s co r ar e   also   an   e x ce llen in d icatio n   t h at  th m o d el  ca n   d is tin g u is h   b etwe en   th m alicio u s   tr af f ic  an d   th r ea I o T   n etwo r k   tr af f ic.   Fig u r 6   s h o w s   th d is tr ib u tio n   o f   r ea lized   an d   esti m ated   class es  o f   a n   a v er ag e   f o ld   o f   th e   c r o s s - v alid atio n   p r o ce d u r e   o f   1 0   f o ld s .   Fig u r e   6   in d icate s   th at  f alse  p o s itiv es  wer lo ( 2 1 , 6 3 5 )   in   th m o d el.   T h is   is   ess en tial  to   I o T   en v ir o n m en s in ce   h ig h   f alse  p o s itiv will   r esu lt  in   b lo ck in g   o f   leg itima te  tr af f ic,   lik s m ar s en s o r s   u p d ates  o r   u s er   co m m an d s ,   to   b r in g   d o wn   th s er v ice.   On   th o th er   h an d ,   th Fals Neg ati v ( 9 , 8 9 6 )   is   lo w,   wh ich   m ea n s   th at  th m ajo r ity   o f   m alicio u s   DDo tr af f ic  will  b ap p lied   p r i o r   to   d am ag i n g   th eq u ip m en t.  E x am in in g   th ese  o v er lo o k ed   in s tan ce s   o f   attac k s   h as  s h o wn   th at  th m o d el  s o m etim es  h as  d if f icu lties   s tealth y   U DP  s ca n n in g ,   wh ic h   esch ews  th h ig h - f r eq u en cy   o f   th h ea r tb ea s ig n als  o f   m al icio u s   I o T   d ev ices.  T h ad a p tiv q u ality   o f   th DQL   ag en h o wev er   ass is ts   i n   r ed u cin g   th is   b y   ch an g i n g   th clu s ter   r ep r esen tatio n   as  th b eh av io r   o f   th e   attac k   ad ap ts .       T ab le  3 .   B in ar y   class if icatio n   p er f o r m an ce   o f   th c o n f u s io n   m atr ix     P r e d i c t e d   p o si t i v e   P r e d i c t e d   n e g a t i v e   A c t u a l   p o si t i v e   4 , 5 0 0 , 0 0 0   9 , 8 9 6   A c t u a l   n e g a t i v e   2 1 , 6 3 5   2 , 0 0 0 , 0 0 0           Fig u r 6 .   C o n f u s io n   m atr i x   o f   th p r o p o s ed   m o d el ,   wh ich   h a s   h ig h   v alu es o f   class if icatio n   an d   lo v al u es in   er r o r s       3 . 2 .     Ana ly s is   o f   t r a ini ng   ev o lutio n a nd   co nv er g ence   On   th s tab ilit y   o f   th p r o p o s ed   DQL   m o d el,   we  p r esen t   m ea s u r es  o f   p er f o r m an ce   th r o u g h   th e   co u r s o f   all  ep is o d es  an d   f o l d s   o f   ten - f o l d   cr o s s - v alid ati o n   s tr ateg y .   Als o ,   in   tr ain in g   an d   test in g ,   ac cu r ac y   an d   F1 - s co r e   wer em p lo y e d   to   test   th d y n am ics  o f   lear n in g   in   th c o u r s o f   tim e.   Fig u r 7   s h o ws  th tr ain in g   d ev elo p m en t,  i n   wh i ch   th ac c u r ac y   an d   F1 - s co r e   lev el  o f f   ea r ly   d u r in g   th e   d e v elo p m en t.   Su ch   a   q u ick   co n v er g en ce   is   ex p lain e d   b y   th co m p o s ite  r ewa r d   f u n ctio n   th at  g iv es r ich   f ee d b ac k   to   th ag en t so   th at   it  ca n   r ap id ly   d if f er en tiate  b et wee n   b en i g n   tr af f ic   an d   d i f f er en DDo attac k   v ec to r s   in   th N - B aI o T   d ata  s et.   T h ese  m ea s u r em en ts   ten d   to   b ex tr em ely   elev ated   in   m aj o r ity   o f   th ep is o d es  an d   g r ea tly   v ar y   b ec au s o f   p er s o n al  d is cr ep an cy   in   th ex p er im en tal  f u n ctio n in g   o f   th DQL   alg o r ith m .   Su cc ess f u lear n in g   ca n   b attr ib u ted   to   th is   tr en d ,   an d   it sh o ws th at  th m o d el  is   co n v er g en t in   n at u r e.           Fig u r 7 .   DQL   t r ain i n g   f o r   a cc u r ac y   an d   F1   s co r e   Evaluation Warning : The document was created with Spire.PDF for Python.
I n d o n esian   J   E lec  E n g   &   C o m p   Sci     I SS N:   2502 - 4 7 5 2       Dee p   lea r n in g   a l g o r ith fo r   d etec tin g   DDo S   a tta ck s   o n   I o T d ev ices  ( La n a   K a mla   A h m ed )   307   Fo llo win g   th tr ain in g   ev alu at io n ,   p er f o r m a n ce   m ea s u r es  o f   test in g   ar illu s tr ated   in   Fig u r 8 ,   wh er e   ac cu r ac y   a n d   F1 - s co r ar e   p lo tted   with   th v ar io u s   f o ld s   o f   c r o s s - v alid atio n   o f   th e   test in g .   Fig u r 8   r ep r esen ts   th test in g   p er f o r m a n ce   o f   th e   1 0 - f o ld   c r o s s - v alid atio n .   T h f ac th at  th v ar ian ce   am o n g   th f o ld s   is   m in im al  ( it  is   o b s er v e d   in   th tig h clu s ter in g   o f   th e   d a ta   p o in ts )   p r o v es  th f ac t h at  th e   m o d el  h as g o o d   g e n er aliza tio n   p o wer .   T h is   will  m ak th s y s tem   ef f ec tiv wh en   n ew,   u n k n o wn   I o T   d e v ices  o r   d if f e r en co n d itio n s   o f   n etwo r k   tr af f ic  ar in v o lv ed ,   wh ich   wer n o t e x p er ie n ce d   d u r in g   th e   tr ain in g   f o ld .   T h an aly s is   o f   th p r o g r ess io n   o f   th m ax im u m   Q - v alu e s   o f   th ag en o v er   ep is o d es   was  al s o   m o n ito r ed   i n   th ex p er im en to   ex am in co n v er g en ce .   C o n v er g e n ce   o f   th lea r n in g   p r o ce s s   is   as s u m ed   to   h av b ee n   s u cc ess f u in   th e   ev en o f   a   co n s tan o f   s tab le   tr en d   in   t h Q - v al u es;  th tr en d   in   th is   ca s is   p r esen ted   in   Fig u r 9 .           Fig u r 8 .   Acc u r ac y   a n d   F1 - s c o r f o l d s   f o r   test in g           Fig u r 9 .   Ma x im u m   Pre d icted   Q - v alu Per   T r ai n in g   E p is o d e       T im d y n a m ics  o f   th m a x i m u m   Q - v alu es  in   s eq u en tial   tr ain in g   s ess io n s   we r s ee n   to   test   th s tab ilit y   o f   th lear n in g   p r o ce s s .   F ig u r 9   s h o ws  th m ax im u m   p r ed icted   Q - v alu es  ar v a r y in g   in tr a - e p is o d as  o p p o s ed   to   s tr aig h lin e.   T h is   is   g o o d   s ig n   th at  th ep s ilo n - d ec ay   m ec h a n is m   is   b eh av in g   p r o p er ly ,   th e   ag en k ee p s   s ea r ch i n g   h i g h - d im en s io n al  s p ac e   o f   I o T   tr af f ic  with o u t   f allin g   i n to   th e   lo ca o p tim o f   t h s tr ateg y ,   s o   th at  th e   r esu ltin g   p o licy   will b r esis tan to   s tealth y ,   a d ap tab le   attac k   e x ec u tio n s .   No n eth eless ,   th e   ac cu r ac y   an d   F1 - s co r cu r v es in   Fig u r 9   d o   n o t d if f er   b etw ee n   th cr o s s - v alid atio n   f o ld s ,   wh ich   in d icate s   th p r esen ce   o f   g o o d   r esu lts   in   ter m s   o f   lear n in g .   T h cr o s s - v alid atio n   f o ld   r e q u ir ed   a p p r o x im ately   6 5 . 6   s ec o n d s   o f   wall - clo ck   tim e;  th e   m ea n   in f er en ce   tim was  0 . 0 5 m s   p e r   in s tan ce ,   a n d   t h m o d el  o cc u p an cy   was  ar o u n d   5 0 0   MB,  in d icatin g   th co m p u tatio n al  ef f icien cy   o f   th m o d el  an d   m ak in g   it  u s ef u in s tan ce   to   b u s ed   in   lo w - m em o r y   I o T   ap p licatio n s .     3 . 3 .   Q - lea rning   re s ults   T h Q - lear n in g   m o d el  p r o p o s ed   h as  r ea ch ed   th ac cu r ac y   o f   th class if icatio n   o f   9 2 . 3 1 ,   p r ec is io n   o f   9 2 . 3 1 ,   r ec all   o f   9 2 . 1 3 ,   an d   F1 - s co r o f   9 2 . 4 0   b y   ad j u s tin g   t h lear n in g   r ate  ( 0 . 0 5 )   an d   d is co u n f ac to r   ( 0 . 9 5 )   o v er   th n u m b er   o f   iter atio n s .   T o   f in d   th b e st - k n o wn   p ar a m eter s ,   th v alu es  o f   th h y p e r p ar am eter s   α   an d   γ   wer o b tain ed   th r o u g h   t h i n tu itiv ( n o n - co m p u tatio n al )   s ea r ch   o f   s elec ted   h y p er p a r am eter   s p ac e,   b y   Evaluation Warning : The document was created with Spire.PDF for Python.
                      I SS N :   2 5 0 2 - 4 7 5 2   I n d o n esian   J   E lec  E n g   &   C o m p   Sci Vo l.  4 3 ,   No .   1 ,   Ju ly   20 2 6 :   299 - 31 3   308   ass es s in g   th v alid atio n   p er f o r m an ce .   Su b s tan tial  en h a n ce m en o f   th e   u n tu n ed   b aselin was  o b tain ed   b y   th is   tu n in g   p r o ce s s .   I n   ad d itio n ,   1 0 - f o ld   cr o s s - v alid atio n   ( K   1 0 )   u s in g   th e   o p tim ized   h y p er p ar am eter s   r esu lted   in   an   ac c u r ac y   o f   9 5 . 2 4 ,   p r ec i s io n   o f   9 5 . 0 2 ,   r ec all  o f   9 5 . 2 4 ,   an d   an   F1 - s co r o f   9 5 . 1 3 .   T a b le  4   also   g iv es  a   m o r s p ec if ic  co m p a r is o n   o f   ev alu atio n   m ea s u r es  o f   th p r o p o s ed   Q - lear n in g   m o d el  as  is   an d   af ter   th e   ap p licatio n   o f   th K - f o ld   an d   K - m ea n s .       T ab le  4 .   Q - lear n in g   m o d els r e s u lt   P e r f o r ma n c e   m e t r i c s   Q - l e a r n i n g   Q - l e a r n i n g   ( K - f o l d   a n d   K - m e a n )   A c c u r a c y   9 2 . 3 1 %   9 5 . 2 4 %   P r e c i s i o n   9 2 . 3 1 %   9 5 . 0 2 %   R e c a l l   9 2 . 1 3 %   9 5 . 2 4 %   F1 - S c o r e   9 2 . 4 0 %   9 5 . 1 3 %       I is   s h o wn   in   th tab le  ab o v th at  th p er f o r m an ce   m ea s u r es  o f   th Q - lear n i n g   m o d el     an d   th m o d i f ied   m o d el  a r co m p ar e d ,   in clu d in g   K - f o ld   cr o s s - v alid atio n   an d   K - m ea n s   clu s ter in g .     T h cr o s s - v alid atio n   p lu s   th K - m ea n s   is   f air ly   s u cc es s f u i n   en h an cin g   all  ass e s s m en m ea s u r es,  esp ec ially   th r ec all  an d   ac cu r ac y .   T h r esu lts   o f   th co m p ar is o n s   ar r ep r esen ted   in   Fig u r 1 0 th er ef o r e,   th f o c u s   is   p u t o n   t h b etter   s en s itiv ity   o f   th m o d if ied   m o d el  t o   id en tif y   th in d iv id u al.   Fig u r 1 0   d em o n s tr ates  th in f lu en ce   o f   h y p er p ar am eter   o p tim izatio n   o n   th p er f o r m an ce   o f   a     Q - lear n in g   m o d el  to   id en tif y   DDo attac k s   s ta tin g   th at  it  g en er ally   in cr ea s es  th ab ili ty   to   d etec th ese   attac k s .   So m ex tr p er f o r m a n ce   wo u ld   b o b tain ed   with   K - f o ld   cr o s s - v alid atio n   a n d   K - Me an s   clu s ter in g   th at  wo u ld   r ed u ce   o v er f itti n g   a n d   p r o v id e   g r ea ter   r eg u la r ity   ac r o s s   d iv is io n s   o f   d ata.   Nev e r th eless ,   th eir   s p ec if ic  d ef icien cies  r em ain   th e   lim its   in   th p r ec is io n   an d   th eir   a b ilit y   to   d etec attac k s ,   w h ich   lea d s   to   th n ec ess ity   to   in tr o d u ce   m o r s o p h is ticated   ap p r o ac h es  th at  will  b ab le   to   wo r k   o n   th n ew  d y n a m ica lly   ev o lv in g   attac k   p atter n s   o n   I o T   n etwo r k s .           Fig u r 1 0 .   Q - lear n in g   p er f o r m an ce   an d   i n teg r atin g   (K - Fo ld   a n d   K - Me an )       3 . 4 .   DQ L   r esu lt s   lear n in g   r ate  ( 0 . 0 0 1 )   was  u s ed   s o   th at  th weig h ch an g ed   with o u ac ce ler atio n h ig h er   v alu led   to   g lo b b in g ,   an d   lo wer   v alu led   to   n o n - lea r n in g .   W h av f ix ed   th d is co u n f ac t o r   ( 0 . 9 5 )   to   en a b le    th ag en to   ca p tu r f u tu r n etwo r k   s tates,  wh ich   is   ap p r o p r iate  to   ca p tu r th tem p o r al  d ep en d en cies  o f   m u lti - s tag DDo attac k s .   T h ep s ilo n - d ec ay   s ch ed u le  b eg i n s   with   an   ep s ilo n   o f   1 . 0   an d   d ec r ea s es  to   0 . 0 1   in   th f ir s 1 0 0   ep is o d es,  g iv in g   p r elim in ar y   b r o a d   s ea r ch   o f   th h ig h - d im e n s io n al  N - B aI o T   p r o p e r ties   af ter   wh ich   th ag e n is   co n s id e r ed   to   s witch   to   an   eq u ilib r iu m   ex p l o itatio n   p o licy .   T h r esu lt  o f   t h ese  o p tim izatio n s   was  b aselin p er f o r m a n ce   f o r   th m o d e o f   9 5 . 8 3 ac c u r ac y ,   9 5 . 8 3 r ec all,   9 3 . 7 5 %   p r ec is io n ,   an d   an   F1 - s co r o f   9 4 . 4 4 %.  L ater   o p tim izatio n s   in clu d ed   a d d in g   K - m ea n s   cl u s ter in g   to   th e   DQL   tr ain in g   p r o ce s s   an d   u s in g   K - f o ld   cr o s s - v alid atio n .   K - m ea n s   clu s ter in g   was  u s ed   to   r ed u ce   n o is an d   m ak e   class es m o r s ep ar ab le,   th er eb y   m ak in g   th d ata  m o r a m en ab le  to   DQL   tr ain in g .   T h c r o s s - v al i d a ti o n   p r o c ess   was  p e r f o r m e d   u s i n g   K   d if f e r e n d ata   p a r tit io n s ,   a n d   th m o d el  w as   v al id at ed ,   p r o v i d i n g   co n s is te n cy   a n d   e n a b l in g   g e n e r a liz ati o n   t o   n ew   d at a.   T h m o s t   p o s iti v e   r es u lts   w er e   ac h ie v e d   wi th   t h e   i n t eg r a te d   m o d el,   w h i ch   c o m b in e d   o p ti m iz ed   h y p er p a r am ete r s ,   K - m ea n s   cl u s te r i n g ,   a n d     K - f o ld   cr o s s - v ali d ati o n .   T h is   s etu p   ac h ie v ed   9 8 . 9 5 %   ±   0 . 1 7 a cc u r a cy ,   9 8 . 7 2 %   ±   0 . 1 7 %   p r ec is i o n ,   9 8 . 9 5 %   ±   0 . 1 7 %   r ec all ,   a n d   a n   F1 - s c o r e   o f   9 8 . 7 3 %   ±   0 . 1 7 % ,   w h ic h   w a s   b y   f a r   b et te r   t h a n   t h e   b as eli n an d   t h im p r o v ed   DQL   m o d e ls .   t - t est   c o m p ar i n g   th F 1   Sc o r es   f r o m   1 0 - f o l d   c r o s s - v a li d at io n   p r o v i d e d   ev id e n c o f   t h e   s tatis t ic al  s i g n i f ic a n ce   o f   th is   i m p r o v e m e n t .   T h e   p - v al u o b ta in e d   f r o m   th t est w as  5 . 0 8 × 1 0 ^ - 5 ,   i n d ica ti n g   t h at   Evaluation Warning : The document was created with Spire.PDF for Python.