Int ern at i onal  Journ al of Ele ctrical  an d  Co mput er  En gin eeri ng   (IJ E C E)   Vo l.   8 , No .   6 ,  Decem ber   201 8,   pp. 4 505~ 4518   IS S N: 20 88 - 8708 ,  DOI: 10 .11 591/ ijece . v8 i 6 . pp 4505 - 45 18          4505       Journ al h om e page :  http: // ia es core .c om/ journa ls /i ndex. ph p/IJECE   SCDT: F C - NNC - structur ed  Com plex Dec ision T echn iqu e fo r  Gene An alys i s Us ing Fuz zy Clust er based  Nearest  Neighb or   Classifie r       Sudha  V . 1 ,  Gi ri ja mm a H .   A . 2   1 Depa rtment of I S&E,   RNS   Insti t ute   of   T ec hnolo g y ,   Indi a   2   Depa rtment   of   Com pute r  Scie n ce   & Engi ne eri n g,   RNS   Instit u te  of  Technol og y ,   I ndia       Art ic le  In f o     ABSTR A CT    Art ic le  history:   Re cei ved   Feb   12 , 201 8   Re vised  Jun  1 7 , 201 8   Accepte d   J un   2 0 , 201 8       In  m an y   d isea s es  cl assifi cation   an  a cc ur at e   ge ne  an aly s is  is  nee ded ,   for   which  select ion  of  m ost  informa ti ve  gen es  is  ver y   importan t  and  it   req uir e  a   te chn ique   of  dec ision  in  complex  cont ext   of  ambiguity .   Th e   tra dit ion al  m et hods  inc lude  for  sele ct ing  m ost  signifi ca nt   gene   inc lud es  som e  of  the   stat isti ca l  an aly s is  namel y   2 - Sam ple - T - te st  (2STT ),   Ent rop y ,   Sign al   to  Noise  Rat io  (SN R).   Thi s  pape r  eva lu a te s  gene   select i on  and  cl assificat ion  on  th e   basis  of  ac cur a te  gene   select ion  using  struct ure d  complex  dec isio n  te chni qu e   (SCD T)  and  cl a ss ifi e s  it   using  f uzzy   c luste r  b ase d  nea r est  ne igh borc la ss ifier   (FC - NN C).   The  eff e ct iv ene ss   of  the  propose d  SC DT  and  FC - NN C  is   eva lu at ed   for  l e ave   on e  out   cr oss   val ida t ion  m et ric (LOOCV )  al ong  wi th   sensiti vity ,   spec i fic ity ,   pr ec ision  and  F1 - score   wi th  four  diff er ent  cl assifi ers   namel y   1)  Radia l  Basis  Functi on   (RBF ),   2)  Multi - lay er  p ercept io n(MLP),  3)  Feed  Forw ard (F F)  and  4)  Suppo rt  vec tor  m ac hin e(SVM )  for  thre e  diffe re n t   dat ase ts  of  DL BCL,   L euke m ia  and  Pros ta t e  t um or.   The   prop osed  SC DT   &FC - NN C  exhi bit s  superior   result   for  bei ng  conside red   m ore   ac cur ate   dec ision   m ec han ism .   Ke yw or d:   Fu zzy  classi fic at ion   Gen e  an al ysi s   Gen e  selec ti on   Ma chine  le a rn i ng   Mi cro  a rr ay   da ta     Copyright   ©   201 8   Instit ut e  o f Ad vanc ed   Engi n ee r ing  and  S cienc e .    Al l   rights re serv ed .   Corres pond in g  Aut h or :   Sudh a  V . ,    Dep a rtm ent o f Info rm at ion  Sc ie nce &   En gine erin g,     RNS  In sti tute  of Tech nolo gy,  Ben galuru, I ndia .   Em a il :  su dh a vi nayakam @g m ai l.com       1.   INTROD U CTION     The  accu racy  of   dia gnos is  is  the  basis  fo r  t he  perfect   treat m ent  pr ocess  t o  be  ad op te d  especial ly   in  the  case  of   fa ta l  disease  li ke  cancer s,  le ukem ia   and   pro strat e  tum or   et c.  Alon g  with   the  hist op at holo gy ,   m edical   rad iol og y  a nd  im aging  te ch niques,   the  m ic r o - ar r ay   data  a naly sis  co uld   be  pr ov e n  quit e  hel pful  as   well   as  rig htf ul   if  eff ic ie nt  te chn i qu e s  of  an al ysi s  are  evo l ved   [ 1].  The  a ccur acy   of   dis ease  cl assifi cat ion   or  early  d ia gn os is  d e pends  up on,  how acc urat el y t he  ge ne o f  si gn i ficance is  s el ect ed.     The  D NA - m i cro a rr ay   data   analy sis  is   chall eng in g  in  bo th  as pects  of   sta ti sti cal l y  and   com pu ta ti on al ly   as  it   po ssess es  non - li near   no ise s  al ong  with  hi gh   dim ensio nalit y  of   low  sam ple  data  [2 ] .  Ma ny  ef forts  towa rd s   diseas e  diag nosis  pa rtic ularly   canc er,  t um or   et c,   cl assifi cat ion  ha ve  bee n  se en  i n  li te ratur e  [ 3 ] - [ 10 ] .  T he  sect io n  2  descr i bes  t he  insig hts  of   r el at ed  wor k.   V ario us   m achine  le arn in g  ap pro ache s   are  us e d  for  th e  cl assifi cat ion  wh ic h  incl ud e s  rad ia l  ba sis  f un ct io n  (RBF) ,  arti fici al   neur al   networ k  ( A NN),   su p port  vecto r   m achine  (SV M)  et c.  by  f orm ing   the  pro blem   as  bin ary  cl assifi cat ion .  T he  pr ob le m   of   dim ension   reducti on   for  sear chin g  m os t  sig nificant  ge ne  is  bein g  form ulate d  as  m any  pr oble m   sp aces  wh i c h  include s  1)  Mi xed  inte ger  pr ogram m ing   ( MIP),  2)  Bi o - i ns p i red  op ti m i zat ion   (BI O),  3)  Mi ning  as s ociat ion   ru le s  (MAR ),  a nd last  but  no t  the lea st 4 )  E nsem ble tec hn i que (ET ) [8 ] .   The  cl inica ll y  com pr ehe ns ive   m et ho d  requi res  ha nd li ng  hi gh   dim ensional   data  with  ver aci ty   and  no ise s  t o  ha nd le   a m big uity   duri ng   t he  rig ht   gen e  ca ndidat e  sel ect ion .  T hi s  pap e r  pr opose s  a  m echan ism   of   Evaluation Warning : The document was created with Spire.PDF for Python.
                          IS S N :   2088 - 8708   In t J  Elec  &  C om p  En g,   V ol.  8 , N o.   6 ,  Dece m ber  2 01 8   :   4505   -   4518   4506   structu re d  c omplex  decisi on  t echn i qu e   ( SCDT)  f or  f uzzy  cl us te rin g  neig hborh ood  cl us t er  ( FC - NC).  S ect ion   3  descr i bes  com plete   syst e m   m od el   f or   SC DT   &  FC - NC,  Se ct ion   4  de scrib es  about  three  diff e r e nt  m ic ro arr ay   dataset s.  Sect i on 5 il lustrate s  resu lt s a nd a na ly sis fo ll owed   by conclu sio n  i n  Sect io n 6.     1.1 .   B ackgr ound   The  acc ur at e  c lusterin g  of  th e  data  is  a  cha ll eng in g  an d  open  resear ch  pro blem   fo r  cl assifi cat ion s  sp eci al ly   us ing   super vise  le arn i ng.   A n  ex te ns ive  sur vey  is  con duct ed   to  under sta nd  the  eff ect ive ness  of   cl us te rin g  te ch niques  pa rtic ul arly   fo r  m edical   data  li ke  m ic ro   ar ray  ge ne   dataset   [11].   Fo r  t he  pur pose  of   tum or   diag no si s,  the  ap proac h  of   prof il in g  th e  gen e  acc urac y  is  co m par at ively   of   higher  reli abili ty   with  m or e  accuracy  tha n  that  of   t he  m e thod  ad opte d  by  the  m edica l  i m aging   te ch nique  of  m or phol og ic al   anal ysi s  of  tum or .  Tra diti on al ly   ad op te d  s up e rv ise d  l earn i ng   a ppr oa ches  fall s  int o  pitfal l  of  ac cur acy   du e  t o  few e r  sam ples  of   cancer  t ypes  exist  into  the  trai ning  dataset   of   ge ne  ex pr essi on  as  well   the  ov e rh ea ds   due  to  higher   data - dim ension al it y du e t o  la rg e  g e ne  e xpre ssion.    In   the  work   of  Lipowan g  et   al   aim s   to  sel e ct   few   num ber s  of   ge nes  to  c la ssify  the  cancer  from   the   m ic ro arr ay   dat a   to  m eet   the  go al   of   bala nci ng   tra de - off  a m on g  the  accu racy  as  well   as  m ini m iz at ion   of   the   com pu ta ti on al   com plexity   or   ov e r head s [ 3].  They  ha ve  use d  “feat ur e  im portance  ranki ng   s chem e”  for  the   accurate  or  sig nificant  ge ne  s el ect ion   a nd  f orm ulate d  the  c la ssi ficat ion   prob le m   as  ty pical   cl us te r  of  bi nar y   cl assifi cat ion   pro blem .  The  m achine  le arn i ng   a ppr oac hes   us e d  in  thei r  work  are  m ix  us e  of   fu zzy   ne ur al   netw ork  (FN N )  an d  SV M.  T he  dim ension   r edu ct io ns   ob ta i ned   wer e  getti ng   sam e  accuracy  on ly   by  sel ect ing   28  ge nes  as   co m par ed  to  16,063  ge nes  of  tra diti on al   m et ho d  of   t hat  tim e.  The  ty pical   dat aset   exp l or e d  f or   t he   ob s er vations  i nclu des  1)  Ly m ph om a  Data,  2)  SRB CT  D at a,  3)  Li ver   C ancer   Data,   an d  4)  GCM  dat a.  T hey  reco m m end ed  consi der i ng   th e  cooper at io n  aspects  betw ee n  the  ge nes  to  m ini m iz e  the  gen e  s ubset   for  m or e  accurate  predic ti on .    Fu rt her, the work which  ha s r efe r  to this includ es the wo rk b y  C hien - Pa ng et al   wh o  ha ve  introdu c e d  a  m et ho d  hybr idize d  us i ng   ge netic   al go rith m   and   dynam i cal ly   setting   up  the  pa ram et e r  for  sig nifican t  gen e   sel ect ion   an d  then  furthe r  use s  SV M  for  ve rificat ion   pur po s es  to  predi ct   gen e  sel ect ion e ff ic ie nc y  [1 2].  T he   dim ension   re duct ion  an d  feat ur e   sel ect ion  is  the  c or e   pr oble m   to  be  ha nd le d  as   ge ne  expressi on  m icr oa rr ay   (G EM A)  co ns i st  of   h undred   to  s om eti m e  th ou s an ds   of  the   featu res  in   a  ver y  sm al l  sa m ple  siz e.  These  hi gh   nu m ber s  of  fe at ur es  i n  a  sm al l  sa m ple  of   GEMA   m akes  it   of   ver y  high  dim ension   da ta .  The  c onve ntion a l   m et ho ds   a dopt ed  f or  feat ur e s el ect ion   w hich  is  al so   cal le d  a s  ge ne  sel ect io n  in  c ase of   t he   GEM A  analy s is  for  the cla ssici zat ion o f  t he diese s inclu des 1 ) G ai n  & Rel ie f , 2 )  Chi  Squa res,   3) Fishe r Sco re , and 4 )  Lass o et c.    The  ge neselect ion   m et ho d  is  cl assifi ed  into  three  cat eg or ie s  1)   s up e rv ise d,   2)  uns up e r vi sed  an d  3)  sem i - su pe rv ise d  on  the  basis   of   c orrespo nding   data  ty pes  of   1)   f ully   la beled,   2)   unla beled  a nd   3)   pa rtia ll y  la beled  res pec ti vely   for  cl as sific at ion   or  predict io n  of  cl asses  as   des cr ibed   by  [ 13 ] .   Fu rt her,  t he  fe at ur e   avail able  into   the  GEMA  s a m ples  are  cat egoriz ed  into   two  crit ic al   sel ect ion s  na m el y   red un da ncy  an d  releva ncy.  The   F ig ure   1 ,  s hows  the  ty pical   cl assifi cat ion   based   on   the   com bin at ion   of   these  tw o  c riti cal   inf or m at ion ’s.           Figure  1 .   Ge ne  f eat uresel ect io n or featu re cla ssici zat ion   basis       Re centl y  the  f ocus  of  re searc h  is  ve ry  act ive   as  w hen  a  key word  of  ‘ ge ne  sel ect ion ’  giv e n  int o  IE EE   Xp l or e  a  dig it al   li br ary  then   approxim at ely  51   jo urnals  was  f ound  onl y  fr om   20 16  t il l  3 rd   Febr ua r y  2018.  Tan g  et   al   in   their  m et ho d  of   feat ur e  se le ct ion   from   GEMA  ha ve  introd uced   a n  i m pr ov ise d  m utu a l   inf or m at ion   co rr el at ion   (MIC )  to  ha ndle   the  distor ti on  due   to  no ise   in  ge ne  an d  chall e nges  of  m ulti va riat e  Evaluation Warning : The document was created with Spire.PDF for Python.
In t J  Elec  &  C om p  En g     IS S N: 20 88 - 8708       SCD T :  FC - NN C - structu red  Complex  Decisi on Tec hn i qu e  for Ge ne   . ..  ( Su dha V. )   4507   distrib ution  est i m ation   by  a doptin g  releva nc e  bo os ti ng  a nd  e nhancem ent  of  the   feat ure  en ha ncem ent   [ 14 ] .   Table  1  li st t he  trends  of the   m et ho ds use d f or the  ge ne  sel ect ion .       Table  1 .   T re nd of the  Met hod A doptio n for  Gen e  Select io n   Sl.  No  and  r ef erences   Gen e Selection  /  C lass if icatio n  M eth o d     Dieses  Class if icati o n  &  Dataset  us ed   [ 1 5 ]  Zhan g  et  al.  ( 2 0 1 6 )   ●   m in i m u m   redu n d an cy  f eatu re  sele cti o n   m eth o d   ( m R MR)   ●   Multip le Ker n el  M achi n e ( MK L)  learn in g   m eth o d   ●   Glio b lasto m a m u lti f o r m e   ●   Can cer  Gen o m e  A tlas(TCGA )  d atab a se   [ 1 6 ]  Azzawi  et  al.  (20 1 6 )   ●     T wo  gen e sele ctio n   m eth o d s     ●   Gen e exp ressio n  pro g ra m m in g   (G EP ) - b ased   m o d el   ●   Lun g  ca n cer   ●   Real  m i croar ray lu n g  cancer  d atasets   [ 1 7 ]  Mallik et al.  ( 2 0 1 7 )   ●   m a x i m a l - re lev an ce  and   m in i m a l - r ed u n d an cy   ●   Epig en etic Bio m ar k er  d isco v ery   ●   Multi - O m ics P ros tate Carc in o m a  ( PC)  d ataset   [ 1 8 ]  Hu erta  et al .  ( 2 0 1 6 )   ●   Gen etic Algo rith m   ●   Tabu  Sear ch   ●   Su p p o rt  Vector   M achi n e   ●   Tu m o r  cl ass if icatio n   ●   Dif f u se Lar g e B - c ell L y m p h o m a   [ 1 9 ]  Sah a et  al.  (20 1 6 )   ●   Fu zzy  C - m e an s   ●   Hy p o th etical con d itio n  of  Yeast   ●   Yeast Sp o rulatio n ,  Yeast Cell  Cy cle,  Arabid o p sis ,  Hu m an  Fibro b last Scru m ,  Ra t  CNS   [ 2 0 ]  Mon tiel ( 2 0 1 6 )   ●   Si m u lated  ann eali n g   ●   Su p p o rt  v ecto r  m a ch in e   ●   Leuk e m ia datab as e   ●   Co lo n  Can cer  d atab ase   [ 2 1 ]  Ng u y en  ( 2 0 1 6 )   ●   Ty p e - 2  Fuz zy  log ic   ●   d if fus e lar g e B - cel l ly m p h o m a ,  l eu k em i a  cancer,  an d  pro state   [ 2 2 ]  Jin  and  W in  ( 2 0 1 6 )   ●   Swar m  intellig en c e   ●   Tu m o r  cl ass if icatio n   ●   Gen e M ic roarr ay  datas et   [ 2 3 ]Ray  et  al.  ( 2 0 1 6 )   ●   Self - Organ izin g  M ap   ●   Gen e M ic roarr ay  datas et   [ 2 4 ]  W an g  et  al.  ( 2 0 1 6 )   ●   Matr ix  f acto riza tio n   ●   Gen e M ic roarr ay  datas et   [ 2 5 ]  Han  et  al.  (20 1 7 )   ●   Particle  Swa r m   Op ti m izatio n   ●   SRB CT Data   [ 2 6 ]  Li  an d  W an g   (20 1 7 )   ●   K - m e an s alg o rithm     ●   ALL ,  GC M,   LY M ,  NC1 6 0 ,  M LL ,  H BC   [ 2 7 ]  Fen g  et  al.  (2 0 1 7 )   ●   Princip le Co m p o n en t Analysis   ●   PDDA - GE  Dataset   [ 2 8 ]  O m ar  et al.  ( 2 0 1 8 )   ●   Featu re  sele ctio n  prin cip le   ●   Gen e exp ressio n  datas et   [ 2 9 ]  Harikiran et a l.  (20 1 5 )   ●   seg m en tatio n  of   m icroarr a y  i m ag es   ●   Gen e M ic roarr ay  datas et   [ 3 0 ]  Ho re  et al.  ( 2 0 1 6 )   ●   I m ag e s eg m en tatio n   ●   Alp ert  d ataset       Ther e   are   va ri ou s   stu dies  be ing   ca rr ie d  out  in  e xisti ng  sy stem   towards  analy zi ng  m ic r oarray   data   us in g  di ff e ren t   form s  of   cl us te rin g  ap proac h.  Existi ng  m ec han ism   of   cl ust ering   a re  im mensely   it erati ve  in  it s   appr oach   wh ic h  evide ntly   cal ls  fo r  com pu t at ion al   com plexity .  Su ch  c om plexit y  issues  hav e  nev e r  bein g  addresse d  by  a ny  resea rc her s  ti ll   day.  On e   of  the  e ff e ct ive  m echan ism s  to  resist   s uch  co m plexit y  prob l e m   is  to  desig n  a nd  dev el op   a  novel  te chn i qu e  with  ve ry  lim i te d  set   of   it er at ion   unli ke  c onve ntion al   m achin e   le arn in g  a pproaches.  As  m ic ro ar ray data co ns ist s of h i gh e r  n um ber   of  in f or m at ion , th er e  is a n eed  of  a  syst e m   that  can  rea d  al l  the  exp li ci t  featur es  of   the  database  in  ord er  to  pe rfor m   a n  eff e ct ive  cl assifi cat ion .  Ado ptio n  of  fuzzy - ba sed   infe ren ce   syst e m   is  on e  s uc h  ap proac h  w he re  acc ur acy   in  classi ficat ion   a nd  com plexity   can b e   balance d.   But  existi ng   a pproaches  to wards  fu zzy   lo gic  al so   doesn ’t  seem   to  of fe r  m uch   convinci ng   ou t com es   towa rd s  cl assif ic at ion   posin g  as  on e  im ped im ent  towards  existi ng   re sear ch  w orks .  The   nex t  sect io n  outl ine s   the syst em   m o del of  pro pose d  s olu ti on.         2.   SY STE M   MO DEL: S C DT  & FC - NNC   The  pro posed  syst e m   m od el s   SCDT  &  FC - NN C  co ns ist   of  DS i {DLBC L (D S 1 ),   Le uk e m ia (D S 2 ),  Pr ost at e  Tum or   (DS 3 )},  w he r e  i =1, 2,3.  T he  ind ivi du al   data set   c har act erist ic s  are  sh own  in  the  Table  1a ,  1b,  and 1 c  of eac h DS 1, DS 2  an d D S 3 .   T he  s na ps hot vis ualiz at ion o f  eac h datase t i s shown i n  T able 2       Ta bl e  1( a) .   Desc ript ion   of  DLBC L(DS 1 )  Data set   Dataset na m e   Total Gen e   Total Sa m p l e   DLBCL   FL   DLBCL( DS 1 )   5470   77   58   19       Table   1(b) .   De scriptio n of Le uk em ia (D S 2)   Dataset na m e   Total Gen e   Total Sa m p l e   ALL   AML   Leuk e m ia( DS 2 )   5328   72   47   25   Evaluation Warning : The document was created with Spire.PDF for Python.
                          IS S N :   2088 - 8708   In t J  Elec  &  C om p  En g,   V ol.  8 , N o.   6 ,  Dece m ber  2 01 8   :   4505   -   4518   4508   Table1   (c ) .   De scriptio n of Pr os ta te  Tu m or   ( DS 3 )   Dataset na m e   Total Gen e   Total Sa m p l e   ALL   AML   Pros tate T u m o r  ( D S 3 )   1 0 5 1 0   102   52   50       Table  2 .   Sn a psho t  of eac h datase t DS 1, DS 2  a nd DS 3     1   2   3   4   1   59   1 7 4 8 0   3   384   2   267   1 2 0 8 6   52   - 325   3   66   8611   - 7   491   4   - 37   2 4 1 9 7   25   - 694   5   109   1 5 1 0 9   38   - 108   6   71   9059   - 23   - 220   7   31   2 9 4 8 0   31   - 5868   8   148   8305   - 21   - 96   9   84   1 0 3 2 1   2   - 4933   10   53   1 0 5 9 9   - 11   - 266   11   72   1 5 8 4 2   - 32   - 5193       1   2   3   4   1   88   1 5 0 9 1   7   311   2   283   1 1 0 3 8   37   134   3   309   1 6 6 9 2   183   378   4   12   1 5 7 6 3   45   268   5   168   1 8 1 2 8   - 28   118   6   71   3 4 2 0 7   65   154   7   55   3 0 8 0 1   43   80   8   - 2   2 5 1 4 7   338   269   9   268   1 5 2 7 2   29   188   10   219   2 1 8 0 1   - 36   - 39   11   82   1 8 1 6 7   - 8   115       DLBCL( DS 1 )     Leuk e m ia( DS 2 )       1   2   3   4   1   6 .10 0 0   - 0 .10 0 0   1 1 .90 0 0   1 4 .40 0 0   2   1   0   2   4   3   22   2   51   52   4   14   6   15   21   5   13   4   39   25   6   20   1   23   29   7   16   8   47   33   8   13   0   29   15   9   32   8   96   37   10   18   14   58   32     Pros tate T u m o u r ( DS 3 )       2.1 .   Gene  Sele ction  M e thod :  Con vent i onal  a n d Pr oposed  SCDT   Thr ee  c onve ntion al   m et ho ds   f or   the  ge ne  sec ti on s  inclu des  1)   T w o  sam ple   T  te st(2S TT ),   2)   E ntr opy   te st(ET)  an d  3)   Sig nal - to - N oise  Ra ti o( SNR )  is  evaluated  with  ra ndom   sa m ple   siz e   sect ion (S s )  for  ge ne   ran king  an d  visu al iz in g  to p - k  ge ne,   where  k = 3.   Al ong  with  pro pose d  Stru ct ur e d  Com plex  D eci sion  Tech nique (SC DT). T he  sect i on 3.1.1  d e scri bes 2ST T.     2.1.1 .   Tw o  s am ple T - te st ( 2ST T)   In   this  proce ss ,  two  in dep e ndent  sam ples  of   data  is  ta ke n  an d  in  order  to  know   the  wh et her   the   aver a ge  diff e re nce  am on g  t he se  two  sam ple s  are  sig nifica nt  or  not,  t he  2 - S am ple - T - te st  (2 S TT)  is  done .  I n  the  co nte xt  of  gen e   sel ect ion,   the  2S TT   is  pe rfor m ed  on  ea ch  gen e   a nd  th e  ex pr es sio n  le vels  are   se par a te d  on   the  basis  of  cl ass  var ia b le .  If   the  value  of  ‘ abs( t)’   is  f ound  m or e  that  ind ic at es  that  the  gen e  is   m or e   i m po rtant. nIf   the  t wo - data  si ze  of   n 1   a nd  n 2   with   th ei r  sa m ple  m ean  as  1   an d  2   as  well   1   and  2   be   t hei r   sam ple stand ar d dev ia ti on,  th en  the  v al ue of  t is com pu te d  by  Eq uatio n 1.                           (1)     2.1.2 En tropy   Te st   (ET )   The  cases  w he re  the  assumpti on   is  that  cl asses  are  nor m al l y  distribu te d  relat ive  ent ropy(RE )  or  Ku ll bac k - Lie bl er d ist ance or d ive rg e nce test  is con duct ed  usi ng   E qu at io n 2.   Th e g en e h a ving h ig hest v a lue  of  entr op y i s  sel ect ed  f or the i np ut of classi ficat ion  m odule.       (2)   Evaluation Warning : The document was created with Spire.PDF for Python.
In t J  Elec  &  C om p  En g     IS S N: 20 88 - 8708       SCD T :  FC - NN C - structu red  Complex  Decisi on Tec hn i qu e  for Ge ne   . ..  ( Su dha V. )   4509   2.1.3 .   Sign al  t o No ise  Rat i o (SNR )   SN R  def i nes  t he  r el at ive cla ss  separ at io n  m et ric b y m eans of  sig nal  qu al it y an d no ise     2.2 .   Gene  R aki ng   Algori th m s   The  pro posed   SCDT  ge ne  sel ect ion   al gorith m   ta kes  input  f ro m   the  th ree - diff e re nt  al gori thm   na m el y  2S TT ,  E T  a nd  SN R  for  great est   ranki ng  sel e ct ion   of  ge ne   f or  the  cl assi ficat ion   pur pose.  On  the   exec ution  of   above al gorith m , th e sn aps ho t of each  in div i du al  m et hod  is  ta ken  a nd s ho wn in t he  ta ble  3     SCDT : Ge ne R ank i ng A l gor it h m : GR - Algo rithm s   Creat e Em pty  vecto r  f or  2S T T, ET , SNR   for  eac h Gi   [v1, v 2] ←f (A ll , F L )   [m u1 m   m u2 ]←f m ean (v 1, v2 )   [sd1, sd 2]←  f st d (v1,v2)   [n1, n 2]←f len ( v1, v2)   2S TT  ← form ula   ET←  form ula   SRN←  f or m ula   Pr oc ess  for  P r opose d SCD T   Norm al iz a ti on   of 2 S TT,  ET,  S RN   [2 S TT,  ET,  SR N]← [ 2S T T /f m ax( 2STT) , E T /fm ax( ET) , SR N  /fm ax( SR N) ]   SCDT← f avg ( 2STT,  ET,  SRN )       Table  3  S na psho t  of Selec te d Ge ne by 2 STT , ET, S NR a nd  SCDT           To p Gen e  selec te d  f ro m  Two  Sam ple  T test (2STT) ,     To p Gen e  selec t ed  f ro m   Entr op y  Test (E T),     To p Gen e  selec te d  f ro m  Signal   to Noise  rati o  e st(SN RT ),         To p Gen e  selec te d  f ro m  p r op os e d  SC DT       2.3 .   Classi fica tion base d  on  th e S el ected  G ene   The  syst em   is   evaluated  f or   five  dif fer e nt   cl assifi cat ion s   in  wh ic h  fou r  nam el y  1)   Ra dial  Ba sis   functi on, 2)  M LP, 3) Fee d for ward,  4) SVM  and fi nally  FC - NN C i s  u se d      Evaluation Warning : The document was created with Spire.PDF for Python.
                          IS S N :   2088 - 8708   In t J  Elec  &  C om p  En g,   V ol.  8 , N o.   6 ,  Dece m ber  2 01 8   :   4505   -   4518   4510   2.3.1 .  R ad ial  Basis Fu ncti on   ( RBF)     The  ra dial  basi s  netw ork  (RB N)   a ppr ox im ates  the  functi on   by  add i ng   a dd it ion al   la ye r  to  the  hidde n  la ye r  of RB N u nless it  r eac hes  or ac hieves sp eci fied  ( or tar ge te d)  m ean squ are  obj ect ives .     f rbn (Input  vecto r,  Ta r get cla ss  value )  :  →RB N     2.3.2 .   Mul tila yer Perce pt i on ( MLP)   It  is  basical ly  a  ty pe  of   ne ur al   net work   t hat  us es  fee d  forw a r d - base d  le arn in g  m ec han ism   with   pr ese nce  of  th r ee  disti nct  la ye rs  of   nodes .  I t  is  al so   kn own  f or  it s  ad op ti on  of  s up e r vised  le ar ning  a ppr oac h  that i s term ed  as Bac k pro pa ga ti on  alg ori thm . MLP is  know n  f or it s ca pab i li ty  to  und ersta nd the  disti nction o f   li near   a nd  non - li near  data.   MLP  is  al s o  know n  for  it s  ut il iz ation   of  si gm oid   f un ct io n  t hat  is  em pirical ly  represe nted by      y(v i )=ta nh(v i ) a nd y( v i )=(1+e - vi ) - 1     The  prel im inary  com po ne nt  is  basical ly   a  hyperb olic  t ang e nt  with  a   ra nge  of  [ - 1  1]  wh il e  the   second   com po ne nt is  ba sic al ly  r epr es entat ion   of lo gi sti c functi on wi th a r a nge  of   [ 0 1].     2.3.3 .   Feed F or w ard   A  fee d  f orward  le arn i ng   pro cess  is  on e  of   t he  freq ue ntly   us e d  trai ning  a lgorit hm s  wh ic h  gove r n  the   or ie ntati on   of   the  inf o rm at ion   restrict ed  t o  a  sing le   direc ti on .  T he  oper at ion   of  fee d  forw a r d  ap pro ach  is  carried  ou t  bot h  in  sing le   an d  m ult iple - la ye r  per ce ptio ns   wh e re  both  of   the m   are  associat ed  with  pros  an d  cons.  The  pro s  factor   of   sin gle  and   m ultip le - la ye re d  pe rcep ti on  is   it s  si m plicity   an d  capa bili ty   to   so lv e   com plex  pro ble m s r especti ve ly . O n  the  oth e r  side, t he  co ns fact or   of   sin gle an d  m ulti ple - la ye red  p e rcep t ion  i s   it s co nsum ption   of h i gh e r  c om pu ta ti on al  ti m e and  i nclu de s incr ea sin g  it erati on s  r es pecti vely .     2.3.4 .   Sup po r t  V ec to r  Machi ne ( S V M)   It  is  al so  a  kind  of  m achine  le arn i ng  co nce pt   that  use s  s up erv ise d  le ar ning  a ppr oach e s  with  a   ta r get   of   a pp ly in g  th e m   fo r  perfor m ing   re gr essi on   or   pe rfor m ing   cl assifi cat ion   operati on.  SV M  is  capa ble  of   perform ing   bot h  li nea r  a nd  non - li near  cl ass ific at ion   qu it e  eff ect ively   ir re sp ect ive  of  it s  input  ty pe  of  hi gh e r  degree  of  di m ension al it y.  In   or der   to  a pp ly   this  al go rithm ,  it  is  require d  for  la beling  al l  the  data.  Im ple m entation   of   the  regres sion,  identific a ti on   of   ou tl ie rs ,  and   cl assifi ca ti on   is  carried  ou t  us i ng   hype r  plane   in  sup port  vec tor  m achine.  This  sc hem e  i s  al so   ca pab le   of  co ntr olli ng  the  com pu ta ti on al   l oad  that  al lows   si m pler  proces sing o f do t  pro du ct   us i ng a  va riable  us in g ke rn e  fun ct io n  k ( x,  y ).       2.3.5 F uz z y  Clust eri n g Neig hbo r hood  C lu ster  (FC - NNC )   The  pri m e   int ention  of   this  is  to  e m bed de d  the  bette r  de gr ee  of  f reedom   in  bo th  th e  infer e nce   (Mam dan i  and  Taka gi  Suge no)  m od el s  in   Fu zzy   lo gic  in  or der   to  e nsure  e nhance  c apab il it y  to  ad dr es s   un ce rtai nties.  This  ve rsion  of   fu zzy   lo gic  ha s  m or e  capa bi li t y  as  co m par ed  to  existi ng  on e  as  it   offers  m or e   pr act ic al it y  in  the  infe ren c e  proces s.  I n  c onven ti onal   f uzz y  log ic   base d  i m ple m entat ion ,  the  cris p  in puts  are   giv e n  to  fu zzi f ie r  w hich  is  f urt her   f orwarde d  as  f uzzy  set s   to  the  in fer e nc e  blo c k  that  is  con t ro ll ed  by  a   set   of  fu zzy   ru le s .  T he  fu zzy   outc om es  are  the n  f orwarde d  to   th e  de fu zzi fie r  i n  order  to  obta in  cris p  ou t pu t s.  T he   pro po se d  FC - NN C   pe rfor m s  the  sim il ar  ste p  ti ll   infe ren ce  b loc k  bu t   after  that  it   is  sig nifi cantl y  am e nd ed.  Th e  fu zzy   in puts  i n  FC - N NC  are   subj ect e d  to   a   sp eci al   f or m   of   outp ut  pr oc essing.  I n  t his  case,  a  ty pe  re du ce r  ob ta in s  the   in put  of  fu zzy   outpu t  set s,  proce sses  it   an d  t he n  forw a r ds   it   to  defuzzifi er   bl ock .   T her e   ar e  tw o  ou t pu ts  obtai ne d  in  FC - N NC  pro cess  i.e.  on e of cris p o utput an d  a nothe r i s ty pe - re duced  set.        3.   MICRO A RRAY D ATA  SE T   Ba sic al ly ,  m icr oa rr ay   ca n  be   sai d  to   be   a  c ollec ti on   of  di ff e ren t  num ber   of  s po ts  of  D NA,  wh e re   these  in form ation   is  util iz ed  for  c om pu ti ng  the  de gr ee   of   e xpres sio n  associat ed  with  gen e .  U su al ly ,  the  process  of   re pr ese ntati on   of  ge ne  ex pr es sion   data  is  car ried  ou t  usi ng   ex pressi on   m at rix,   wh er e  the   inf or m at ion   re ta ining   c olu m ns   re present  s ing le   ex pe rim ental   data  w hi le   al l  the  ro w s  ex hib it s  co m ple te  colle ct ion   of  e xp e rim ental   data.  Ba sic al ly ,  it  is  an  arch iv e  of  var i ous  f or m s  of   data  in  m icr oa rr ay   that  co ns i sts  of   esse ntial ly   the  inf or m at ion   of  ge ne  ex pressi on.  T he  pri m e   pu r pose  of   this  databa se  is  to  perfor m   an  eff ect ive   m anag em ent  of  the d at a w th  a n  ai d  of  the  d at a  in de x  assist in g  i n  gen e rati ng  a   query.   Di ff e ren t   f or m s  of   m ic ro - ar ray   data  are  A rr ay Track,  Im m Gen   data base,  ArrayEx pr es s,  G eneN et wor k,   MUSC,  U PSC - BASE ,  Stanf ord  Mi cr oarray   data bas e.  All  these  da ta bases  offer   volum ino us   i nfor m at ion   of  ge ne  e xpressio n  that  is   Evaluation Warning : The document was created with Spire.PDF for Python.
In t J  Elec  &  C om p  En g     IS S N: 20 88 - 8708       SCD T :  FC - NN C - structu red  Complex  Decisi on Tec hn i qu e  for Ge ne   . ..  ( Su dha V. )   4511   us e d  f or  public   util iz at ion .   Co ns ist ing  of  m or e  tha n  60,   00 0  sam ples  of  da ta   there   are   m or e   tha n  m il l ion s  of  prof il es i n gene  expres sio n.         4.   RESU LT S  AND A N ALYSIS   This  sect io n  di scusses  a bout   the  re su lt s  be ing   obta ined   f ro m   the  pro po sed  syst em .  The  c om plete   analy sis  of   t he   ou tc om e  is  carried  out  wi th  an  ai d  of   F1 - Score;   Lea ve  one  ou t  cr os s  valida ti on  m et ric,  Sens it ivit y,  Spec ific it y,  and   Pr eci sio n.   T he   sect ion   al so   el aborates  ab out  these  m e tho ds   i nd i viduall y  and   il lustrate s the  pro posed  an al ys is of res ults     4.1 .   Analysis  of F 1 - Score   In  bi nar y  cl ass ific at ion ,  t he  F 1 - sc ore  gen e ra ll y  m easur es  th e  acc uracy   of  the  te st  w hich  i s  com pu te d  on   t he  basis  of  pr e cessi on  a nd   recall   value s.  The  value  of  F1 - sc or e  is  c om pu te d  by  E qu at io n  3,   w he re  P=  pr eci sio n, S=  S ensiti vity . Th e   best  value of  F 1 - sc ore is c ons idere d  as  1  a nd  the  worst  on e   as 0.     F1 - Score = 2 × [ ( × ) ( + ) ]                 (3)     The  outc om e  sh ow n  in   Fig ure  2  hi gh li gh ts   that  F 1 - Sc ore   of  pr opos e d  SCDT  is  m uch   bette r  tha n  existi ng  ap pro aches  with  re sp ect   to   S NR,   entr opy,  a nd  t - te st.  Wh e re as,  pr opos e d  FC - NCC  offer  bette r   perform ance  in  com par iso n  to  existi ng  m a chin e  le ar ning   te chn i qu es  i. e.  RB F,  MLP,   Feed - f orward ,  a nd   su pp or t  vecto r   m achine,  et c.  A  cl os e r  lo ok  into  t he  pe rfo r m ance  will   on l y  sh ow  that  propose d  SC DT  offers   bette r  F1 - sc or e   in  com par iso n  to  FC - NCC.   T he  val ue  of  F 1 - Sc or e  f or  Pro po s ed  SC DT  is   found  at   1  f or   le as t   nu m ber   of  ge ne  from   the  gen e  ra ng e  of   5  - 30.  Eve n  at   the  increm ent al   values  of   t he   gen e,  th e  F1 - sco re   reduces  bu t  a ga in  bec om es  c on sist e nt  at   30  gen e  at   highe st  le vel  of   sc ore  1.  That  s ho ws  that  the  propose s   SCDT  ac hieve s  best  value  of  F1 - sco re.  A t  th e  sam e  t i m e  at   the  lo wer   ge ne   the  2STT,   S N R  and  the n  E ntropy  exh i bit  the  bette r  perform ance  in  reducin g  order.  Wh erea s  on   the  highe r  sel ect ion   of  gen e  S NR  get s  bette r   resu lt s  as  c ompare d  to  t he  2STT  an d  e ntr opy  ( bo t h  ex hibi t  sa m e  value ).   Figure  3   sho w s   perf or m ance  of   F 1 - scor e  vs c hang ing   value  of  nu m ber s o f  g e ne .           Figure  2 .   Per f orm ance o f F1 - s cor e  v s  ch a nging val ue of  nu m ber s o f  g e ne  ( a ) 2S TT,  ( b)   Entr op y t est ,   ( c)  SN RT ,   (d ) Prop os e d  SC D T       Evaluation Warning : The document was created with Spire.PDF for Python.
                          IS S N :   2088 - 8708   In t J  Elec  &  C om p  En g,   V ol.  8 , N o.   6 ,  Dece m ber  2 01 8   :   4505   -   4518   4512       Figure  3 Per for m ance of F1 - s cor e  v s  ch a nging va l ue of  nu m ber s o f  g e ne  ( a )  RB F ,  ( b) M LP,  ( c)  Fee d   forw a r d ,   ( 4)  S VM ,   ( 5) P r opose d  FC - NCC       4.2 .   Analysis  of Lea ve One  Out Cr os s  V alidat i on   Met ri c   (LOO C V)   Usu al ly ,  in  ge ne   exp re ssio n  da ta set   of   m ic ro arr ay   the  num ber   of   sam ples  a re  ver y  sm all,  t her e fore  to   pro vid e  ex ha ust ive  trai ning  le ave  one  out  cro ss  validat io n  m et ho d(LO O CV)  is  us e d.   I n  LO OCV  th e   entire  dataset   is  di vide d  int o  ‘ K ’  ra ndom   and  disti nc t  su bse t.  T he  K - 1  is  us e d  f or  trai ning  a nd  kt h  sam ple  is  us e d  f or   the  te sti ng   pur po s e.  T he  accu racy  of   LO OC V  is  com pu te d  by  Eq uatio n  4,   w her e  A  is  th e  count  of  c orr ect ly  cl assifi ed  sam ples .       =                   (4)     A  com par iso n  in  the  trend s  of   SC DT  an d  FC - NCC  sho w s  that  SCDT  offer s  inc reasi ng   value  of  LOO C V  in   co m par ison   to  F C - NCC  ove r  i ncr easi ng  num ber   of  ge nes .  This  is  a no t he r  cl ear  in dicat ing   t hat  pro po se d  SC D T  cou l d  offe r  bette r  de gr ee  of   in form at ion   wh il e  at tem pt ing   to  perform   cl assifi cat ion   of   th e   dis ease  or  any   oth e r  form   of  ab norm al i ty   i n  m ic ro arr ay   da ta .  Til l  now,  the  tre nd  of  SC DT  is  f ound  to   offe r   si m il ar  fo rm   of   co ns ist ency  f or   both  F1 - sc ore  an d  LO OC V .  Per form ance  of  LO OC V  vs   c hangin g  va lue  of  Nu m ber s  of  G ene  as  show n  in  Figure  4.   F igure  5  sho ws   per f orm ance  of   L OO C V  vs  changin g  val ue  of  nu m ber s  of  ge ne           Figure  4 .   Per f orm ance o f LO OCV vs  ch a ng ing   value  of  nu m ber s o f  g e ne  ( a ) 2S TT,  ( b)   Entr op y  Test ,  ( c)  SN RT ,   ( 4) Pro po s ed  SCD T   Evaluation Warning : The document was created with Spire.PDF for Python.
In t J  Elec  &  C om p  En g     IS S N: 20 88 - 8708       SCD T :  FC - NN C - structu red  Complex  Decisi on Tec hn i qu e  for Ge ne   . ..  ( Su dha V. )   4513       Figure  5 .   Per f orm ance o f LO OCV vs  ch a ng ing   value  of  nu m ber s o f  g e n e  ( a )  RB F ,  ( b) M LP,  ( c)   Fee d  f orwa r d,   ( 4)  S VM ,   ( 5) P r opose d  FC - NCC       4.3 .   Analysis  of Precessi on   Pr eci sio n  is  one  of   the  el e m e ntary  par am et e r  us e d  in  patte rn   rec ogniti on  as  well   as  in  classificat io n  pro blem s.  The c om pu ta ti on   of the  pr eci sio n i s car ried o ut   as  foll ow s   E qu at i on 5.       (5)     The  a bove   ex pr essi on  s how s  that  pr eci sio n  P   is  cal culat ed  by  di vid in g  the  diff e re nce   of  rele van t   inf or m at ion   RI  and   e xtracte d  inf or m at ion   EI  with  extr act ed  inform ation   EI .  This  expressi on   is  al ways  interp reted  w it h resp ect   t o prob a bili t y. The ou tc om es o btained  are   as   s ho wn in Fi gure  6  and Fig ure  7.           Figure  6 .  Per f orm ance o f p rec isi on   vs  c ha ng i ng v al ue of  nu m ber s o f  g e ne  ( a)  2STT,  ( b)   Entr op y  t est ,   ( c )  SN RT ,   ( 4) Pro po s ed  SCD T   Evaluation Warning : The document was created with Spire.PDF for Python.
                          IS S N :   2088 - 8708   In t J  Elec  &  C om p  En g,   V ol.  8 , N o.   6 ,  Dece m ber  2 01 8   :   4505   -   4518   4514       Figure  7 .   Per f orm ance o f p rec isi on   vs  c ha ng i ng v al ue of  nu m ber s o f  g e ne  ( a) RBF ,  ( b) M LP,  ( c)   FeedFo rw a rd,   ( 4) S VM , ( 5)   P rop os ed  FC - N CC       A  cl ose r  l ook  into  the   pa tt er n  of   t he  c urve  sh ow n  in   Fig ure  6  a nd  Fig ure  7  s how s  that   both  SC D T   and  FC - NCC  offe rs  sim il ar  lin ear   patte r n  of  preci sion  with   increasi ng  nu m ber   of   ge nes.  This  outc om e  sh ow s   that  irresp ect i ve  of  any  nu m ber   of   ge nes ,  the  pro posed   syst e m   us ing   any  form   of   fu zzy   lo gic  (not  the   conve ntion al  si ng le to n  one )  w il l a lway s yi el d  si m i la r  con sist ency in it s o utcom e, w hich  is qu it e predict ab le  in   it sel f.   The  pr e dicta bili ty   in  pr eci sio n  pe rfor m ance  offers  value  ad ded  per f or m ance  wh e n  at tem pti ng   t o  perform  classi ficat ion  of a ny  f or m  o f  cl inica l  abno rm aliti es i n  m ic r oar ray  da ta .      4.4 .   Analysis  of S en sitivit y   Sens it ivit y  is  ano t her   fr e qu e ntly   us e d  perform ance  par am et er  fo r  asse ssin g  cl assifi cat ion   perform ance.  I t  is  us ed  for  c al culat ing   am ou nt  of   posit ive   ou tc om e  con s idere d  to  be  a ccur at el y  ident ifie d.   The  cal c ulati on   of sen sit ivit y i s car ried o ut in fo ll owin g  m ann e r:       (6)     In  Eq uatio n 6 , Sensit ivit y i s co m pu te d by co ns ide rin g X which  is t r ue p osi ti ve  identific at ion   of so m e   cl inica l  abn or m al i ty   and   Y  wh ic h  is  false  neg at ive  i den ti ficat ion .  T he  grap hical   outc om e  of   s ensiti vi ty   is  a s   fo ll ows   Fig ur e   8  a nd Fig ure  9.           Figure  8 .  Per f orm ance o f Sen sit ivit y vs  cha ngin g value  of  Nu m ber s  of  G ene  ( a ) 2S TT,  ( b) En t ropy Tes t,  ( c)   SN RT ,   ( 4) Pro po s ed  SCD T   Evaluation Warning : The document was created with Spire.PDF for Python.