Int
ern
at
i
onal
Journ
al of Ele
ctrical
an
d
Co
mput
er
En
gin
eeri
ng
(IJ
E
C
E)
Vo
l.
8
, No
.
6
,
Decem
ber
201
8,
pp. 4
505~
4518
IS
S
N: 20
88
-
8708
,
DOI: 10
.11
591/
ijece
.
v8
i
6
.
pp
4505
-
45
18
4505
Journ
al h
om
e
page
:
http:
//
ia
es
core
.c
om/
journa
ls
/i
ndex.
ph
p/IJECE
SCDT: F
C
-
NNC
-
structur
ed
Com
plex Dec
ision T
echn
iqu
e fo
r
Gene An
alys
i
s Us
ing Fuz
zy Clust
er based
Nearest
Neighb
or
Classifie
r
Sudha
V
.
1
,
Gi
ri
ja
mm
a H
.
A
.
2
1
Depa
rtment of I
S&E,
RNS
Insti
t
ute
of
T
ec
hnolo
g
y
,
Indi
a
2
Depa
rtment
of
Com
pute
r
Scie
n
ce
& Engi
ne
eri
n
g,
RNS
Instit
u
te
of
Technol
og
y
,
I
ndia
Art
ic
le
In
f
o
ABSTR
A
CT
Art
ic
le
history:
Re
cei
ved
Feb
12
, 201
8
Re
vised
Jun
1
7
, 201
8
Accepte
d
J
un
2
0
, 201
8
In
m
an
y
d
isea
s
es
cl
assifi
cation
an
a
cc
ur
at
e
ge
ne
an
aly
s
is
is
nee
ded
,
for
which
select
ion
of
m
ost
informa
ti
ve
gen
es
is
ver
y
importan
t
and
it
req
uir
e
a
te
chn
ique
of
dec
ision
in
complex
cont
ext
of
ambiguity
.
Th
e
tra
dit
ion
al
m
et
hods
inc
lude
for
sele
ct
ing
m
ost
signifi
ca
nt
gene
inc
lud
es
som
e
of
the
stat
isti
ca
l
an
aly
s
is
namel
y
2
-
Sam
ple
-
T
-
te
st
(2STT
),
Ent
rop
y
,
Sign
al
to
Noise
Rat
io
(SN
R).
Thi
s
pape
r
eva
lu
a
te
s
gene
select
i
on
and
cl
assificat
ion
on
th
e
basis
of
ac
cur
a
te
gene
select
ion
using
struct
ure
d
complex
dec
isio
n
te
chni
qu
e
(SCD
T)
and
cl
a
ss
ifi
e
s
it
using
f
uzzy
c
luste
r
b
ase
d
nea
r
est
ne
igh
borc
la
ss
ifier
(FC
-
NN
C).
The
eff
e
ct
iv
ene
ss
of
the
propose
d
SC
DT
and
FC
-
NN
C
is
eva
lu
at
ed
for
l
e
ave
on
e
out
cr
oss
val
ida
t
ion
m
et
ric
(LOOCV
)
al
ong
wi
th
sensiti
vity
,
spec
i
fic
ity
,
pr
ec
ision
and
F1
-
score
wi
th
four
diff
er
ent
cl
assifi
ers
namel
y
1)
Radia
l
Basis
Functi
on
(RBF
),
2)
Multi
-
lay
er
p
ercept
io
n(MLP),
3)
Feed
Forw
ard
(F
F)
and
4)
Suppo
rt
vec
tor
m
ac
hin
e(SVM
)
for
thre
e
diffe
re
n
t
dat
ase
ts
of
DL
BCL,
L
euke
m
ia
and
Pros
ta
t
e
t
um
or.
The
prop
osed
SC
DT
&FC
-
NN
C
exhi
bit
s
superior
result
for
bei
ng
conside
red
m
ore
ac
cur
ate
dec
ision
m
ec
han
ism
.
Ke
yw
or
d:
Fu
zzy
classi
fic
at
ion
Gen
e
an
al
ysi
s
Gen
e
selec
ti
on
Ma
chine
le
a
rn
i
ng
Mi
cro
a
rr
ay
da
ta
Copyright
©
201
8
Instit
ut
e
o
f Ad
vanc
ed
Engi
n
ee
r
ing
and
S
cienc
e
.
Al
l
rights re
serv
ed
.
Corres
pond
in
g
Aut
h
or
:
Sudh
a
V
.
,
Dep
a
rtm
ent o
f Info
rm
at
ion
Sc
ie
nce &
En
gine
erin
g,
RNS
In
sti
tute
of Tech
nolo
gy,
Ben
galuru, I
ndia
.
Em
a
il
:
su
dh
a
vi
nayakam
@g
m
ai
l.com
1.
INTROD
U
CTION
The
accu
racy
of
dia
gnos
is
is
the
basis
fo
r
t
he
perfect
treat
m
ent
pr
ocess
t
o
be
ad
op
te
d
especial
ly
in
the
case
of
fa
ta
l
disease
li
ke
cancer
s,
le
ukem
ia
and
pro
strat
e
tum
or
et
c.
Alon
g
with
the
hist
op
at
holo
gy
,
m
edical
rad
iol
og
y
a
nd
im
aging
te
ch
niques,
the
m
ic
r
o
-
ar
r
ay
data
a
naly
sis
co
uld
be
pr
ov
e
n
quit
e
hel
pful
as
well
as
rig
htf
ul
if
eff
ic
ie
nt
te
chn
i
qu
e
s
of
an
al
ysi
s
are
evo
l
ved
[
1].
The
a
ccur
acy
of
dis
ease
cl
assifi
cat
ion
or
early
d
ia
gn
os
is
d
e
pends
up
on,
how acc
urat
el
y t
he
ge
ne o
f
si
gn
i
ficance is
s
el
ect
ed.
The
D
NA
-
m
i
cro
a
rr
ay
data
analy
sis
is
chall
eng
in
g
in
bo
th
as
pects
of
sta
ti
sti
cal
l
y
and
com
pu
ta
ti
on
al
ly
as
it
po
ssess
es
non
-
li
near
no
ise
s
al
ong
with
hi
gh
dim
ensio
nalit
y
of
low
sam
ple
data
[2
]
.
Ma
ny
ef
forts
towa
rd
s
diseas
e
diag
nosis
pa
rtic
ularly
canc
er,
t
um
or
et
c,
cl
assifi
cat
ion
ha
ve
bee
n
se
en
i
n
li
te
ratur
e
[
3
]
-
[
10
]
.
T
he
sect
io
n
2
descr
i
bes
t
he
insig
hts
of
r
el
at
ed
wor
k.
V
ario
us
m
achine
le
arn
in
g
ap
pro
ache
s
are
us
e
d
for
th
e
cl
assifi
cat
ion
wh
ic
h
incl
ud
e
s
rad
ia
l
ba
sis
f
un
ct
io
n
(RBF)
,
arti
fici
al
neur
al
networ
k
(
A
NN),
su
p
port
vecto
r
m
achine
(SV
M)
et
c.
by
f
orm
ing
the
pro
blem
as
bin
ary
cl
assifi
cat
ion
.
T
he
pr
ob
le
m
of
dim
ension
reducti
on
for
sear
chin
g
m
os
t
sig
nificant
ge
ne
is
bein
g
form
ulate
d
as
m
any
pr
oble
m
sp
aces
wh
i
c
h
include
s
1)
Mi
xed
inte
ger
pr
ogram
m
ing
(
MIP),
2)
Bi
o
-
i
ns
p
i
red
op
ti
m
i
zat
ion
(BI
O),
3)
Mi
ning
as
s
ociat
ion
ru
le
s
(MAR
),
a
nd last
but
no
t
the lea
st 4
)
E
nsem
ble tec
hn
i
que (ET
) [8
]
.
The
cl
inica
ll
y
com
pr
ehe
ns
ive
m
et
ho
d
requi
res
ha
nd
li
ng
hi
gh
dim
ensional
data
with
ver
aci
ty
and
no
ise
s
t
o
ha
nd
le
a
m
big
uity
duri
ng
t
he
rig
ht
gen
e
ca
ndidat
e
sel
ect
ion
.
T
hi
s
pap
e
r
pr
opose
s
a
m
echan
ism
of
Evaluation Warning : The document was created with Spire.PDF for Python.
IS
S
N
:
2088
-
8708
In
t J
Elec
&
C
om
p
En
g,
V
ol.
8
, N
o.
6
,
Dece
m
ber
2
01
8
:
4505
-
4518
4506
structu
re
d
c
omplex
decisi
on
t
echn
i
qu
e
(
SCDT)
f
or
f
uzzy
cl
us
te
rin
g
neig
hborh
ood
cl
us
t
er
(
FC
-
NC).
S
ect
ion
3
descr
i
bes
com
plete
syst
e
m
m
od
el
f
or
SC
DT
&
FC
-
NC,
Se
ct
ion
4
de
scrib
es
about
three
diff
e
r
e
nt
m
ic
ro
arr
ay
dataset
s.
Sect
i
on 5 il
lustrate
s
resu
lt
s a
nd a
na
ly
sis fo
ll
owed
by conclu
sio
n
i
n
Sect
io
n 6.
1.1
.
B
ackgr
ound
The
acc
ur
at
e
c
lusterin
g
of
th
e
data
is
a
cha
ll
eng
in
g
an
d
open
resear
ch
pro
blem
fo
r
cl
assifi
cat
ion
s
sp
eci
al
ly
us
ing
super
vise
le
arn
i
ng.
A
n
ex
te
ns
ive
sur
vey
is
con
duct
ed
to
under
sta
nd
the
eff
ect
ive
ness
of
cl
us
te
rin
g
te
ch
niques
pa
rtic
ul
arly
fo
r
m
edical
data
li
ke
m
ic
ro
ar
ray
ge
ne
dataset
[11].
Fo
r
t
he
pur
pose
of
tum
or
diag
no
si
s,
the
ap
proac
h
of
prof
il
in
g
th
e
gen
e
acc
urac
y
is
co
m
par
at
ively
of
higher
reli
abili
ty
with
m
or
e
accuracy
tha
n
that
of
t
he
m
e
thod
ad
opte
d
by
the
m
edica
l
i
m
aging
te
ch
nique
of
m
or
phol
og
ic
al
anal
ysi
s
of
tum
or
.
Tra
diti
on
al
ly
ad
op
te
d
s
up
e
rv
ise
d
l
earn
i
ng
a
ppr
oa
ches
fall
s
int
o
pitfal
l
of
ac
cur
acy
du
e
t
o
few
e
r
sam
ples
of
cancer
t
ypes
exist
into
the
trai
ning
dataset
of
ge
ne
ex
pr
essi
on
as
well
the
ov
e
rh
ea
ds
due
to
higher
data
-
dim
ension
al
it
y du
e t
o
la
rg
e
g
e
ne
e
xpre
ssion.
In
the
work
of
Lipowan
g
et
al
aim
s
to
sel
e
ct
few
num
ber
s
of
ge
nes
to
c
la
ssify
the
cancer
from
the
m
ic
ro
arr
ay
dat
a
to
m
eet
the
go
al
of
bala
nci
ng
tra
de
-
off
a
m
on
g
the
accu
racy
as
well
as
m
ini
m
iz
at
ion
of
the
com
pu
ta
ti
on
al
com
plexity
or
ov
e
r
head
s
[
3].
They
ha
ve
use
d
“feat
ur
e
im
portance
ranki
ng
s
chem
e”
for
the
accurate
or
sig
nificant
ge
ne
s
el
ect
ion
a
nd
f
orm
ulate
d
the
c
la
ssi
ficat
ion
prob
le
m
as
ty
pical
cl
us
te
r
of
bi
nar
y
cl
assifi
cat
ion
pro
blem
.
The
m
achine
le
arn
i
ng
a
ppr
oac
hes
us
e
d
in
thei
r
work
are
m
ix
us
e
of
fu
zzy
ne
ur
al
netw
ork
(FN
N
)
an
d
SV
M.
T
he
dim
ension
r
edu
ct
io
ns
ob
ta
i
ned
wer
e
getti
ng
sam
e
accuracy
on
ly
by
sel
ect
ing
28
ge
nes
as
co
m
par
ed
to
16,063
ge
nes
of
tra
diti
on
al
m
et
ho
d
of
t
hat
tim
e.
The
ty
pical
dat
aset
exp
l
or
e
d
f
or
t
he
ob
s
er
vations
i
nclu
des
1)
Ly
m
ph
om
a
Data,
2)
SRB
CT
D
at
a,
3)
Li
ver
C
ancer
Data,
an
d
4)
GCM
dat
a.
T
hey
reco
m
m
end
ed
consi
der
i
ng
th
e
cooper
at
io
n
aspects
betw
ee
n
the
ge
nes
to
m
ini
m
iz
e
the
gen
e
s
ubset
for
m
or
e
accurate
predic
ti
on
.
Fu
rt
her, the work which
ha
s r
efe
r
to this includ
es the wo
rk b
y
C
hien
-
Pa
ng et al
wh
o
ha
ve
introdu
c
e
d
a
m
et
ho
d
hybr
idize
d
us
i
ng
ge
netic
al
go
rith
m
and
dynam
i
cal
ly
setting
up
the
pa
ram
et
e
r
for
sig
nifican
t
gen
e
sel
ect
ion
an
d
then
furthe
r
use
s
SV
M
for
ve
rificat
ion
pur
po
s
es
to
predi
ct
gen
e
sel
ect
ion
e
ff
ic
ie
nc
y
[1
2].
T
he
dim
ension
re
duct
ion
an
d
feat
ur
e
sel
ect
ion
is
the
c
or
e
pr
oble
m
to
be
ha
nd
le
d
as
ge
ne
expressi
on
m
icr
oa
rr
ay
(G
EM
A)
co
ns
i
st
of
h
undred
to
s
om
eti
m
e
th
ou
s
an
ds
of
the
featu
res
in
a
ver
y
sm
al
l
sa
m
ple
siz
e.
These
hi
gh
nu
m
ber
s
of
fe
at
ur
es
i
n
a
sm
al
l
sa
m
ple
of
GEMA
m
akes
it
of
ver
y
high
dim
ension
da
ta
.
The
c
onve
ntion
a
l
m
et
ho
ds
a
dopt
ed
f
or
feat
ur
e s
el
ect
ion
w
hich
is
al
so
cal
le
d
a
s
ge
ne
sel
ect
io
n
in
c
ase of
t
he
GEM
A
analy
s
is
for
the cla
ssici
zat
ion o
f
t
he diese
s inclu
des 1
) G
ai
n
& Rel
ie
f
, 2
)
Chi
Squa
res,
3) Fishe
r Sco
re
, and 4
)
Lass
o et
c.
The
ge
neselect
ion
m
et
ho
d
is
cl
assifi
ed
into
three
cat
eg
or
ie
s
1)
s
up
e
rv
ise
d,
2)
uns
up
e
r
vi
sed
an
d
3)
sem
i
-
su
pe
rv
ise
d
on
the
basis
of
c
orrespo
nding
data
ty
pes
of
1)
f
ully
la
beled,
2)
unla
beled
a
nd
3)
pa
rtia
ll
y
la
beled
res
pec
ti
vely
for
cl
as
sific
at
ion
or
predict
io
n
of
cl
asses
as
des
cr
ibed
by
[
13
]
.
Fu
rt
her,
t
he
fe
at
ur
e
avail
able
into
the
GEMA
s
a
m
ples
are
cat
egoriz
ed
into
two
crit
ic
al
sel
ect
ion
s
na
m
el
y
red
un
da
ncy
an
d
releva
ncy.
The
F
ig
ure
1
,
s
hows
the
ty
pical
cl
assifi
cat
ion
based
on
the
com
bin
at
ion
of
these
tw
o
c
riti
cal
inf
or
m
at
ion
’s.
Figure
1
.
Ge
ne
f
eat
uresel
ect
io
n or featu
re cla
ssici
zat
ion
basis
Re
centl
y
the
f
ocus
of
re
searc
h
is
ve
ry
act
ive
as
w
hen
a
key
word
of
‘
ge
ne
sel
ect
ion
’
giv
e
n
int
o
IE
EE
Xp
l
or
e
a
dig
it
al
li
br
ary
then
approxim
at
ely
51
jo
urnals
was
f
ound
onl
y
fr
om
20
16
t
il
l
3
rd
Febr
ua
r
y
2018.
Tan
g
et
al
in
their
m
et
ho
d
of
feat
ur
e
se
le
ct
ion
from
GEMA
ha
ve
introd
uced
a
n
i
m
pr
ov
ise
d
m
utu
a
l
inf
or
m
at
ion
co
rr
el
at
ion
(MIC
)
to
ha
ndle
the
distor
ti
on
due
to
no
ise
in
ge
ne
an
d
chall
e
nges
of
m
ulti
va
riat
e
Evaluation Warning : The document was created with Spire.PDF for Python.
In
t J
Elec
&
C
om
p
En
g
IS
S
N: 20
88
-
8708
SCD
T
:
FC
-
NN
C
-
structu
red
Complex
Decisi
on Tec
hn
i
qu
e
for Ge
ne
.
..
(
Su
dha V.
)
4507
distrib
ution
est
i
m
ation
by
a
doptin
g
releva
nc
e
bo
os
ti
ng
a
nd
e
nhancem
ent
of
the
feat
ure
en
ha
ncem
ent
[
14
]
.
Table
1
li
st t
he
trends
of the
m
et
ho
ds use
d f
or the
ge
ne
sel
ect
ion
.
Table
1
.
T
re
nd of the
Met
hod A
doptio
n for
Gen
e
Select
io
n
Sl.
No
and
r
ef
erences
Gen
e Selection
/
C
lass
if
icatio
n
M
eth
o
d
Dieses
Class
if
icati
o
n
&
Dataset
us
ed
[
1
5
]
Zhan
g
et
al.
(
2
0
1
6
)
●
m
in
i
m
u
m
redu
n
d
an
cy
f
eatu
re
sele
cti
o
n
m
eth
o
d
(
m
R
MR)
●
Multip
le Ker
n
el
M
achi
n
e (
MK
L)
learn
in
g
m
eth
o
d
●
Glio
b
lasto
m
a
m
u
lti
f
o
r
m
e
●
Can
cer
Gen
o
m
e
A
tlas(TCGA
)
d
atab
a
se
[
1
6
]
Azzawi
et
al.
(20
1
6
)
●
T
wo
gen
e sele
ctio
n
m
eth
o
d
s
●
Gen
e exp
ressio
n
pro
g
ra
m
m
in
g
(G
EP
)
-
b
ased
m
o
d
el
●
Lun
g
ca
n
cer
●
Real
m
i
croar
ray lu
n
g
cancer
d
atasets
[
1
7
]
Mallik et al.
(
2
0
1
7
)
●
m
a
x
i
m
a
l
-
re
lev
an
ce
and
m
in
i
m
a
l
-
r
ed
u
n
d
an
cy
●
Epig
en
etic Bio
m
ar
k
er
d
isco
v
ery
●
Multi
-
O
m
ics P
ros
tate Carc
in
o
m
a
(
PC)
d
ataset
[
1
8
]
Hu
erta
et al
.
(
2
0
1
6
)
●
Gen
etic Algo
rith
m
●
Tabu
Sear
ch
●
Su
p
p
o
rt
Vector
M
achi
n
e
●
Tu
m
o
r
cl
ass
if
icatio
n
●
Dif
f
u
se Lar
g
e B
-
c
ell L
y
m
p
h
o
m
a
[
1
9
]
Sah
a et
al.
(20
1
6
)
●
Fu
zzy
C
-
m
e
an
s
●
Hy
p
o
th
etical con
d
itio
n
of
Yeast
●
Yeast Sp
o
rulatio
n
,
Yeast Cell
Cy
cle,
Arabid
o
p
sis
,
Hu
m
an
Fibro
b
last Scru
m
,
Ra
t
CNS
[
2
0
]
Mon
tiel (
2
0
1
6
)
●
Si
m
u
lated
ann
eali
n
g
●
Su
p
p
o
rt
v
ecto
r
m
a
ch
in
e
●
Leuk
e
m
ia datab
as
e
●
Co
lo
n
Can
cer
d
atab
ase
[
2
1
]
Ng
u
y
en
(
2
0
1
6
)
●
Ty
p
e
-
2
Fuz
zy
log
ic
●
d
if
fus
e lar
g
e B
-
cel
l ly
m
p
h
o
m
a
,
l
eu
k
em
i
a
cancer,
an
d
pro
state
[
2
2
]
Jin
and
W
in
(
2
0
1
6
)
●
Swar
m
intellig
en
c
e
●
Tu
m
o
r
cl
ass
if
icatio
n
●
Gen
e M
ic
roarr
ay
datas
et
[
2
3
]Ray
et
al.
(
2
0
1
6
)
●
Self
-
Organ
izin
g
M
ap
●
Gen
e M
ic
roarr
ay
datas
et
[
2
4
]
W
an
g
et
al.
(
2
0
1
6
)
●
Matr
ix
f
acto
riza
tio
n
●
Gen
e M
ic
roarr
ay
datas
et
[
2
5
]
Han
et
al.
(20
1
7
)
●
Particle
Swa
r
m
Op
ti
m
izatio
n
●
SRB
CT Data
[
2
6
]
Li
an
d
W
an
g
(20
1
7
)
●
K
-
m
e
an
s alg
o
rithm
●
ALL
,
GC
M,
LY
M
,
NC1
6
0
,
M
LL
,
H
BC
[
2
7
]
Fen
g
et
al.
(2
0
1
7
)
●
Princip
le Co
m
p
o
n
en
t Analysis
●
PDDA
-
GE
Dataset
[
2
8
]
O
m
ar
et al.
(
2
0
1
8
)
●
Featu
re
sele
ctio
n
prin
cip
le
●
Gen
e exp
ressio
n
datas
et
[
2
9
]
Harikiran et a
l.
(20
1
5
)
●
seg
m
en
tatio
n
of
m
icroarr
a
y
i
m
ag
es
●
Gen
e M
ic
roarr
ay
datas
et
[
3
0
]
Ho
re
et al.
(
2
0
1
6
)
●
I
m
ag
e s
eg
m
en
tatio
n
●
Alp
ert
d
ataset
Ther
e
are
va
ri
ou
s
stu
dies
be
ing
ca
rr
ie
d
out
in
e
xisti
ng
sy
stem
towards
analy
zi
ng
m
ic
r
oarray
data
us
in
g
di
ff
e
ren
t
form
s
of
cl
us
te
rin
g
ap
proac
h.
Existi
ng
m
ec
han
ism
of
cl
ust
ering
a
re
im
mensely
it
erati
ve
in
it
s
appr
oach
wh
ic
h
evide
ntly
cal
ls
fo
r
com
pu
t
at
ion
al
com
plexity
.
Su
ch
c
om
plexit
y
issues
hav
e
nev
e
r
bein
g
addresse
d
by
a
ny
resea
rc
her
s
ti
ll
day.
On
e
of
the
e
ff
e
ct
ive
m
echan
ism
s
to
resist
s
uch
co
m
plexit
y
prob
l
e
m
is
to
desig
n
a
nd
dev
el
op
a
novel
te
chn
i
qu
e
with
ve
ry
lim
i
te
d
set
of
it
er
at
ion
unli
ke
c
onve
ntion
al
m
achin
e
le
arn
in
g
a
pproaches.
As
m
ic
ro
ar
ray data co
ns
ist
s of h
i
gh
e
r
n
um
ber
of
in
f
or
m
at
ion
, th
er
e
is a n
eed
of
a
syst
e
m
that
can
rea
d
al
l
the
exp
li
ci
t
featur
es
of
the
database
in
ord
er
to
pe
rfor
m
a
n
eff
e
ct
ive
cl
assifi
cat
ion
.
Ado
ptio
n
of
fuzzy
-
ba
sed
infe
ren
ce
syst
e
m
is
on
e
s
uc
h
ap
proac
h
w
he
re
acc
ur
acy
in
classi
ficat
ion
a
nd
com
plexity
can b
e
balance
d.
But
existi
ng
a
pproaches
to
wards
fu
zzy
lo
gic
al
so
doesn
’t
seem
to
of
fe
r
m
uch
convinci
ng
ou
t
com
es
towa
rd
s
cl
assif
ic
at
ion
posin
g
as
on
e
im
ped
im
ent
towards
existi
ng
re
sear
ch
w
orks
.
The
nex
t
sect
io
n
outl
ine
s
the syst
em
m
o
del of
pro
pose
d
s
olu
ti
on.
2.
SY
STE
M
MO
DEL: S
C
DT
& FC
-
NNC
The
pro
posed
syst
e
m
m
od
el
s
SCDT
&
FC
-
NN
C
co
ns
ist
of
DS
i
{DLBC
L
(D
S
1
),
Le
uk
e
m
ia
(D
S
2
),
Pr
ost
at
e
Tum
or
(DS
3
)},
w
he
r
e
i
=1,
2,3.
T
he
ind
ivi
du
al
data
set
c
har
act
erist
ic
s
are
sh
own
in
the
Table
1a
,
1b,
and 1
c
of eac
h DS
1,
DS
2
an
d D
S
3
.
T
he
s
na
ps
hot vis
ualiz
at
ion o
f
eac
h datase
t i
s shown i
n
T
able 2
Ta
bl
e
1(
a)
.
Desc
ript
ion
of
DLBC
L(DS
1
)
Data
set
Dataset na
m
e
Total Gen
e
Total Sa
m
p
l
e
DLBCL
FL
DLBCL(
DS
1
)
5470
77
58
19
Table
1(b)
.
De
scriptio
n of Le
uk
em
ia
(D
S
2)
Dataset na
m
e
Total Gen
e
Total Sa
m
p
l
e
ALL
AML
Leuk
e
m
ia(
DS
2
)
5328
72
47
25
Evaluation Warning : The document was created with Spire.PDF for Python.
IS
S
N
:
2088
-
8708
In
t J
Elec
&
C
om
p
En
g,
V
ol.
8
, N
o.
6
,
Dece
m
ber
2
01
8
:
4505
-
4518
4508
Table1
(c
)
.
De
scriptio
n of Pr
os
ta
te
Tu
m
or
(
DS
3
)
Dataset na
m
e
Total Gen
e
Total Sa
m
p
l
e
ALL
AML
Pros
tate T
u
m
o
r
(
D
S
3
)
1
0
5
1
0
102
52
50
Table
2
.
Sn
a
psho
t
of eac
h datase
t DS
1,
DS
2
a
nd DS
3
1
2
3
4
1
59
1
7
4
8
0
3
384
2
267
1
2
0
8
6
52
-
325
3
66
8611
-
7
491
4
-
37
2
4
1
9
7
25
-
694
5
109
1
5
1
0
9
38
-
108
6
71
9059
-
23
-
220
7
31
2
9
4
8
0
31
-
5868
8
148
8305
-
21
-
96
9
84
1
0
3
2
1
2
-
4933
10
53
1
0
5
9
9
-
11
-
266
11
72
1
5
8
4
2
-
32
-
5193
1
2
3
4
1
88
1
5
0
9
1
7
311
2
283
1
1
0
3
8
37
134
3
309
1
6
6
9
2
183
378
4
12
1
5
7
6
3
45
268
5
168
1
8
1
2
8
-
28
118
6
71
3
4
2
0
7
65
154
7
55
3
0
8
0
1
43
80
8
-
2
2
5
1
4
7
338
269
9
268
1
5
2
7
2
29
188
10
219
2
1
8
0
1
-
36
-
39
11
82
1
8
1
6
7
-
8
115
DLBCL(
DS
1
)
Leuk
e
m
ia(
DS
2
)
1
2
3
4
1
6
.10
0
0
-
0
.10
0
0
1
1
.90
0
0
1
4
.40
0
0
2
1
0
2
4
3
22
2
51
52
4
14
6
15
21
5
13
4
39
25
6
20
1
23
29
7
16
8
47
33
8
13
0
29
15
9
32
8
96
37
10
18
14
58
32
Pros
tate T
u
m
o
u
r
(
DS
3
)
2.1
.
Gene
Sele
ction
M
e
thod
:
Con
vent
i
onal
a
n
d Pr
oposed
SCDT
Thr
ee
c
onve
ntion
al
m
et
ho
ds
f
or
the
ge
ne
sec
ti
on
s
inclu
des
1)
T
w
o
sam
ple
T
te
st(2S
TT
),
2)
E
ntr
opy
te
st(ET)
an
d
3)
Sig
nal
-
to
-
N
oise
Ra
ti
o(
SNR
)
is
evaluated
with
ra
ndom
sa
m
ple
siz
e
sect
ion
(S
s
)
for
ge
ne
ran
king
an
d
visu
al
iz
in
g
to
p
-
k
ge
ne,
where
k
=
3.
Al
ong
with
pro
pose
d
Stru
ct
ur
e
d
Com
plex
D
eci
sion
Tech
nique (SC
DT). T
he
sect
i
on 3.1.1
d
e
scri
bes 2ST
T.
2.1.1
.
Tw
o
s
am
ple T
-
te
st (
2ST
T)
In
this
proce
ss
,
two
in
dep
e
ndent
sam
ples
of
data
is
ta
ke
n
an
d
in
order
to
know
the
wh
et
her
the
aver
a
ge
diff
e
re
nce
am
on
g
t
he
se
two
sam
ple
s
are
sig
nifica
nt
or
not,
t
he
2
-
S
am
ple
-
T
-
te
st
(2
S
TT)
is
done
.
I
n
the
co
nte
xt
of
gen
e
sel
ect
ion,
the
2S
TT
is
pe
rfor
m
ed
on
ea
ch
gen
e
a
nd
th
e
ex
pr
es
sio
n
le
vels
are
se
par
a
te
d
on
the
basis
of
cl
ass
var
ia
b
le
.
If
the
value
of
‘
abs(
t)’
is
f
ound
m
or
e
that
ind
ic
at
es
that
the
gen
e
is
m
or
e
i
m
po
rtant.
nIf
the
t
wo
-
data
si
ze
of
n
1
a
nd
n
2
with
th
ei
r
sa
m
ple
m
ean
as
1
an
d
2
as
well
1
and
2
be
t
hei
r
sam
ple stand
ar
d dev
ia
ti
on,
th
en
the
v
al
ue of
t is com
pu
te
d
by
Eq
uatio
n 1.
(1)
2.1.2 En
tropy
Te
st
(ET
)
The
cases
w
he
re
the
assumpti
on
is
that
cl
asses
are
nor
m
al
l
y
distribu
te
d
relat
ive
ent
ropy(RE
)
or
Ku
ll
bac
k
-
Lie
bl
er d
ist
ance or d
ive
rg
e
nce test
is con
duct
ed
usi
ng
E
qu
at
io
n 2.
Th
e g
en
e h
a
ving h
ig
hest v
a
lue
of
entr
op
y i
s
sel
ect
ed
f
or the i
np
ut of classi
ficat
ion
m
odule.
(2)
Evaluation Warning : The document was created with Spire.PDF for Python.
In
t J
Elec
&
C
om
p
En
g
IS
S
N: 20
88
-
8708
SCD
T
:
FC
-
NN
C
-
structu
red
Complex
Decisi
on Tec
hn
i
qu
e
for Ge
ne
.
..
(
Su
dha V.
)
4509
2.1.3
.
Sign
al
t
o No
ise
Rat
i
o (SNR
)
SN
R
def
i
nes
t
he
r
el
at
ive cla
ss
separ
at
io
n
m
et
ric b
y m
eans of
sig
nal
qu
al
it
y an
d no
ise
2.2
.
Gene
R
aki
ng
Algori
th
m
s
The
pro
posed
SCDT
ge
ne
sel
ect
ion
al
gorith
m
ta
kes
input
f
ro
m
the
th
ree
-
diff
e
re
nt
al
gori
thm
na
m
el
y
2S
TT
,
E
T
a
nd
SN
R
for
great
est
ranki
ng
sel
e
ct
ion
of
ge
ne
f
or
the
cl
assi
ficat
ion
pur
pose.
On
the
exec
ution
of
above al
gorith
m
, th
e sn
aps
ho
t of each
in
div
i
du
al
m
et
hod
is
ta
ken
a
nd s
ho
wn in t
he
ta
ble
3
SCDT
: Ge
ne R
ank
i
ng A
l
gor
it
h
m
: GR
-
Algo
rithm
s
Creat
e Em
pty
vecto
r
f
or
2S
T
T, ET
, SNR
for
eac
h Gi
[v1, v
2]
←f
(A
ll
, F
L
)
[m
u1
m
m
u2
]←f
m
ean
(v
1, v2
)
[sd1, sd
2]←
f
st
d
(v1,v2)
[n1, n
2]←f
len
(
v1,
v2)
2S
TT
←
form
ula
ET←
form
ula
SRN←
f
or
m
ula
Pr
oc
ess
for
P
r
opose
d SCD
T
Norm
al
iz
a
ti
on
of 2
S
TT,
ET,
S
RN
[2
S
TT,
ET,
SR
N]← [
2S
T
T /f
m
ax(
2STT)
, E
T /fm
ax(
ET)
, SR
N
/fm
ax(
SR
N)
]
SCDT←
f
avg
(
2STT,
ET,
SRN
)
Table
3
S
na
psho
t
of Selec
te
d Ge
ne by 2
STT
, ET, S
NR a
nd
SCDT
To
p Gen
e
selec
te
d
f
ro
m
Two
Sam
ple
T test
(2STT)
,
To
p Gen
e
selec
t
ed
f
ro
m
Entr
op
y
Test
(E
T),
To
p Gen
e
selec
te
d
f
ro
m
Signal
to Noise
rati
o
e
st(SN
RT
),
To
p Gen
e
selec
te
d
f
ro
m
p
r
op
os
e
d
SC
DT
2.3
.
Classi
fica
tion base
d
on
th
e S
el
ected
G
ene
The
syst
em
is
evaluated
f
or
five
dif
fer
e
nt
cl
assifi
cat
ion
s
in
wh
ic
h
fou
r
nam
el
y
1)
Ra
dial
Ba
sis
functi
on, 2)
M
LP, 3) Fee
d for
ward,
4) SVM
and fi
nally
FC
-
NN
C i
s
u
se
d
Evaluation Warning : The document was created with Spire.PDF for Python.
IS
S
N
:
2088
-
8708
In
t J
Elec
&
C
om
p
En
g,
V
ol.
8
, N
o.
6
,
Dece
m
ber
2
01
8
:
4505
-
4518
4510
2.3.1
.
R
ad
ial
Basis Fu
ncti
on
(
RBF)
The
ra
dial
basi
s
netw
ork
(RB
N)
a
ppr
ox
im
ates
the
functi
on
by
add
i
ng
a
dd
it
ion
al
la
ye
r
to
the
hidde
n
la
ye
r
of RB
N u
nless it
r
eac
hes
or ac
hieves sp
eci
fied
(
or tar
ge
te
d)
m
ean squ
are
obj
ect
ives
.
f
rbn
(Input
vecto
r,
Ta
r
get cla
ss
value
)
:
→RB
N
2.3.2
.
Mul
tila
yer Perce
pt
i
on (
MLP)
It
is
basical
ly
a
ty
pe
of
ne
ur
al
net
work
t
hat
us
es
fee
d
forw
a
r
d
-
base
d
le
arn
in
g
m
ec
han
ism
with
pr
ese
nce
of
th
r
ee
disti
nct
la
ye
rs
of
nodes
.
I
t
is
al
so
kn
own
f
or
it
s
ad
op
ti
on
of
s
up
e
r
vised
le
ar
ning
a
ppr
oac
h
that i
s term
ed
as Bac
k pro
pa
ga
ti
on
alg
ori
thm
. MLP is
know
n
f
or it
s ca
pab
i
li
ty
to
und
ersta
nd the
disti
nction o
f
li
near
a
nd
non
-
li
near
data.
MLP
is
al
s
o
know
n
for
it
s
ut
il
iz
ation
of
si
gm
oid
f
un
ct
io
n
t
hat
is
em
pirical
ly
represe
nted by
y(v
i
)=ta
nh(v
i
) a
nd y(
v
i
)=(1+e
-
vi
)
-
1
The
prel
im
inary
com
po
ne
nt
is
basical
ly
a
hyperb
olic
t
ang
e
nt
with
a
ra
nge
of
[
-
1
1]
wh
il
e
the
second
com
po
ne
nt is
ba
sic
al
ly
r
epr
es
entat
ion
of lo
gi
sti
c functi
on wi
th a r
a
nge
of
[
0 1].
2.3.3
.
Feed F
or
w
ard
A
fee
d
f
orward
le
arn
i
ng
pro
cess
is
on
e
of
t
he
freq
ue
ntly
us
e
d
trai
ning
a
lgorit
hm
s
wh
ic
h
gove
r
n
the
or
ie
ntati
on
of
the
inf
o
rm
at
ion
restrict
ed
t
o
a
sing
le
direc
ti
on
.
T
he
oper
at
ion
of
fee
d
forw
a
r
d
ap
pro
ach
is
carried
ou
t
bot
h
in
sing
le
an
d
m
ult
iple
-
la
ye
r
per
ce
ptio
ns
wh
e
re
both
of
the
m
are
associat
ed
with
pros
an
d
cons.
The
pro
s
factor
of
sin
gle
and
m
ultip
le
-
la
ye
re
d
pe
rcep
ti
on
is
it
s
si
m
plicity
an
d
capa
bili
ty
to
so
lv
e
com
plex
pro
ble
m
s r
especti
ve
ly
. O
n
the
oth
e
r
side, t
he
co
ns fact
or
of
sin
gle an
d
m
ulti
ple
-
la
ye
red
p
e
rcep
t
ion
i
s
it
s co
nsum
ption
of h
i
gh
e
r
c
om
pu
ta
ti
on
al
ti
m
e and
i
nclu
de
s incr
ea
sin
g
it
erati
on
s
r
es
pecti
vely
.
2.3.4
.
Sup
po
r
t
V
ec
to
r
Machi
ne (
S
V
M)
It
is
al
so
a
kind
of
m
achine
le
arn
i
ng
co
nce
pt
that
use
s
s
up
erv
ise
d
le
ar
ning
a
ppr
oach
e
s
with
a
ta
r
get
of
a
pp
ly
in
g
th
e
m
fo
r
perfor
m
ing
re
gr
essi
on
or
pe
rfor
m
ing
cl
assifi
cat
ion
operati
on.
SV
M
is
capa
ble
of
perform
ing
bot
h
li
nea
r
a
nd
non
-
li
near
cl
ass
ific
at
ion
qu
it
e
eff
ect
ively
ir
re
sp
ect
ive
of
it
s
input
ty
pe
of
hi
gh
e
r
degree
of
di
m
ension
al
it
y.
In
or
der
to
a
pp
ly
this
al
go
rithm
,
it
is
require
d
for
la
beling
al
l
the
data.
Im
ple
m
entation
of
the
regres
sion,
identific
a
ti
on
of
ou
tl
ie
rs
,
and
cl
assifi
ca
ti
on
is
carried
ou
t
us
i
ng
hype
r
plane
in
sup
port
vec
tor
m
achine.
This
sc
hem
e
i
s
al
so
ca
pab
le
of
co
ntr
olli
ng
the
com
pu
ta
ti
on
al
l
oad
that
al
lows
si
m
pler
proces
sing o
f do
t
pro
du
ct
us
i
ng a
va
riable
us
in
g ke
rn
e
fun
ct
io
n
k
(
x,
y
).
2.3.5 F
uz
z
y
Clust
eri
n
g Neig
hbo
r
hood
C
lu
ster
(FC
-
NNC
)
The
pri
m
e
int
ention
of
this
is
to
e
m
bed
de
d
the
bette
r
de
gr
ee
of
f
reedom
in
bo
th
th
e
infer
e
nce
(Mam
dan
i
and
Taka
gi
Suge
no)
m
od
el
s
in
Fu
zzy
lo
gic
in
or
der
to
e
nsure
e
nhance
c
apab
il
it
y
to
ad
dr
es
s
un
ce
rtai
nties.
This
ve
rsion
of
fu
zzy
lo
gic
ha
s
m
or
e
capa
bi
li
t
y
as
co
m
par
ed
to
existi
ng
on
e
as
it
offers
m
or
e
pr
act
ic
al
it
y
in
the
infe
ren
c
e
proces
s.
I
n
c
onven
ti
onal
f
uzz
y
log
ic
base
d
i
m
ple
m
entat
ion
,
the
cris
p
in
puts
are
giv
e
n
to
fu
zzi
f
ie
r
w
hich
is
f
urt
her
f
orwarde
d
as
f
uzzy
set
s
to
the
in
fer
e
nc
e
blo
c
k
that
is
con
t
ro
ll
ed
by
a
set
of
fu
zzy
ru
le
s
.
T
he
fu
zzy
outc
om
es
are
the
n
f
orwarde
d
to
th
e
de
fu
zzi
fie
r
i
n
order
to
obta
in
cris
p
ou
t
pu
t
s.
T
he
pro
po
se
d
FC
-
NN
C
pe
rfor
m
s
the
sim
il
ar
ste
p
ti
ll
infe
ren
ce
b
loc
k
bu
t
after
that
it
is
sig
nifi
cantl
y
am
e
nd
ed.
Th
e
fu
zzy
in
puts
i
n
FC
-
N
NC
are
subj
ect
e
d
to
a
sp
eci
al
f
or
m
of
outp
ut
pr
oc
essing.
I
n
t
his
case,
a
ty
pe
re
du
ce
r
ob
ta
in
s
the
in
put
of
fu
zzy
outpu
t
set
s,
proce
sses
it
an
d
t
he
n
forw
a
r
ds
it
to
defuzzifi
er
bl
ock
.
T
her
e
ar
e
tw
o
ou
t
pu
ts
obtai
ne
d
in
FC
-
N
NC
pro
cess
i.e.
on
e of cris
p o
utput an
d
a
nothe
r i
s ty
pe
-
re
duced
set.
3.
MICRO
A
RRAY D
ATA
SE
T
Ba
sic
al
ly
,
m
icr
oa
rr
ay
ca
n
be
sai
d
to
be
a
c
ollec
ti
on
of
di
ff
e
ren
t
num
ber
of
s
po
ts
of
D
NA,
wh
e
re
these
in
form
ation
is
util
iz
ed
for
c
om
pu
ti
ng
the
de
gr
ee
of
e
xpres
sio
n
associat
ed
with
gen
e
.
U
su
al
ly
,
the
process
of
re
pr
ese
ntati
on
of
ge
ne
ex
pr
es
sion
data
is
car
ried
ou
t
usi
ng
ex
pressi
on
m
at
rix,
wh
er
e
the
inf
or
m
at
ion
re
ta
ining
c
olu
m
ns
re
present
s
ing
le
ex
pe
rim
ental
data
w
hi
le
al
l
the
ro
w
s
ex
hib
it
s
co
m
ple
te
colle
ct
ion
of
e
xp
e
rim
ental
data.
Ba
sic
al
ly
,
it
is
an
arch
iv
e
of
var
i
ous
f
or
m
s
of
data
in
m
icr
oa
rr
ay
that
co
ns
i
sts
of
esse
ntial
ly
the
inf
or
m
at
ion
of
ge
ne
ex
pressi
on.
T
he
pri
m
e
pu
r
pose
of
this
databa
se
is
to
perfor
m
an
eff
ect
ive
m
anag
em
ent
of
the d
at
a w
th
a
n
ai
d
of
the
d
at
a
in
de
x
assist
in
g
i
n
gen
e
rati
ng
a
query.
Di
ff
e
ren
t
f
or
m
s
of
m
ic
ro
-
ar
ray
data
are
A
rr
ay
Track,
Im
m
Gen
data
base,
ArrayEx
pr
es
s,
G
eneN
et
wor
k,
MUSC,
U
PSC
-
BASE
,
Stanf
ord
Mi
cr
oarray
data
bas
e.
All
these
da
ta
bases
offer
volum
ino
us
i
nfor
m
at
ion
of
ge
ne
e
xpressio
n
that
is
Evaluation Warning : The document was created with Spire.PDF for Python.
In
t J
Elec
&
C
om
p
En
g
IS
S
N: 20
88
-
8708
SCD
T
:
FC
-
NN
C
-
structu
red
Complex
Decisi
on Tec
hn
i
qu
e
for Ge
ne
.
..
(
Su
dha V.
)
4511
us
e
d
f
or
public
util
iz
at
ion
.
Co
ns
ist
ing
of
m
or
e
tha
n
60,
00
0
sam
ples
of
da
ta
there
are
m
or
e
tha
n
m
il
l
ion
s
of
prof
il
es i
n gene
expres
sio
n.
4.
RESU
LT
S
AND A
N
ALYSIS
This
sect
io
n
di
scusses
a
bout
the
re
su
lt
s
be
ing
obta
ined
f
ro
m
the
pro
po
sed
syst
em
.
The
c
om
plete
analy
sis
of
t
he
ou
tc
om
e
is
carried
out
wi
th
an
ai
d
of
F1
-
Score;
Lea
ve
one
ou
t
cr
os
s
valida
ti
on
m
et
ric,
Sens
it
ivit
y,
Spec
ific
it
y,
and
Pr
eci
sio
n.
T
he
sect
ion
al
so
el
aborates
ab
out
these
m
e
tho
ds
i
nd
i
viduall
y
and
il
lustrate
s the
pro
posed
an
al
ys
is of res
ults
4.1
.
Analysis
of F
1
-
Score
In
bi
nar
y
cl
ass
ific
at
ion
,
t
he
F
1
-
sc
ore
gen
e
ra
ll
y
m
easur
es
th
e
acc
uracy
of
the
te
st
w
hich
i
s
com
pu
te
d
on
t
he
basis
of
pr
e
cessi
on
a
nd
recall
value
s.
The
value
of
F1
-
sc
or
e
is
c
om
pu
te
d
by
E
qu
at
io
n
3,
w
he
re
P=
pr
eci
sio
n, S=
S
ensiti
vity
. Th
e
best
value of
F
1
-
sc
ore is c
ons
idere
d
as
1
a
nd
the
worst
on
e
as 0.
F1
-
Score
=
2
×
[
(
×
)
(
+
)
]
(3)
The
outc
om
e
sh
ow
n
in
Fig
ure
2
hi
gh
li
gh
ts
that
F
1
-
Sc
ore
of
pr
opos
e
d
SCDT
is
m
uch
bette
r
tha
n
existi
ng
ap
pro
aches
with
re
sp
ect
to
S
NR,
entr
opy,
a
nd
t
-
te
st.
Wh
e
re
as,
pr
opos
e
d
FC
-
NCC
offer
bette
r
perform
ance
in
com
par
iso
n
to
existi
ng
m
a
chin
e
le
ar
ning
te
chn
i
qu
es
i.
e.
RB
F,
MLP,
Feed
-
f
orward
,
a
nd
su
pp
or
t
vecto
r
m
achine,
et
c.
A
cl
os
e
r
lo
ok
into
t
he
pe
rfo
r
m
ance
will
on
l
y
sh
ow
that
propose
d
SC
DT
offers
bette
r
F1
-
sc
or
e
in
com
par
iso
n
to
FC
-
NCC.
T
he
val
ue
of
F
1
-
Sc
or
e
f
or
Pro
po
s
ed
SC
DT
is
found
at
1
f
or
le
as
t
nu
m
ber
of
ge
ne
from
the
gen
e
ra
ng
e
of
5
-
30.
Eve
n
at
the
increm
ent
al
values
of
t
he
gen
e,
th
e
F1
-
sco
re
reduces
bu
t
a
ga
in
bec
om
es
c
on
sist
e
nt
at
30
gen
e
at
highe
st
le
vel
of
sc
ore
1.
That
s
ho
ws
that
the
propose
s
SCDT
ac
hieve
s
best
value
of
F1
-
sco
re.
A
t
th
e
sam
e
t
i
m
e
at
the
lo
wer
ge
ne
the
2STT,
S
N
R
and
the
n
E
ntropy
exh
i
bit
the
bette
r
perform
ance
in
reducin
g
order.
Wh
erea
s
on
the
highe
r
sel
ect
ion
of
gen
e
S
NR
get
s
bette
r
resu
lt
s
as
c
ompare
d
to
t
he
2STT
an
d
e
ntr
opy
(
bo
t
h
ex
hibi
t
sa
m
e
value
).
Figure
3
sho
w
s
perf
or
m
ance
of
F
1
-
scor
e
vs c
hang
ing
value
of
nu
m
ber
s o
f
g
e
ne
.
Figure
2
.
Per
f
orm
ance o
f F1
-
s
cor
e
v
s
ch
a
nging val
ue of
nu
m
ber
s o
f
g
e
ne
(
a
) 2S
TT,
(
b)
Entr
op
y t
est
,
(
c)
SN
RT
,
(d
) Prop
os
e
d
SC
D
T
Evaluation Warning : The document was created with Spire.PDF for Python.
IS
S
N
:
2088
-
8708
In
t J
Elec
&
C
om
p
En
g,
V
ol.
8
, N
o.
6
,
Dece
m
ber
2
01
8
:
4505
-
4518
4512
Figure
3 Per
for
m
ance of F1
-
s
cor
e
v
s
ch
a
nging va
l
ue of
nu
m
ber
s o
f
g
e
ne
(
a
)
RB
F
,
(
b) M
LP,
(
c)
Fee
d
forw
a
r
d
,
(
4)
S
VM
,
(
5) P
r
opose
d
FC
-
NCC
4.2
.
Analysis
of Lea
ve One
Out Cr
os
s
V
alidat
i
on
Met
ri
c
(LOO
C
V)
Usu
al
ly
,
in
ge
ne
exp
re
ssio
n
da
ta
set
of
m
ic
ro
arr
ay
the
num
ber
of
sam
ples
a
re
ver
y
sm
all,
t
her
e
fore
to
pro
vid
e
ex
ha
ust
ive
trai
ning
le
ave
one
out
cro
ss
validat
io
n
m
et
ho
d(LO
O
CV)
is
us
e
d.
I
n
LO
OCV
th
e
entire
dataset
is
di
vide
d
int
o
‘
K
’
ra
ndom
and
disti
nc
t
su
bse
t.
T
he
K
-
1
is
us
e
d
f
or
trai
ning
a
nd
kt
h
sam
ple
is
us
e
d
f
or
the
te
sti
ng
pur
po
s
e.
T
he
accu
racy
of
LO
OC
V
is
com
pu
te
d
by
Eq
uatio
n
4,
w
her
e
A
is
th
e
count
of
c
orr
ect
ly
cl
assifi
ed
sam
ples
.
=
(4)
A
com
par
iso
n
in
the
trend
s
of
SC
DT
an
d
FC
-
NCC
sho
w
s
that
SCDT
offer
s
inc
reasi
ng
value
of
LOO
C
V
in
co
m
par
ison
to
F
C
-
NCC
ove
r
i
ncr
easi
ng
num
ber
of
ge
nes
.
This
is
a
no
t
he
r
cl
ear
in
dicat
ing
t
hat
pro
po
se
d
SC
D
T
cou
l
d
offe
r
bette
r
de
gr
ee
of
in
form
at
ion
wh
il
e
at
tem
pt
ing
to
perform
cl
assifi
cat
ion
of
th
e
dis
ease
or
any
oth
e
r
form
of
ab
norm
al
i
ty
i
n
m
ic
ro
arr
ay
da
ta
.
Til
l
now,
the
tre
nd
of
SC
DT
is
f
ound
to
offe
r
si
m
il
ar
fo
rm
of
co
ns
ist
ency
f
or
both
F1
-
sc
ore
an
d
LO
OC
V
.
Per
form
ance
of
LO
OC
V
vs
c
hangin
g
va
lue
of
Nu
m
ber
s
of
G
ene
as
show
n
in
Figure
4.
F
igure
5
sho
ws
per
f
orm
ance
of
L
OO
C
V
vs
changin
g
val
ue
of
nu
m
ber
s
of
ge
ne
Figure
4
.
Per
f
orm
ance o
f LO
OCV vs
ch
a
ng
ing
value
of
nu
m
ber
s o
f
g
e
ne
(
a
) 2S
TT,
(
b)
Entr
op
y
Test
,
(
c)
SN
RT
,
(
4) Pro
po
s
ed
SCD
T
Evaluation Warning : The document was created with Spire.PDF for Python.
In
t J
Elec
&
C
om
p
En
g
IS
S
N: 20
88
-
8708
SCD
T
:
FC
-
NN
C
-
structu
red
Complex
Decisi
on Tec
hn
i
qu
e
for Ge
ne
.
..
(
Su
dha V.
)
4513
Figure
5
.
Per
f
orm
ance o
f LO
OCV vs
ch
a
ng
ing
value
of
nu
m
ber
s o
f
g
e
n
e
(
a
)
RB
F
,
(
b) M
LP,
(
c)
Fee
d
f
orwa
r
d,
(
4)
S
VM
,
(
5) P
r
opose
d
FC
-
NCC
4.3
.
Analysis
of Precessi
on
Pr
eci
sio
n
is
one
of
the
el
e
m
e
ntary
par
am
et
e
r
us
e
d
in
patte
rn
rec
ogniti
on
as
well
as
in
classificat
io
n
pro
blem
s.
The c
om
pu
ta
ti
on
of the
pr
eci
sio
n i
s car
ried o
ut
as
foll
ow
s
E
qu
at
i
on 5.
(5)
The
a
bove
ex
pr
essi
on
s
how
s
that
pr
eci
sio
n
P
is
cal
culat
ed
by
di
vid
in
g
the
diff
e
re
nce
of
rele
van
t
inf
or
m
at
ion
RI
and
e
xtracte
d
inf
or
m
at
ion
EI
with
extr
act
ed
inform
ation
EI
.
This
expressi
on
is
al
ways
interp
reted
w
it
h resp
ect
t
o prob
a
bili
t
y. The ou
tc
om
es o
btained
are
as
s
ho
wn in Fi
gure
6
and Fig
ure
7.
Figure
6
.
Per
f
orm
ance o
f p
rec
isi
on
vs
c
ha
ng
i
ng v
al
ue of
nu
m
ber
s o
f
g
e
ne
(
a)
2STT,
(
b)
Entr
op
y
t
est
,
(
c
)
SN
RT
,
(
4) Pro
po
s
ed
SCD
T
Evaluation Warning : The document was created with Spire.PDF for Python.
IS
S
N
:
2088
-
8708
In
t J
Elec
&
C
om
p
En
g,
V
ol.
8
, N
o.
6
,
Dece
m
ber
2
01
8
:
4505
-
4518
4514
Figure
7
.
Per
f
orm
ance o
f p
rec
isi
on
vs
c
ha
ng
i
ng v
al
ue of
nu
m
ber
s o
f
g
e
ne
(
a) RBF
,
(
b) M
LP,
(
c)
FeedFo
rw
a
rd,
(
4) S
VM
, (
5)
P
rop
os
ed
FC
-
N
CC
A
cl
ose
r
l
ook
into
the
pa
tt
er
n
of
t
he
c
urve
sh
ow
n
in
Fig
ure
6
a
nd
Fig
ure
7
s
how
s
that
both
SC
D
T
and
FC
-
NCC
offe
rs
sim
il
ar
lin
ear
patte
r
n
of
preci
sion
with
increasi
ng
nu
m
ber
of
ge
nes.
This
outc
om
e
sh
ow
s
that
irresp
ect
i
ve
of
any
nu
m
ber
of
ge
nes
,
the
pro
posed
syst
e
m
us
ing
any
form
of
fu
zzy
lo
gic
(not
the
conve
ntion
al
si
ng
le
to
n
one
)
w
il
l a
lway
s yi
el
d
si
m
i
la
r
con
sist
ency in it
s o
utcom
e, w
hich
is qu
it
e predict
ab
le
in
it
sel
f.
The
pr
e
dicta
bili
ty
in
pr
eci
sio
n
pe
rfor
m
ance
offers
value
ad
ded
per
f
or
m
ance
wh
e
n
at
tem
pti
ng
t
o
perform
classi
ficat
ion
of a
ny
f
or
m
o
f
cl
inica
l
abno
rm
aliti
es i
n
m
ic
r
oar
ray
da
ta
.
4.4
.
Analysis
of S
en
sitivit
y
Sens
it
ivit
y
is
ano
t
her
fr
e
qu
e
ntly
us
e
d
perform
ance
par
am
et
er
fo
r
asse
ssin
g
cl
assifi
cat
ion
perform
ance.
I
t
is
us
ed
for
c
al
culat
ing
am
ou
nt
of
posit
ive
ou
tc
om
e
con
s
idere
d
to
be
a
ccur
at
el
y
ident
ifie
d.
The
cal
c
ulati
on
of sen
sit
ivit
y i
s car
ried o
ut in fo
ll
owin
g
m
ann
e
r:
(6)
In
Eq
uatio
n 6
, Sensit
ivit
y i
s co
m
pu
te
d by co
ns
ide
rin
g X which
is t
r
ue p
osi
ti
ve
identific
at
ion
of so
m
e
cl
inica
l
abn
or
m
al
i
ty
and
Y
wh
ic
h
is
false
neg
at
ive
i
den
ti
ficat
ion
.
T
he
grap
hical
outc
om
e
of
s
ensiti
vi
ty
is
a
s
fo
ll
ows
Fig
ur
e
8
a
nd Fig
ure
9.
Figure
8
.
Per
f
orm
ance o
f Sen
sit
ivit
y vs
cha
ngin
g value
of
Nu
m
ber
s
of
G
ene
(
a
) 2S
TT,
(
b) En
t
ropy Tes
t,
(
c)
SN
RT
,
(
4) Pro
po
s
ed
SCD
T
Evaluation Warning : The document was created with Spire.PDF for Python.