Indonesian
J
our
nal
of
Electrical
Engineering
and
Computer
Science
V
ol.
42,
No.
3,
June
2026,
pp.
827
∼
834
ISSN:
2502-4752,
DOI:
10.11591/ijeecs.v42.i3.pp827-834
❒
827
On
exploring
text
mining
appr
oaches
to
sentiment
analysis
based
on
the
combination
of
w
ord-based
and
ontology-based
appr
oaches
Suthira
Plansangk
et,
Supapor
n
Kansomk
eat,
Supasit
Kajkamhaeng
Di
vision
of
Computational
Science,
F
aculty
of
Science,
Prince
of
Songkla
Uni
v
ersity
,
Songkhla,
Thailand
Article
Inf
o
Article
history:
Recei
v
ed
Jul
28,
2025
Re
vised
Mar
21,
2026
Accepted
May
26,
2026
K
eyw
ords:
CSDF
ontoCSDF
Sentiment
analysis
T
e
xt
mining
TF-IDF
ABSTRA
CT
Currently
,
sentiment
analysis
plays
an
important
role
in
b
usiness.
Entrepreneurs
try
to
understand
customer
needs
for
products
and
services.
If
the
y
kno
w
about
the
needs,
the
y
can
create
the
mark
eting
plans
or
strate
gy
plans
in
their
b
usiness
that
help
impro
v
e
products
and
services.
Therefore,
this
study
e
xplores
tw
o
no
v
el
approaches
to
impro
v
e
the
classication
accurac
y
of
sentiment
analysis
data
using
a
combination
of
a
w
ord-based
approach
(TF-IDF
or
CSDF)
and
an
ontology-based
approach
(ontoSen)
to
pro
vide
tw
o
ne
w
methods
,
called
ontoTF-
IDF
and
ontoCSDF
.
The
e
xperimental
r
esults
sho
w
that
CSDF
method
had
the
best
classication
accurac
y
among
all
the
methods
in
this
study:
ontoCSDF
did
not
impro
v
e
further
the
classi
cation
accurac
y
of
sentiment
analysis
data.
Furthermore,
ontoTFIDF
method
impro
v
ed
the
classication
by
IBk
algorithm
signicantly
(p
<
0.05).
This
is
an
open
access
article
under
the
CC
BY
-SA
license
.
Corresponding
A
uthor:
Suthira
Plansangk
et
Di
vision
of
Computational
Science,
F
aculty
of
Science,
Prince
of
Songkla
Uni
v
ersity
Songkhla,
Thailand
Email:
suthira.p@psu.ac.th
1.
INTR
ODUCTION
Sentiment
analysis
plays
an
important
role
in
b
usiness.
T
o
understand
customer
needs
for
products
and
services
is
concerned
by
the
entrepreneurs.
F
or
e
xample,
ho
w
people
think
about
a
particular
mo
vie
or
a
hotel,
etc.
If
a
b
usiness
is
a
w
are
of
the
needs,
it
can
create
mark
eting
or
st
rate
gy
plans
to
impro
v
e
products
and
services.
In
addition,
if
we
can
impro
v
e
sentiment
analysis
for
feedback
of
feelings
by
the
users,
then
the
products
and
services
can
be
impro
v
ed,
gi
ving
a
competiti
v
e
adv
antage
to
the
b
usiness.
This
study
aims
to
e
xplore
tw
o
no
v
el
methods
to
impro
v
e
the
classication
accurac
y
of
sentim
ent
analysis
data,
using
combinations
of
w
ord-based
and
ontology-based
approaches.
Normally
,
w
ord-based
ap-
proaches
use
a
bag-of-w
ords
representation
with
statistics.
This
ignores
both
order
and
meaning
of
the
w
ords.
On
the
other
hand,
an
ontology-based
approach
focuses
on
the
order
,
semantics,
and
relationships
of
w
ords.
Therefore,
this
study
aimed
to
e
xplore
the
combination
these
tw
o
approaches
for
an
impro
v
ed
sentiment
anal-
ysis
by
te
xt
mining.
2.
LITERA
TURE
REVIEW
Af
fecti
v
e
computing
is
study
for
de
v
elopi
n
g
systems
and
de
vices
that
can
recognize,
process,
inte
rpret,
and
simulate
human
af
fects.
It
is
an
interdisciplinary
eld
combining
computer
science,
psychology
,
and
J
ournal
homepage:
http://ijeecs.iaescor
e
.com
Evaluation Warning : The document was created with Spire.PDF for Python.
828
❒
ISSN:
2502-4752
cogniti
v
e
science
[1].
One
of
the
moti
v
ations
for
such
research
is
the
ability
to
gi
v
e
machines
emotional
intelligence,
including
simulated
empath
y
.
The
machine
should
interpret
the
emotional
states
of
humans
and
adapt
its
beha
viour
to
them,
gi
ving
appropriate
responses
to
those
emotions.
T
o
detect
the
emotions,
we
start
from
the
sensors
that
collect
data
on
the
status
or
ph
ysical
beha
viour
of
humans,
without
interpretation.
The
data
from
the
sens
ors
are
c
omparable
to
the
data
from
human
s
enses
reecting
other
human’
s
emotions.
F
or
e
xample,
a
camera
can
detec
t
f
acial
e
xpressions,
posture,
and
ph
ysical
mo
v
ements;
and
a
microphone
can
detect
speech.
Furthermore,
there
are
other
sensors
that
can
directly
collect
ph
ysiological
data,
such
as
body
temperature,
brai
n
w
a
v
es,
and
electrical
conducti
vity
of
the
skin.
In
addition,
emotional
recognition
needs
to
e
xtract
meaningful
patterns
from
the
collected
signals,
which
can
be
learned
by
v
arious
machine
learning
algorithms,
such
as
those
for
speech
recognition,
f
ace
detection,
and
natural
language
process
ing
(NLP).
After
that,
these
can
be
interpreted
as
emotions.
Sentiment
analysis
is
a
subset
of
af
fecti
v
e
computing.
It
is
about
NLP
that
learns
from
the
sentences
in
documents,
and
then
translates
to
e
xpress
feelings,
thinking,
or
attitudes.
Sentiment
analysis
is
separated
into
tw
o
major
types:
positi
v
e
and
ne
g
ati
v
e
[2].
Some
research
also
considers
a
type
called
neutral
[3].
Sentiment
analysis
consists
of
tw
o
major
techniques.
First,
machine
learning
or
data
mining
is
used
to
nd
patterns
and
relationships
inside
a
lar
ge
dataset
for
sentiment
analysis,
by
using
mathematical,
statistical,
and
patt
ern
recog-
nition
approaches
[3]-[5].
Second,
the
types
of
w
ords
in
a
document
f
all
into
tw
o
cate
gories:
positi
v
e
w
ords
and
ne
g
ati
v
e
w
ords.
After
that,
each
w
ord
can
be
gi
v
en
a
score
and
s
u
m
med
to
a
total
score.
If
the
positi
v
e
w
ords
dominate
o
v
er
the
ne
g
ati
v
e
w
ords,
then
the
sum
is
positi
v
e
and
we
can
conclude
that
this
document
is
positi
v
e
[6].
Natural
language
processing
plays
an
important
role
in
sentiment
analysis.
Benamara
et
al.
[7]
pro-
posed
using
a
combination
of
adjecti
v
es
and
adv
erbs
in
sentiment
analysis.
The
results
sho
w
that
this
combina-
tion
helps
impro
v
e
the
performance
of
sentiment
analysis
using
pearson
correlation.
Furthermore,
K
ouloumpis
et
al.
[5]
studied
the
parts
of
speech;
nouns,
v
erbs,
adjecti
v
es,
and
adv
erbs
in
sentiment
analysis
by
data
mining.
The
data
representation
in
this
research
used
a
bag
of
w
ords
[8],
which
accounts
for
each
w
ord
separately
with-
out
considering
their
order
.
The
method
of
data
representation
by
term
frequenc
y–in
v
erse
document
frequenc
y
(TF-IDF)
[9]
is
well-kno
w
and
a
standard
method.
Ho
we
v
er
,
Plansangk
et
and
Gan
[10]
ha
v
e
proposed
a
no
v
el
data
representation
called
the
class
specic
document
frequenc
y
(CSDF).
It
is
a
data
representation
method
which
focuses
on
the
important
w
ords
in
te
xt
mining.
These
are
the
w
ords
that
often
appear
in
documents
of
the
same
class,
b
ut
rarely
appear
in
documents
of
a
dif
ferent
class.
The
e
xperimental
results
sho
w
that
CSDF
gi
v
es
a
better
classication
accurac
y
than
the
TF-IDF
.
TF-IDF
and
CSDF
are
a
bag
of
w
ords
model.
Each
w
ord
i
s
independently
accounted
without
reference
to
w
ord
order
.
Ho
we
v
er
,
the
order
of
w
ords
is
important
in
meaningful
sentences.
A
dif
ferent
order
of
w
ords
gi
v
es
a
semantic
dif
ference.
F
or
e
xample,
“
the
rabbit
ran
f
aster
than
the
turtle
”
and
“
the
turtle
ran
f
aster
than
the
rabbit
”
are
totally
dif
ferent
semantically
,
although
made
up
of
the
s
ame
w
ords.
Therefore,
some
studies
ha
v
e
used
ontology
in
te
xt
mining
for
information
retrie
v
al
[11]
and
sentiment
analysis
[12].
Ontology
is
a
study
of
object
characteristics,
re
g
arding
what
is
this
object,
ho
w
man
y
types
of
this
object
are
there,
and
structure,
properties,
e
v
ents,
processes
of
this
object,
and
relations
between
objects
in
the
real
w
orld.
In
addition,
ontology
is
the
reliable
reality
in
the
real
w
orld
[13].
Furt
hermore,
ontology
is
the
principal
foundation
of
kno
wledge
representation
that
collects
all
kno
wledge,
se
mantics,
and
relationships
from
dif
ferent
sources
[14].
Ontology
separates
things
into
dif
ferent
types;
therefore,
it
relates
to
te
xt
mining
and
sentiment
analysis.
No
w
adays,
e
xtensi
v
e
research
focuses
on
inte
grating
ontology
with
machine
learning
and
deep
learn-
ing
techniques.
F
or
e
xample,
Sharma
and
K
umar
[15]
proposed
a
ne
w
h
ybrid
semantic
inde
xing
approach
for
unstructured
te
xt
documents
w
as
introduced
by
inte
grating
machine
learning
with
domain
ontology
.
The
method
enhanced
its
capability
to
identify
concepts
t
hat
are
semantically
related
to
document
content
through
the
use
of
a
machine
learning–based
skip-gram
model.
Experimental
results
demonstrated
that
the
proposed
method
outperformed
state-of-the-art
techniques
on
the
e
v
aluated
datasets,
achie
ving
an
a
v
erage
accurac
y
im-
pro
v
ement
of
29%.
Moreo
v
er
,
Jain
et
al.
[16]
presented
a
ne
w
ontology-based
natural
language
processing
techniques
that
combine
feature
e
xtraction
with
deep
learning–based
classication.
A
comparati
v
e
analysis
w
as
conducted
between
the
proposed
approach
and
e
xisting
techniques,
with
the
proposed
method
achie
ving
superior
results.
Indonesian
J
Elec
Eng
&
Comp
Sci,
V
ol.
42,
No.
3,
June
2026:
827–834
Evaluation Warning : The document was created with Spire.PDF for Python.
Indonesian
J
Elec
Eng
&
Comp
Sci
ISSN:
2502-4752
❒
829
Although
the
use
of
machine
learning
and
deep
learning
can
yield
good
results,
their
comple
xity
and
the
signicant
amount
of
time
required
for
learning
led
the
researchers
to
choose
a
simpler
and
more
con
v
enient
approach
for
learning.
This
study
aims
to
e
xplore
tw
o
no
v
el
approaches
to
impro
v
e
the
classication
accurac
y
of
sentiment
analysis
data
using
com
bined
w
ord-based
approac
h
,
namel
y
TF-IDF
or
CSDF
,
and
ontology-based
approach,
namely
ontoSen.
The
tw
o
ne
w
methods
called
ontoTF-IDF
and
ontoCSDF
may
help
to
impro
v
e
the
performance
of
sentiment
analysis.
3.
METHOD
Methodologies
in
this
research
which
aims
to
impro
v
e
the
performance
of
sentiment
analysis
were
separated
into
three
major
approaches.
3.1.
Data
r
epr
esentation
Data
representation
usi
ng
alternati
v
ely
TF-IDF
and
CSDF
method
is
used
in
both
training
and
tes
ting
set.
First,
the
TF-IDF
scores
[17],
[18]
are
calculated
as
(1):
T
F
I
D
F
j
i
=
(
(1
+
l
og
2
T
F
j
i
)
×
l
og
2
N
D
F
i
,
if
T
F
j
i
>
0
0
,
otherwise
(1)
where
T
F
I
D
F
j
i
is
a
score
for
term
i
in
document
j,
T
F
j
i
is
the
frequenc
y
for
term
i
in
document
j,
N
is
the
total
documents
in
the
training
set,
and
D
F
i
is
the
number
of
documents
that
term
i
appears
in,
in
the
training
set.
Second,
CSDF
score
[10]
of
term
i
in
class
k
is
dened
as
(2):
C
S
D
F
ik
=
(
D
F
ik
N
k
/
(
D
F
i
−
D
F
ik
)
(
N
−
N
k
)+1
,
if
T
F
ik
>
0
0
,
otherwise
(2)
where
D
F
ik
is
the
count
of
term
i
appearing
in
the
training
set
and
in
class
k,
D
F
i
is
the
total
number
of
documents
that
term
i
appears
in
in
the
training
set,
N
k
is
the
total
number
of
documents
in
the
training
set
and
in
class
k,
and
N
is
the
total
nu
m
ber
of
documents
in
the
training
set.
Ho
we
v
er
,
we
do
not
kno
w
the
v
alues
of
D
F
ik
and
N
k
in
documents
of
the
test
set.
Therefore,
CSDF
v
alues
of
term
i
in
both
training
and
testing
documents
are
dened
by
the
v
ariance
v
alue
of
original
CSDF
for
term
i
in
class
k,
as
sho
wn
in
(3).
C
S
D
F
i
=
v
ar
(
C
S
D
F
ik
)
(3)
3.2.
The
ontology
The
data
representation
using
ontology
in
this
researc
h
w
as
applied
as
in
the
research
of
W
´
ojcik
and
T
ucho
wski
[19].
The
ontology
or
kno
wledge
base
is
the
collection
of
all
f
acts
about
the
objects
in
the
real
w
orld,
including
relationships
between
them.
The
structure
of
this
collection
is
a
hierarchical
structure
that
is
rank
ed
one
abo
v
e
the
other
by
importance.
F
or
e
xample,
the
root
class
is
thing,
with
phone
as
its
direct
subclass.
This
class
serv
es
as
the
parent
for
cate
gories
that
represent
the
main
features
of
smartphones.
Subsequent
le
v
els
in
the
hierarch
y
describe
less
signicant
or
more
detailed
phone
char
acteristics.
Each
class
is
assigned
a
single
le
v
el
based
on
its
position
in
the
hierarch
y
.
The
phone
class
is
designated
as
the
rst-le
v
el
class.
Its
immediate
subclasses
form
the
second
le
v
el,
while
their
descendants
correspond
to
the
follo
wing
hierarchical
le
v
els.
Therefore,
the
total
sentiment
is
dened
as
follo
ws:
O
ntoS
e
n
i
=
P
n
j
=1
h
O
ntoS
en
i
(
j
)
l
ev
el
i
(
j
)
i
n
(4)
where
O
ntoS
en
i
is
the
total
ontology
score
of
term
i
from
each
document
j
,
n
is
the
total
number
of
document
s,
O
ntoS
e
n
i
(
j
)
is
an
ontology
score
of
term
i
of
document
j
,
and
l
ev
el
i
(
j
)
is
the
relationship
le
v
el
of
term
i
of
document
j
.
On
e
xploring
te
xt
mining
appr
oac
hes
to
sentiment
analysis
based
on
the
...
(Suthir
a
Plansangk
et)
Evaluation Warning : The document was created with Spire.PDF for Python.
830
❒
ISSN:
2502-4752
3.3.
ontoTFIDF
and
ontoCSDF
ontoTFIDF
and
ontoCSDF
are
tw
o
no
v
el
methods
proposed
in
this
research.
The
y
are
methods
that
combine
w
ord-based
approach
and
ontology-based
approach,
aiming
to
calculate
combinations
of
statistical
as
semantic
scores
for
a
w
ord.
ontoTFIDF
and
ontoCSDF
are
dened
as
follo
ws:
ontoT
F
I
D
F
i
=
α
×
T
F
I
D
F
i
+
(1
−
α
)
×
O
ntoS
en
i
(5)
ontoC
S
D
F
i
=
α
×
C
S
D
F
i
+
(1
−
α
)
×
O
ntoS
en
i
(6)
where
ontoT
F
I
D
F
i
and
ontoC
S
D
F
i
are
the
sums
of
T
F
I
D
F
i
or
C
S
D
F
i
scores
and
ontoS
en
i
score,
respecti
v
ely
,
and
α
is
a
weight
used
to
balance
(tune)
the
scores
of
tw
o
methods
from
10%
to
90%
to
nd
the
best
combination.
4.
EXPERIMENT
AL
DESIGN
The
data
set
used
in
this
research
is
sentiment
analysis
data
on
opinions
re
g
arding
Amazon
products,
collected
by
Johns
Hopkins
Uni
v
ersity’
s
Department
of
Computer
Science
[20]
.
There
are
four
cate
gories
in
this
dataset;
books,
D
VDs,
electronics,
and
kitchenw
are.
In
the
book
cate
gory
,
there
are
1,025
documents
in
training
set
and
80
documents
in
the
testing
set.
F
or
D
VD
cate
gories,
there
are
1,017
documents
in
the
training
set
and
106
documents
in
the
testing
set.
F
or
electronics,
there
are
4,721
documents
in
the
training
set
and
1,182
documents
in
the
testing
set.
Finally
,
in
kitchenw
are
cate
gory
,
there
are
4,120
documents
in
the
training
set
and
1,031
documents
in
the
testing
set.
Therefore,
the
o
v
erall
number
of
documents
in
the
training
set
is
10,883
documents
and
2,399
documents
in
the
testing
set.
The
split
between
training
and
testing
sets
80%
–20%.
Each
cate
gory
in
thes
e
data
is
separated
into
four
classes:
v
ery
satisfying
(best),
satisf
action
(good),
lo
wer
than
e
xpectation
(bad),
dissatisf
action
(w
orst).
There
is
no
neutral
class
in
this
study
.
The
bag
of
w
ords
model
is
applied
in
this
research.
There
are
351,672
w
ords
in
the
training
set
and
100,066
w
ords
in
the
testing
set.
Data
cleansing
or
data
preprocessing
is
an
important
process
remo
ving
all
the
s
top
w
ords
and
special
characters,
and
retaining
only
necessary
data
for
the
analysis.
The
training
set
w
as
subjected
to
preprocessing.
Starti
ng
with
remo
ving
all
meaningless
w
ords,
stop
w
ords,
and
special
characters
from
351,672
w
ords
the
set
came
do
wn
to
about
200,000
w
ords.
After
that,
an
NL
TK
library
in
Python
called
nltkstopw
ord
w
as
used
to
remo
v
e
all
stop
w
ords;
bringing
the
w
ord
count
to
45,613
w
ords.
A
challenge
of
te
xt
mining
is
to
select
the
right
features
for
classication
[17],
[21],
[22].
Feature
section
decreases
the
size
of
the
data,
remo
v
es
duplicate
data,
and
remo
v
es
all
unrelated
w
ords.
It
helps
to
easily
nd
the
hidden
patterns
in
data
for
classi
cation.
Ov
erall
there
were
10,883
d
oc
u
m
ents;
and
there
were
o
v
erall
45,613
w
ords.
The
o
v
erall
document
count
is
four
times
less
than
the
o
v
erall
feature
count,
so
that
feature
section
[22]
is
v
ery
necessary
in
this
case.
Starting
with
choosing
meaningful
w
ords
both
in
English
dictionary
and
in
ontology
.
This
e
xperiment
uses
NL
TK
library
in
Python
called
PyEnchant
t
o
select
w
ords
in
the
English
dictionary
from
45,613
w
ords
this
selected
10,403
w
ords.
After
that,
the
features
are
reduced
by
choosing
part
of
speech.
The
study
[7]
found
that
sentiment
analysis
related
to
nouns,
v
erbs,
adjecti
v
es,
and
adv
erbs.
This
e
xperiment
uses
a
NL
TK
library
in
Python
namely
nltkpos–tag
to
choose
nouns,
v
erbs,
adjecti
v
es,
and
adv
erbs.
This
step
decreased
the
number
of
features
from
10,403
w
ords
to
3,080
w
ords.
Normally
,
there
are
tw
o
feature
selection
approaches;
lter
approach
and
wrapper
approach.
Ho
we
v
er
,
due
to
the
limited
time,
this
study
chose
the
lter
approach
ignoring
the
wrapper
approach
that
w
ould
need
a
lot
of
tim
e
to
nd
the
best
features.
Function
CfsSubsetEv
al
[23]
in
the
W
eka
package
[24]
w
as
applied
to
the
lter
approach
in
this
study
.
It
uses
correlation-based
feature
subset
selection
[25]
algorithm
to
select
the
features
that
relate
and
ef
fecti
v
ely
might
classify
the
data,
and
to
remo
v
e
duplicate
data.
This
function
returns
a
ranking
of
the
features
from
the
best
features
that
are
associated
with
the
classes,
in
descending
order
.
Therefore,
from
the
3,080
features,
only
the
top
400
features
were
chosen.
After
that,
there
are
some
documents
that
are
not
related
to
these
features.
Therefore,
some
documents
were
remo
v
ed
from
this
e
xperiment.
T
o
summarize,
there
4,903
documents
remaining
from
the
10,883
documents
in
the
training
set.
These
are
separated
into
four
classes;
v
ery
satisf
action
1,142
documents,
satisf
action
1,171
documents,
lo
wer
than
e
xpected
1,257
documents,
and
dissatisf
action
1,331
documents.
On
the
other
hand,
in
test
data
from
initial
2,399
documents
1,241
were
retained,
separated
into
four
classes;
v
ery
satisf
action
283
documents,
satisf
action
315
documents,
lo
wer
than
e
xpected
317
documents,
and
dissatisf
action
326
documents
as
sho
wn
in
T
able
1.
Indonesian
J
Elec
Eng
&
Comp
Sci,
V
ol.
42,
No.
3,
June
2026:
827–834
Evaluation Warning : The document was created with Spire.PDF for Python.
Indonesian
J
Elec
Eng
&
Comp
Sci
ISSN:
2502-4752
❒
831
The
selected
400
retained
features
were
used
in
this
e
xperiment.
The
data
in
these
les
is
the
frequenc
y
data
from
the
ori
g
i
nal
dataset,
both
in
training
and
testing
sets.
After
that,
data
representation
w
as
applie
d
by
calculating
the
scores
for
TF-IDF
,
CSDF
,
ontoTFIDF
,
and
ontoCSDF
,
follo
wing
the
equations
in
methodology
section.
The
ontoTFIDF
and
ontoCSDF
scores
were
calculated
for
alpha
v
alues
from
0.1
to
0.9.
The
nal
step
is
classifying
the
data,
and
comparing
the
classication
accuracies
of
v
e
alternati
v
e
classication
methods;
SMO
or
support
v
ector
machine
(SVM),
Na
¨
ıv
e
bayes,
J48
or
decision
tree,
IBk
or
k-nearest
neighbors
(kNN)
algorithm,
and
logistic
re
gression.
T
able
1.
Statistics
of
training
and
testing
data
Sentiment
Best
Good
W
orse
Bad
T
otal
T
raining
set
1142
1171
1259
1331
4903
T
esting
set
326
283
317
315
1241
T
otal
1468
1454
1576
1646
6144
5.
EXPERIMENT
AL
RESUL
TS
The
e
xperimental
results
are
sho
wn
in
T
able
2,
T
able
3,
and
Figure
1.
T
o
ensure
that
our
predicti
v
e
models
are
correct,
these
were
tested
by
the
training
data.
The
e
xperiment
al
results
sho
w
that
t
he
TF-IDF
g
a
v
e
poor
classic
ation
accurac
y
already
with
training
data
of
less
than
60%.
This
sho
ws
that
the
TF-IDF
data
representation
is
not
appropriate
for
creating
a
good
predicti
v
e
model,
in
this
case
study
.
On
the
other
hand,
the
e
xperimental
results
with
CSDF
g
a
v
e
classication
accurac
y
of
prediction
in
training
data
as
100%
with
IBk
or
kNN
approaches,
95.38%
with
J48
or
decision
tree
approach,
and
o
v
er
44%
b
ut
less
than
60%
with
Na
¨
ıv
e
bayes,
logistic
re
gression,
and
SMO
approaches.
Therefore,
CSDF
data
representation
may
allo
w
an
appropriate
prediction
model
by
using
IBk
or
J48
methods.
T
able
2.
Classication
accuracies
on
using
TF-IDF
,
CSDF
,
ontoTFIDF
,
and
ontoCSDF
data
Classication
accurac
y
Classication
method
Document
representation
method
Prediction
with
training
data
Prediction
with
testing
data
Na
¨
ıv
e
Bayes
TFIDF
41.79
40.53
CSDF
44.54
43.23
Logistic
TFIDF
50.68
41.5
CSDF
58.31
51.68
SMO
TFIDF
48.91
41.02
CSDF
59.93
55.7
IBk
TFIDF
59
37.39
CSDF
100
96.47
J48
TFIDF
44.1
38.2
CSDF
95.38
96.56
T
able
3.
Classication
accuracies
on
using
TF-IDF
,
CSDF
,
ontoTFIDF
,
and
ontoCSDF
data
(Continued)
Classication
accurac
y
Classication
method
Document
representation
method
α
v
alues
of
ontoTFIDF
and
ontoCSDF
0.1
0.2
0.3
0.4
0.5
0.6
0.7
0.8
0.9
Na
¨
ıv
e
Bayes
TFIDF
40.85
40.94
40.94
41.34
40.94
38.76
37.79
35.38
36.1
CSDF
34.04
39.9
41.43
41.43
40.12
40.77
40.77
41.26
41.26
Logistic
TFIDF
41.82
41.74
41.74
41.1
41.18
41.5
41.42
41.66
41.02
CSDF
44.46
44.63
44.87
44.71
45.12
45.2
45.37
45.86
48.07
SMO
TFIDF
40.69
41.42
41.1
41.02
40.61
40.77
40.69
40.85
40.94
CSDF
42.49
42.66
42.74
43
43.15
44.46
47.17
47.74
48.24
IBk
TFIDF
38.03
37.47
37.31
37.71
38.11
38.2
38.03
37.95
37.79
CSDF
63.5
78.5
80.23
82.61
86.4
87.04
88.27
88.52
90.4
J48
TFIDF
38.2
38.2
38.2
38.2
38.2
38.2
38.2
38.2
38.2
CSDF
35.03
37.87
36.26
39.21
39.71
39.62
39.71
40.2
40.28
On
e
xploring
te
xt
mining
appr
oac
hes
to
sentiment
analysis
based
on
the
...
(Suthir
a
Plansangk
et)
Evaluation Warning : The document was created with Spire.PDF for Python.
832
❒
ISSN:
2502-4752
The
e
xperimental
results
from
using
testing
set
g
a
v
e
better
classication
accurac
y
for
CSDF
data
than
for
TF-IDF
data
with
an
y
classication
method.
F
or
J48
and
IBk
approaches,
CSDF
data
g
a
v
e
classication
accurac
y
of
about
96%.
Ho
we
v
er
,
the
classication
accuracies
of
SMO,
Na
¨
ıv
e
bayes,
and
logistic
re
gression
were
belo
w
56%
with
all
data
representation
methods,
including
CSDF
.
ontoTFIDF
combines
the
scores
from
ontoS
en
and
TF-IDF
,
while
ontoCSDF
combines
the
scores
from
ontoSen
and
CSDF
.
The
combination
uses
alpha
as
a
weighting
f
actor
to
balance
the
tw
o
combined
scores,
with
v
alues
from
0.1
to
0.9.
The
e
xperimental
results
sho
w
that
the
classication
accurac
y
of
ontoCSDF
data
with
all
alpha
v
alues
were
lo
wer
than
with
CSDF
data.
Increasing
the
ontoSen
scores
de
graded
classication
accurac
y
belo
w
that
with
CSDF
data.
Therefore,
ontoSen
did
not
impro
v
e
the
performance
of
classication.
On
the
other
hand,
the
e
xperimental
resul
ts
with
ontoTFIDF
which
combines
ontoSen
and
TF-IDF
in
the
dif
ferent
proportions
are
sho
wn
in
Figure
2.
It
w
as
found
that
the
classication
accuracies
of
logistic
re
gression,
SMO,
and
J48
methods
were
almost
similar
to
each
other
with
TF-IDF
data.
In
addition,
the
classication
accurac
y
of
ontoTFIDF
by
Na
¨
ıv
e
Bayes
method
gradually
decreased;
while
the
classication
accurac
y
of
ontoTFIDF
by
IBk
method
gradually
increased
from
37.39%
of
original
TF-IDF
to
the
maximum
of
38.2%
at
alpha
=
0.6.
W
e
used
a
t-test
assuming
unequal
v
ariances
to
determine
if
there
is
a
signicant
dif
ference
between
the
means
of
the
classication
accurac
y
of
ontoTFIDF
.
The
results
are
sho
wn
in
T
able
4:
the
classication
accurac
y
with
ontoTFIDF
data
w
as
signicantly
better
than
that
with
TF-IDF
data
on
using
IBk
method,
at
a
signicance
le
v
el
of
0.05.
Therefore,
ontoTFIDF
helped
signicantly
impro
v
e
the
performance
of
sentiment
analysis
by
IBk
classication
method,
compared
to
TF-IDF
scoring.
Figure
1.
The
classication
accuracies
with
TF-IDF
,
CSDF
,
ontoTFIDF
,
and
ontoCSDF
data
Figure
2.
The
classication
accuracies
with
TF-IDF
and
ontoTFIDF
data
Indonesian
J
Elec
Eng
&
Comp
Sci,
V
ol.
42,
No.
3,
June
2026:
827–834
Evaluation Warning : The document was created with Spire.PDF for Python.
Indonesian
J
Elec
Eng
&
Comp
Sci
ISSN:
2502-4752
❒
833
T
able
4.
t-T
est
comparison
of
using
TFIDF
and
ontoTFIDF
(T
w
o-sample
t-test
assuming
unequal
v
ariances)
TFIDF
ontoTFIDF
Mean
37.35
37.91125
V
ariance
0.0032
0.057498
Observ
ations
h
ypothesized
mean
dif
ference
0
df
8
t
Stat
-5.98727
P(T¡=t)
one-tail
0.000164
t
Critical
one-tail
1.859548
P(T¡=t)
tw
o-tail
0.000328
t
Critical
tw
o-tail
2.306004
6.
CONCLUSION
This
study
e
xplored
tw
o
no
v
el
approaches
to
impro
v
e
the
classication
accurac
y
in
sentiment
analysi
s
by
using
combinations
of
a
w
ord-based
approach
(TF-IDF
or
CSDF)
with
an
ontology-based
approach
(on-
toSen)
gi
ving
tw
o
ne
w
methods,
called
ontoTF-IDF
and
ontoCSDF
.
The
e
xperimental
results
sho
w
that
based
on
classication
accurac
y
the
CSDF
w
as
the
best
data
representation
m
ethod
for
sentiment
analysis,
using
J48
or
decision
tree,
and
IBk
or
kNN
classication
methods.
The
best
classication
performance
w
as
about
96%.
Ho
we
v
er
,
a
combination
of
CSDF
and
ontology-based
approach,
ontoCSDF
,
did
not
impro
v
e
the
classication
of
sentiment
analysis
data.
On
the
other
hand,
the
classication
of
TF-IDF
data
w
as
not
good
enough;
there-
fore,
the
combination
of
TF-IDF
and
ontoSen
(ontoTFIDF)
impro
v
ed
the
classication
accurac
y
signicantly
(p
<
0.05)
on
using
IBk
classier
.
7.
DISCUSSION
AND
FUTURE
W
ORK
From
the
observ
ations,
CSDF
method
with
normalization
adjusts
the
v
alues
measured
on
dif
ferent
scales
to
a
no
t
ionally
common
scale.
The
CSDF
v
alues
are
only
from
0
to
1.
This
may
ha
v
e
impro
v
ed
the
performance
of
classication
of
sentiment
analysis
data.
Furthermore,
instead
of
using
the
ontoSen
method,
there
are
alternati
v
e
ontology-based
approaches
that
could
impro
v
e
the
performance
of
sentiment
analysis
in
the
future.
REFERENCES
[1]
J.
T
ao
and
T
.
T
an,
“
Af
fecti
v
e
computing:
a
re
vie
w
,
”
in
International
Confer
ence
on
Af
fective
computing
and
intellig
ent
inter
action
,
J.
T
ao,
T
.
T
an,
and
R.
W
.
Picard,
Eds.,
in
Lecture
Notes
in
Computer
Science.
Berlin,
Heidelber
g:
Springer
Berlin
Heidelber
g,
2005,
pp.
981–995.
doi:
10.1007/11573548
125.
[2]
T
.
Nasuka
w
a
and
J.
Y
i,
“Sentiment
analysis:
capturing
f
a
v
orability
us
ing
natural
language
processing,
”
in
Pr
oceedings
of
the
2nd
International
Confer
ence
on
Knowledg
e
Captur
e
,
K-CAP
2003
,
Ne
w
Y
ork,
NY
,
USA:
A
CM,
Oct.
2003,
pp.
70–77,
doi:
10.1145/945645.945658.
[3]
R.
Prabo
w
o
and
M.
Thel
w
all,
“Sentiment
analysis:
a
combined
approach,
”
J
ournal
of
Informetrics
,
v
ol.
3,
no.
2,
pp.
143–157,
Apr
.
2009,
doi:
10.1016/j.joi.2009.01.003.
[4]
A.
Ag
arw
al,
B.
Xie,
I.
V
o
vsha,
O.
Rambo
w
,
and
R.
P
assonneau,
“Sentiment
analysis
of
T
witter
data,
”
in
Pr
oceedings
of
the
W
orkshop
on
Langua
g
e
in
Social
Media
(LSM
2011)
,
2011,
pp.
30–38.
[5]
E.
K
ouloumpis,
T
.
W
ilson,
and
J.
Moore,
“T
witter
sentiment
analysis:
the
good
the
bad
and
the
OMG!,
”
in
Pr
oceed-
ings
of
the
5th
Internati
onal
AAAI
Confer
ence
on
W
eblo
gs
and
Social
Media,
ICWSM
2011
,
Aug.
2011,
pp.
538–541,
doi:
10.1609/icwsm.v5i1.14185.
[6]
S.
T
an
and
J.
Zhang,
“
An
empirical
study
of
sentiment
analysis
for
chinese
documents,
”
Expert
Systems
with
Applications
,
v
ol.
34,
no.
4,
pp.
2622–2629,
May
2008,
doi:
10.1016/j.esw
a.2007.05.028.
[7]
F
.
Benamara,
C.
Cesarano,
A.
Picariello,
D.
Refor
giato,
and
V
.
S.
Subrahmanian,
“Sentiment
analysis:
adjecti
v
es
and
adv
erbs
are
better
than
adjecti
v
es
alone,
”
ICWSM
2007
-
International
Confer
ence
on
W
eblo
gs
and
Social
Media
,
v
ol.
7,
pp.
203–206,
2007.
[8]
Y
.
Zhang,
R.
Jin,
and
Z.
H.
Zhou,
“Understanding
bag-of-w
ords
model:
a
statistical
frame
w
ork,
”
International
J
ournal
of
Mac
hine
Learning
and
Cybernetics
,
v
ol.
1,
no.
1–4,
pp.
43–52,
Dec.
2010,
doi:
10.1007/s13042-010-0001-0.
[9]
G.Salt
on
and
M.
J.
McGill,
Intr
oduction
to
modern
information
r
etrie
val
,
1983.
[10]
S.
Plansangk
et
and
J.
Q.
Gan,
“
A
ne
w
term
weighting
scheme
based
on
class
specic
document
frequenc
y
for
document
repre-
sentation
and
classicati
on,
”
in
2015
7th
Computer
Science
and
El
ectr
onic
Engineering
Confer
ence
,
CEEC
2015
-
Confer
ence
Pr
oceedings
,
IEEE,
Sep.
2015,
pp.
5–8,
doi:
10.1109/CEEC.2015.7332690.
[11]
K.
Munir
and
M.
S.
Anjum
,
“The
use
of
ontologies
for
ef
fecti
v
e
kno
wledge
modelling
and
information
retrie
v
al,
”
Applied
Computing
and
Informatics
,
v
ol.
14,
no.
2,
pp.
116–126,
Jul.
2018,
doi:
10.1016/j.aci.2017.07.003.
On
e
xploring
te
xt
mining
appr
oac
hes
to
sentiment
analysis
based
on
the
...
(Suthir
a
Plansangk
et)
Evaluation Warning : The document was created with Spire.PDF for Python.
834
❒
ISSN:
2502-4752
[12]
M
.
Dragoni,
S.
Poria,
and
E.
Cambria,
“OntoSenticNet:
a
commonsense
ontology
for
sentiment
analysis,
”
IEEE
Intellig
ent
Systems
,
v
ol.
33,
no.
3,
pp.
77–85,
May
2018,
doi:
10.1109/MIS.2018.033001419.
[13]
L.
Floridi,
The
Blac
kwell
Guide
to
the
Philosophy
of
Computing
and
Information
,
John
W
ile
y
.
2008,
doi:
10.1111/b
.9780631229193.2003.00008.x.
[14]
A.
A.
Salati
no,
T
.
Thanapalasing
am,
A.
Mannocci,
A.
Biruk
ou,
F
.
Osborne,
and
E.
Motta,
“The
computer
science
ontology:
a
comprehensi
v
e
automati
cally-generated
taxonomy
of
research
areas,
”
Data
Intellig
ence
,
v
ol.
2,
no.
3,
pp.
379–416,
Jul.
2020,
doi:
10.1162/dint
a
00055.
[15]
A.
Sharma
and
S.
K
umar
,
“Machine
learning
and
ontology-based
no
v
el
semantic
document
i
nde
xi
ng
for
information
retrie
v
al,
”
Computer
s
and
Industrial
Engineering
,
v
ol.
176,
p.
108940,
Feb
.
2023,
doi:
10.1016/j.cie.2022.108940.
[16]
D.
K.
Jain,
S.
Qamar
,
S.
R.
Sangw
an,
W
.
Ding,
and
A.
J.
K
ulkarni,
“Ontology-based
natural
language
processing
for
sentimen-
tal
kno
wledge
analysis
using
deep
learning
architectures,
”
A
CM
T
r
ansactions
on
Asian
and
Low-Resour
ce
Langua
g
e
Information
Pr
ocessing
,
v
ol.
23,
no.
1,
pp.
1–17,
2024,
doi:
10.1145/3624012.
[17]
R
.
Baeza-Y
ates
and
B.
Ribeiro-Neto,
Modern
information
r
etrie
val
,
v
ol.
463.
Ne
w
Y
ork,
1999.
[18]
D.
Jurafsk
y
and
J.
H.
Martin,
Speec
h
and
langua
g
e
pr
ocessing:
an
intr
oduction
to
natur
al
langua
g
e
pr
ocessing
,
computational
linguistics,
and
speec
h
r
eco
gnition
.
2000.
[19]
K
.
W
´
ojcik
and
J.
T
ucho
wski,
“Ontology
based
approach
to
sentiment
analysis,
”
2014,
[Online].
A
v
ailable:
http://www
.researchg
ate.net/publication/267324473
Ontology
Based
Approach
to
Sentiment
Analysis.
[20]
J
.
Blitzer
,
M.
Dredze,
and
F
.
Pereira,
“Multidomain
sentiment
dataset
(v
ersion
2.0).,
”
Johns
Hopkins
Uni
v
ersity
,
2009.
Accessed:
Mar
.
23,
2009.
[Online].
A
v
ailable:
https://www
.cs.jhu.edu/
mdredze/datasets/sentiment/
[21]
R
.
O.
Duda
and
P
.
E.
Hart,
P
attern
classication
.
ne
w
york,
2006.
[22]
W
.
B.
P
o
well,
Appr
oximate
dynamic
pr
o
gr
amming:
solving
the
cur
ses
of
dimensionality
,
v
ol.
703.
2007.
[23]
Uni
v
ers
ity
of
W
aikato,
“Class
CfsSubsetEv
al,
”
W
eka
Application
Programming
Interf
ace
Document.
[Online].
A
v
ailable:
http://weka.sourcefor
ge.net/doc.de
v/weka/attrib
uteSelection/CfsSubsetEv
al.html
[24]
M
achine
Learning
Group
Uni
v
ersity
of
W
aikato,
“WEKA
-
The
w
orkbench
for
machine
learning,
”
Uni
v
ersity
of
W
aikato.
[Online].
A
v
ailable:
https://www
.cs.w
aikato.ac.nz/ml/weka/inde
x.html
[25]
M
.
A.
Hall,
“Correlation-based
feature
subset
selection
for
machine
learning,
”
Uni
v
er
sity
of
W
aikato,
1998.
BIOGRAPHIES
OF
A
UTHORS
Suthira
Plansangk
et
is
a
lecturer
at
Di
vision
of
Computational
Science,
F
aculty
of
Sci-
ence,
Prince
of
Songkla
Uni
v
ersity
,
Songkhla,
Thailand.
She
recei
v
ed
the
B.S.
(Computer
Science)
de
gree
in
2002,
M.S.
(Computer
Science)
in
2006
from
Prince
of
Songkla
Uni
v
ersity
,
Songkhla,
Thailand
and
Ph.D.
(Computer
Science)
in
2017
from
Uni
v
ersity
of
Esse
x,
Colchester
,
UK.
Her
re-
search
interests
are
in
machine
learning,
data
science,
and
information
retrie
v
al.
She
can
be
contacted
at
email:
suthira.p@psu.ac.th.
Supapor
n
Kansomk
eat
recei
v
ed
the
B.S.
(Mathematics)
de
gre
e
in
1991,
M.S.
(Com-
puter
Science)
in
1995
from
Prince
of
Songkla
Uni
v
ersity
and
D.
(Computer
Engineering)
in
2007
from
Chulalongk
orn
Uni
v
ersity
.
Since
1996,
she
has
been
instructor
at
Di
vision
of
Computational
Science,
F
aculty
of
Science,
Princ
e
of
Songkla
Uni
v
ersity
,
Songkhla,
Thailand.
Her
research
is
con-
cerned
with
softw
are
testing,
image
processing,
and
machine
learning.
She
can
be
contacted
at
email:
supaporn.k@psu.ac.th.
Supasit
Kajkamhaeng
is
a
lecturer
in
the
Information
and
Communication
T
echnol-
ogy
Programme,Di
vision
of
Computational
Science,
F
aculty
of
Science,
Prince
of
Songkla
Uni-
v
ersity
,
Songkhla,
Thailand.
He
recei
v
ed
the
master’
s
de
gree
in
the
Department
of
Computer
En-
gineering,
F
aculty
of
Engineering,
Kasetsart
Uni
v
ersity
,
Thailand.
His
research
interests
include
parallel/distrib
uted
computing
and
high
performance
computing.
He
can
be
contacted
at
email:
supasit.k@psu.ac.th.
Indonesian
J
Elec
Eng
&
Comp
Sci,
V
ol.
42,
No.
3,
June
2026:
827–834
Evaluation Warning : The document was created with Spire.PDF for Python.