Word-Level Embedding to Improve Performance of Representative Spatio-temporal Document Classification

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

Tokenization is the process of segmenting the input text into smaller units of text, and it is a preprocessing taskthat is mainly performed to improve the efficiency of the machine learning process. Various tokenizationmethods have been proposed for application in the field of natural language processing, but studies haveprimarily focused on efficiently segmenting text. Few studies have been conducted on the Korean language toexplore what tokenization methods are suitable for document classification task. In this paper, an exploratorystudy was performed to find the most suitable tokenization method to improve the performance of arepresentative spatio-temporal document classifier in Korean. For the experiment, a convolutional neuralnetwork model was used, and for the final performance comparison, tasks were selected for documentclassification where performance largely depends on the tokenization method. As a tokenization method forcomparative experiments, commonly used Jamo, Character, and Word units were adopted. As a result of theexperiment, it was confirmed that the tokenization of word units showed excellent performance in the case ofrepresentative spatio-temporal document classification task where the semantic embedding ability of the tokenitself is important.

키워드

Spatio-temporal Document ClassificationTokenizationWord -Level EmbeddingTEXT CLASSIFICATION
제목
Word-Level Embedding to Improve Performance of Representative Spatio-temporal Document Classification
저자
Kim, ByoungwookJang, Hong-Jun
DOI
10.3745/JIPS.04.0296
발행일
2023-12
유형
Article
저널명
JIPS(Journal of Information Processing Systems)
19
6
페이지
830 ~ 841