Woźniak Rafał (Lodz University of Technology), Ożdżyński Piotr (Lodz University of Technology), Zakrzewska Danuata (Lodz University of Technology)
Cluster Analysis of Medical Text Documents by Using Semi-Clustering Approach Based on GRAPH Representation
Information Systems in Management, 2018, vol. 7, nr 3, s. 213-224, rys., tab., bibliogr. 14 poz.
Analiza skupień, Eksploracja tekstu
Cluster analysis, Text mining
The development of Internet resulted in an increasing number of online text repositories. In many cases, documents are assigned to more than one class and automatic multi-label classification needs to be used. When the number of labels exceeds the number of the documents, effective label space dimension reduction may significantly improve classification accuracy, what is a major priority in the medical field. In the paper, we propose document clustering for label selection. We use semi-clustering method, by considering graph representation, where documents are represented by vertices and edge weights are calculated according to their mutual similarity. Assigning documents to semi-clusters helps in reducing number of labels, further used in multi-label classification process. The performance of the method is examined by experiments conducted on real medical datasets. (original abstract)
