상세 보기
뉴스 문서 계층적·점증적 군집화 성능 향상을 위한 언어 모델 기반 문서 표현 방법 고찰
- 김병욱;
- 정경환;
- 이승우;
- 양영욱;
- 장홍준
초록
This study proposes a language model-based summarization approach for document representation to enhance the performance of hierarchical and incremental clustering of news articles. While using the BERT model for document representation, the input token length is limited to 512 tokens, which can lead to information loss in lengthy documents. To address this issue, this study explores an alternative approach that utilizes article summaries instead of full-text articles for document representation. Experimental results confirm that the summarization-based document representation method significantly improves hierarchical and incremental clustering performance compared to using the full article text as input. Furthermore, it is empirically verified that enhancing the quality of summaries generated by LLMs through prompt engineering further improves clustering performance. Among the evaluated document representation strategies, the most effective approach was using the original title and the first sentence together. The second-best performing method involved incorporating the original title, the LLM-generated title, and the first sentence. This study contributes to improving the performance of hierarchical and incremental clustering for news articles and is expected to be applicable to various document clustering tasks in the future.
키워드
- 제목
- 뉴스 문서 계층적·점증적 군집화 성능 향상을 위한 언어 모델 기반 문서 표현 방법 고찰
- 제목 (타언어)
- Investigation of a Language Model-Based Document Representation Method for Improving Hierarchical and Incremental Clustering Performance of News Articles
- 저자
- 김병욱; 정경환; 이승우; 양영욱; 장홍준
- 발행일
- 2025-04
- 유형
- Y
- 저널명
- 데이타베이스연구
- 권
- 41
- 호
- 1
- 페이지
- 53 ~ 75