Large Language Models for Semantic Join: A Comprehensive Survey

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

3

초록

Semantic join, the operation of integrating information siloed across heterogeneous data sources, is critical for modern data science, yet traditional methods have long been hampered by the persistent challenges of deep semantic ambiguity, poor scalability, and prohibitive manual intervention. This survey posits that Large Language Models (LLMs) represent a promising new paradigm, offering the potential to overcome these long-standing hurdles. By replacing brittle rules and syntactic analysis with deep contextual understanding, LLMs leverage their core capabilities in contextual representation learning and in-context learning (ICL) to automate and significantly improve the accuracy of linking records based on their underlying conceptual relatedness. We provide a comprehensive and structured review of this burgeoning field, synthesizing the state-of-the-art from foundational methodologies-such as data textualization and bi/cross-encoder architectures-to advanced techniques including Retrieval-Augmented Generation (RAG), structured prompting for complex reasoning, and optimizations for scalability. Furthermore, we survey transformative applications across enterprise data management, knowledge graph (KG) construction, and scientific research. By consolidating current knowledge, structuring the landscape of techniques, and identifying key open questions, this survey aims to catalyze future research and guide the development of the next generation of more powerful, reliable, and responsible semantic data integration solutions.

키워드

Large language modelsLarge language modelssemantic joinsemantic joindata integrationdata integrationSIMILARITYEMBEDDINGS
제목
Large Language Models for Semantic Join: A Comprehensive Survey
저자
Hong, KijaePark, Yeonsu
DOI
10.1109/ACCESS.2025.3625753
발행일
2025
유형
Article
저널명
IEEE Access
13
페이지
184478 ~ 184493