TY - CHAP
T1 - Using pseudo-synonyms to generate embeddings for clinical terms
AU - Purja Pun, Santosh
AU - Obst, Oliver
AU - Basilakis, Jim
AU - Ginige, Jeewani Anupama
PY - 2025
Y1 - 2025
N2 - Existing approaches attempt to explicitly learn clinical term embedding from clinical datasets by training a model, such as word2vec and recurrent neural network or fine-tuning a pre-trained large language model (LLM). While the corpus-based methods require exposure to a rich vocabulary in the training corpus, insufficient contextual information, in clinical terms, makes LLMs prone to failure to generate meaningful embeddings. In this regard, we propose a novel method to generate embeddings for clinical terms using pseudo-synonyms - terms that might be associated with a clinical term but not the exact synonyms. The proposed method uses an LLM as a black-box tool and requires no training or fine-tuning. To demonstrate the effectiveness of the learned embeddings, we compared our approach with existing corpus-based embedding approaches on semantic textual similarity (STS) tasks on five benchmark datasets. Our proposed method outperformed all existing approaches (https://github.com/Xujan24/pseudo-synonyms-for-clinical-term-embedding).
AB - Existing approaches attempt to explicitly learn clinical term embedding from clinical datasets by training a model, such as word2vec and recurrent neural network or fine-tuning a pre-trained large language model (LLM). While the corpus-based methods require exposure to a rich vocabulary in the training corpus, insufficient contextual information, in clinical terms, makes LLMs prone to failure to generate meaningful embeddings. In this regard, we propose a novel method to generate embeddings for clinical terms using pseudo-synonyms - terms that might be associated with a clinical term but not the exact synonyms. The proposed method uses an LLM as a black-box tool and requires no training or fine-tuning. To demonstrate the effectiveness of the learned embeddings, we compared our approach with existing corpus-based embedding approaches on semantic textual similarity (STS) tasks on five benchmark datasets. Our proposed method outperformed all existing approaches (https://github.com/Xujan24/pseudo-synonyms-for-clinical-term-embedding).
KW - Clinical Term Embedding
KW - Large Language Models
KW - Semantic Textual Similarity
UR - https://www.scopus.com/pages/publications/105009239605
UR - https://go.openathens.net/redirector/westernsydney.edu.au?url=https://doi.org/10.1007/978-981-96-8298-0_17
U2 - 10.1007/978-981-96-8298-0_17
DO - 10.1007/978-981-96-8298-0_17
M3 - Chapter
AN - SCOPUS:105009239605
SN - 9789819682973
T3 - Lecture Notes in Computer Science
SP - 209
EP - 220
BT - Data Science: Foundations and Applications: 29th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2025, Sydney, Australia, June 10-13, 2025, Proceedings, Part VII
A2 - Wu, Xintao
A2 - Spiliopoulou, Myra
A2 - Wang, Can
A2 - Kumar, Vipin
A2 - Cao, Longbing
A2 - Zhou, Xiangmin
A2 - Pang, Guansong
A2 - Gama, Joao
PB - Springer
CY - Singapore
T2 - 29th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2025
Y2 - 10 June 2025 through 13 June 2025
ER -