TY - CHAP
T1 - Managing data uncertainty in automatic mapping of clinical classification systems
AU - Purja Pun, Santosh
AU - Obst, Oliver
AU - Basilakis, Jim
AU - Ginige, Jeewani Anupama
PY - 2025
Y1 - 2025
N2 - Mapping clinical classification systems, like the International Classification of Disease (ICD) across different versions and other external clinical classifications systems, is challenging and often done manually by trained professionals. Among others, variation in the code descriptions to describe the same clinical condition in different versions poses a unique challenge to implementing automated mapping systems. We call this data uncertainty. Existing lexical-based methods attempt to solve this problem by generating alternative terms using synonyms. This work addresses the data uncertainty by learning a probabilistic embedding for each code description using similar terms and paraphrases. A valid code pair must exhibit proximity in the embedding space and have a comparable distribution. Additionally, we propose a new evaluation metric that considers the hierarchical structure of ICD to evaluate the performance of an automated mapping system. We demonstrate the effectiveness of our approach by mapping ICD-9-CM (Clinical Modification) and ICD-10-CM, ICD-10-AM (Australian Modification) and ICD-11 in both directions. The source code will be available at: https://github.com/Xujan24/wt-KL
AB - Mapping clinical classification systems, like the International Classification of Disease (ICD) across different versions and other external clinical classifications systems, is challenging and often done manually by trained professionals. Among others, variation in the code descriptions to describe the same clinical condition in different versions poses a unique challenge to implementing automated mapping systems. We call this data uncertainty. Existing lexical-based methods attempt to solve this problem by generating alternative terms using synonyms. This work addresses the data uncertainty by learning a probabilistic embedding for each code description using similar terms and paraphrases. A valid code pair must exhibit proximity in the embedding space and have a comparable distribution. Additionally, we propose a new evaluation metric that considers the hierarchical structure of ICD to evaluate the performance of an automated mapping system. We demonstrate the effectiveness of our approach by mapping ICD-9-CM (Clinical Modification) and ICD-10-CM, ICD-10-AM (Australian Modification) and ICD-11 in both directions. The source code will be available at: https://github.com/Xujan24/wt-KL
KW - Data Uncertainty
KW - International Classification of Disease
KW - Mapping Tables
KW - Probabilistic Embedding
UR - https://www.scopus.com/pages/publications/105009267187
UR - https://go.openathens.net/redirector/westernsydney.edu.au?url=https://doi.org/10.1007/978-981-96-8298-0_23
U2 - 10.1007/978-981-96-8298-0_23
DO - 10.1007/978-981-96-8298-0_23
M3 - Chapter
AN - SCOPUS:105009267187
SN - 9789819682973
T3 - Lecture Notes in Computer Science
SP - 284
EP - 295
BT - Data Science: Foundations and Applications: 29th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2025, Sydney, Australia, June 10-13, 2025, Proceedings, Part VII
A2 - Wu, Xintao
A2 - Spiliopoulou, Myra
A2 - Wang, Can
A2 - Kumar, Vipin
A2 - Cao, Longbing
A2 - Zhou, Xiangmin
A2 - Pang, Guansong
A2 - Gama, Joao
PB - Springer
CY - Singapore
T2 - 29th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2025
Y2 - 10 June 2025 through 13 June 2025
ER -