Comparative analyses of multilingual drug entity recognition systems for clinical case reports in cardiology
File(s)ICUE_BioASQ.pdf (639.5 KB)
Published version
Author(s)
Lee, C
Simpson, TI
Posma, JM
Lain, AD
Type
Conference Paper
Abstract
Performance disparities exist in Named Entity Recognition (NER) systems across languages due to variations in available human-annotated data. We participated in the MultiDrug subtask of MultiCardioNER, a shared task focusing on multilingual NER for cardiology, to compare the effectiveness of fine-tuning BERT-based monolingual and multilingual language models, and prompting Large Language Models (LLMs) for drug entity recognition across multiple languages. Our findings demonstrate that monolingual BERT models pretrained on biomedical corpora generally outperform their multilingual counterparts. However, for languages lacking access to a broader range of pretrained models, combining the translation capability of LLM [1, 2, 3, 4] with the best-performing pretrained monolingual BERT model yielded superior results. This approach effectively reduces the resource disparity while leveraging domain-specific knowledge captured by the monolingual BERT model. Our best systems in the MultiCardioNER track yielded F1-scores of 0.9277 for Spanish, 0.9107 for English, and 0.8776 for Italian. We highlight the comparative advantages of domain-specific fine-tuning and LLM-powered language translation for multilingual drug NER.
Date Issued
2024-09-01
Date Acceptance
2024-09-01
Citation
CEUR Workshop Proceedings, 2024, 3740, pp.159-167
ISSN
1613-0073
Start Page
159
End Page
167
Journal / Book Title
CEUR Workshop Proceedings
Volume
3740
Copyright Statement
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0)
License URL
Source
Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2024)
Publication Status
Published
Start Date
2024-09-09
Finish Date
2024-09-12
Coverage Spatial
Grenoble, France