Contrastive learning of adverse events to provide effective and interpretable vector representations for machine-assisted pharmacovigilance

Balogh Olivér Márton; Pétervári Mátyás; Csernák Áron Márk; Puhl Eszter; Horváth András; Ferdinandy Péter; Ágg Bence: Contrastive learning of adverse events to provide effective and interpretable vector representations for machine-assisted pharmacovigilance.
BRIEFINGS IN BIOINFORMATICS, 27 (5). ISSN 1467-5463 (2026)

[thumbnail of bbag463.pdf] Szöveg
bbag463.pdf - Megjelent verzió

Download (4MB)
Mű típusa: Folyóiratcikk
Szerző azonosítók:
NévORCIDMTMT szerző azonosító
Balogh Olivér Márton0000-0002-6296-638310082699
Pétervári Mátyás0000-0002-4816-074610065329
Csernák Áron Márk0000-0002-6040-461910100490
Puhl Eszter0000-0001-8501-147810072501
Horváth András0000-0001-5855-418610029872
Ferdinandy Péter0000-0002-6424-680610010125
Ágg Bence0000-0002-6492-042610057401
Absztrakt (kivonat): Post-marketing surveillance is crucial for drug safety, yet the tools of pharmacovigilance rely solely on text-based data that may limit contemporary machine learning methodologies in the support of decision-making. With the recent surge of employing large language models (LLMs) for text-based tasks, there also arises an unmet need for a different approach which is not grounded in the linguistic patterns of unfiltered natural text, like LLMs, but rather based on real-world drug safety data. Here, we adapt contrastive learning algorithms to generate adverse event vector representations from spontaneous adverse event reports to serve as machine-readable (i.e. numerical) resources for downstream pharmacovigilance applications, such as drug-event association prediction for signal detection or causality assessment. We present comprehensive interpretability analyses of the resulting representations through density-based clustering, semantic evaluation, and comparison of multivariate dispersions, revealing patterns that reflect both functional and causal relations of the adverse events while also capturing drug-safety-related information better than existing medical terminologies and encoder-only LLMs. Furthermore, we demonstrate the applicability of our representations as input features in our downstream classifier model, outperforming the reporting odds ratio method, commonly used by regulatory agencies, and also LLM-generated representations (area under the receiver operating characteristic curve: 0.88 versus 0.76-0.83) on drug-event association prediction benchmarks. Therefore, we propose an interpretable adverse event vector representation, serving as a general resource that could enable the development of a wide array of machine learning applications to support decision-making in pharmacovigilance and facilitate patient safety.
Folyóirat címe: BRIEFINGS IN BIOINFORMATICS
Megjelenés éve: 2026
Kötet: 27
Szám: 5
ISSN: 1467-5463
Intézmény: Pázmány Péter Katolikus Egyetem
Kar: Információs Technológiai és Bionikai Kar (2013.07.-)
Nyelv: angol
MTMT rekordazonosító: 37504014
DOI azonosító: 10.1093/bib/bbag463
Scopus azonosító: 105049008705
WoS azonosító: 001866637100001
Dátum: 2026. Okt. 01. 16:14
Utolsó módosítás: 2026. Okt. 01. 16:14
URI: https://publikacio.ppke.hu/id/eprint/3747

Actions (login required)

Tétel nézet Tétel nézet