Misinformation Detection in Low-Resource Languages and Health Domains: Review and Evaluation Framework

Authors

  • Bassey Isong North-West University
  • Rose Linah Computer Science Department, North-West University, Mafikeng, South Africa

DOI:

https://doi.org/10.33022/ijcs.v15i3.5129

Keywords:

Misinformation detection, LRLs, Health misinformation, Explanable AI, Multimodal detection.

Abstract

Misinformation on digital platforms harms public health decisions, electoral processes, and institutional trust. Its detection in low-resource languages (LRLs) remains structurally neglected. No prior survey applies a scoring framework to assess study quality or enable cross-study comparison. This review examines 51 peer-reviewed studies published between 2021 and 2026, following the PRISMA 2020 protocol across four dimensions: detection methodology, dataset coverage, health-domain adaptation, and explainable AI (XAI) integration, and proposes the LRL misinformation evaluation framework (LRLM-EF). The five-criterion evaluation framework was applied retrospectively to all reviewed studies. The findings reveal that dataset construction is the dominant research activity, with new corpora built for Amharic, Bangla, Bengali, Luganda, Sepedi, Sesotho, Xitsonga, isiZulu, Kurdish Sorani, and several Arabic dialects. Transformer-based models outperform classical and deep learning baselines in most settings; classical classifiers achieve comparable results where annotated data is scarce. Health-domain coverage is narrow. Research mostly concentrates on COVID-19 and vaccine misinformation, while HIV, malaria, and reproductive health appear in no reviewed studies. Multimodal fusion improves detection in all five studies where it was tested, yet audio-based detection in any LRL setting is absent. XAI is applied in a few studies, exclusively through post-hoc LIME, with no study evaluating its effect on user decisions. LRLM-EF scoring reveals that most studies address fewer than half the framework criteria, with adversarial evaluation and standardised reporting as the weakest dimensions. However, two contradictions exist in the evidence. Classical retrieval outperforms neural similarity on rare-terminology datasets, and augmentation volume shows no reliable accuracy gain, which further expose absence of a shared benchmarking standard.

Downloads

Published

29-06-2026