The Minimal isiNdebele Explicit and Implicit Question-answering Text Datasets for Zero-shot Learning
DOI:
https://doi.org/10.33022/ijcs.v15i3.5112Abstract
The use of Large Language Models (LLMs) for task specific has gained more popularity for resourced languages like English due to availability of text data. There are no readily available datasets to fine tune for task specific in low resourced languages like isiNdebele. Data scarcity remains the main obstacle for many low resourced languages as it hinders on the development of language models. In this study we proposed the creation of two isiNdebele Question-answering (QA) text datasets, for explicit and implicit QA. Forty-five matric question papers from the South African Department of Basic Education website were downloaded. A data verification process was performed using human-in-the-loop technique to validate the created datasets. Further data augmentation was performed to increase the implicit dataset from 3012 to 36560 context-question-answer triplets. The augmented implicit dataset was then used to fine-tune the mT5 model on a zero-shot. The two datasets were accepted with a 98.33% acceptance percentage by the participants. The mT5 model performed exceptionally with ROUGE-L and BERTScore F1 of 0,992 and 0,999 respectively. The model made 94.7% accurate predictions with a perplexity score of 1,157. The results indicate that the multilingual model and transfer learning have great potential of dealing with low resourced languages in such as QA.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Promise Malatji, Thipe Modipa

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.





