AI-DRIVEN TEXT AND AUDIO SYNTHESIS FOR LOW-RESOURCE LANGUAGE DATASETS

Authors

DOI:

https://doi.org/10.26577/jpcsit4220268

Keywords:

low-resource languages, methodology for creating datasets, parallel text corpora, AI systems, audio synthesis

Abstract

This paper presents an AI-driven methodology for constructing parallel text corpora and synthesizing parallel audio datasets for low-resource Turkic language pairs, namely Kazakh-Kyrgyz and Kazakh-Uzbek. The scarcity of high-quality linguistic and speech resources for these languages poses significant challenges for the development of neural machine translation and automatic speech recognition systems. The paper proposes and validates a methodology for parallel dataset construction driven by artificial intelligence, in which the selection of a translation system is guided by a set of criteria encompassing free accessibility, translation quality, and processing efficiency, assessed through experiments on a 2000-sentence test set.  Based on this evaluation, large-scale parallel corpora for the Kazakh–Kyrgyz and Kazakh–Uzbek language pairs were generated using the selected AI system. A subsequent manual error analysis revealed that approximately 3.7% (Kazakh–Kyrgyz) and 4.3% (Kazakh–Uzbek) of the translations contained inaccuracies, indicating the need for post-editing and further model refinement. Audio synthesis experiments using MMS-TTS and TurkicTTS systems demonstrated that synthetic speech of near-natural quality can be generated for both languages, with NISQA scores reaching up to 4.44 for Uzbek. The findings confirm that the proposed methodology provides a practical and scalable foundation for expanding linguistic resources for low-resource Turkic languages and supporting further research in machine translation and speech processing.

Downloads

Download data is not yet available.

Author Biographies

  • Aidana Karibayeva, Al-Farabi Kazakh National University, Almaty, Kazakhstan

    Aidana Karibayeva, PhD, is an Acting Associate Professor at the Department of Information Systems of al-Farabi Kazakh National University (Almaty, Kazakhstan; karibayeva.aidana@kaznu.edu.kz). Dr. Karibayeva has more than 13 years of experience in artificial intelligence and machine learning. Her research interests include neural networks, deep learning applications, data mining, and speech recognition.

  • Balzhan Abduali, Al-Farabi Kazakh National University, Almaty, Kazakhstan

    Balzhan Abduali is  Master of Engineering Science, Senior Lecturer at the Department of Artificial Intelligence and Big Data of al-Farabi Kazakh National University (Almaty, Kazakhstan; abduali.balzhan@kaznu.kz). She has more than 12 years of experience in machine learning. Her research interests include artificial intelligence, machine learning, and data mining.

Downloads

Published

2026-06-19

How to Cite

AI-DRIVEN TEXT AND AUDIO SYNTHESIS FOR LOW-RESOURCE LANGUAGE DATASETS. (2026). Journal of Problems in Computer Science and Information Technologies, 4(2), 74-85. https://doi.org/10.26577/jpcsit4220268