Artificial intelligence multilingual image-to-speech for accessibility and text recognition

Rosalina Rosalina

Hasanul Fahmi

Genta Sahuri

International Journal of Artificial Intelligence

Artificial intelligence multilingual image-to-speech for accessibility and text recognition

Abstract

The primary challenge for visually impaired and illiterate individuals is accessing and understanding visual content, which hinders their ability to navigate environments and engage with text-based information. This research addresses this problem by implementing an artificial intelligence (AI)-powered multilingual image-to-speech technology that converts text from images into audio descriptions. The system combines optical character recognition (OCR) and text-to-speech (TTS) synthesis, using natural language processing (NLP) and digital signal processing (DSP) to generate spoken outputs in various languages. Tested for accuracy, the system demonstrated high precision, recall, and an average accuracy rate of 0.976, proving its effectiveness in real-world applications. This technology enhances accessibility, significantly improving the quality of life for visually impaired individuals and offering scalable solutions for illiterate populations. The results also provide insights for refining OCR accuracy and expanding multilingual support.

Cite

Full View

DOI

10.11591/ijai.v14.i3.pp1743-1751

ISSN Information

2089-4872

Pages

1743-1751

More Information

Volume 14

Issue 3

Publish at 2025-06-01

Discover Our Library

Embark on a journey through our expansive collection of articles and let curiosity lead your path to innovation.

Explore Now