Application of Deep Learning in Khmer Invoice Document Image Analysis: A Comparative Study of CNN, CNN-RNN Hybrid, and Transformer Models

Authors

Sidavid Sin
Paragon International University, Phnom Penh, Cambodia

Synopsis

This paper explores the application of deep learning (DL) techniques to the document image analysis of Khmer invoices, a niche but increasingly important area within the field of optical character recognition (OCR). Given the unique challenges posed by the Khmer script and the complexity of invoice layouts, this study reviews three pioneering DL approaches to understand their strengths and limitations in addressing these challenges. The first approach employs Convolutional Neural Networks (CNNs) for their superior text recognition capabilities, particularly adept at deciphering the intricate patterns of Khmer characters. The second method integrates CNNs with Recurrent Neural Networks (RNNs), leveraging the latter's proficiency in contextual and sequential data analysis to enhance layout understanding and text extraction. The third strategy introduces Transformer-based models, capitalizing on their self-attention mechanisms to foster a deeper comprehension of document context and relationships. Through a comparative analysis, this paper delineates the advantages of each model, such as the CNN's accuracy in character recognition, the CNN-RNN hybrid's effective layout and text analysis, and the Transformer's comprehensive document understanding. Conversely, it also discusses their disadvantages, including issues related to computational demands, training data requirements, and adaptability to diverse invoice formats. Concluding with a synthesis of findings, the study proposes a hybrid model that combines the robust feature extraction of CNNs with the contextual awareness of Transformers as a promising solution for Khmer invoice processing. This approach aims to mitigate the identified limitations and suggests directions for future research, emphasizing the need for lightweight architectures and the development of comprehensive benchmark datasets.

null
Published
March 12, 2025
Online ISSN
2582-3922