Tamil Palm-Leaf Manuscripts Restoration Using a Transformer–CNN Model

Authors

  • Baghavathi Priya Sankaralingam Department of Computer Science and Engineering, Amrita Vishwa Vidyapeetham-Chennai, India https://orcid.org/0000-0001-9168-7810
  • Krithikha Sanju Saravanan Department of Information Technology, Sri Sivasubramaniya Nadar College of Engineering-Chennai, India https://orcid.org/0000-0003-0120-3355
  • Dhanush Kumar Vijayakumar Department of Information Technology, Sri Sivasubramaniya Nadar College of Engineering-Chennai, India
  • Veda Chatiyode Department of Computer Science and Engineering, Amrita Vishwa Vidyapeetham-Chennai, India

DOI:

https://doi.org/10.47852/bonviewJCCE62028671

Keywords:

Transformer–CNN model, Vision Transformer (ViT), image restoration, synthetic data generation, Tamil palm-leaf manuscripts

Abstract

Palm-leaf manuscripts written in ancient Tamil represent invaluable cultural heritage but have suffered significant degradation due to physical damage, environmental exposure, and biological factors. These challenges, including cracks, ink fading, and text obliteration, hinder scholarly analysis and preservation efforts. Deep learning offers promising solutions for automated restoration; however, its effectiveness is constrained by the scarcity of large, annotated datasets tailored to Tamil palm-leaf degradation patterns. To address this limitation, this research proposes a scalable methodology for generating a synthetic dataset of degraded manuscript images. A custom Python-based pipeline simulates realistic deterioration using techniques such as random masking, structural crack generation, and linear occlusions, implemented with OpenCV and NumPy. Additional variability is introduced through data augmentation, including brightness adjustment, rotation, and scaling. Each degraded image is paired with its corresponding clean version to form a high-quality training dataset. This dataset is used to train a hybrid Transformer-Convolutional Neural Network (CNN) model, combining a Vision Transformer encoder with a lightweight CNN decoder. The proposed approach significantly enhances restoration performance, achieving a Peak Signal-to-Noise Ratio of 29.85 dB and a Structural Similarity Index of 0.887. These results demonstrate the model’s effectiveness in reconstructing degraded Tamil manuscript images, thereby improving their readability, accessibility, and long-term preservation.



Received: 3 December 2025 | Revised: 8 June 2026 | Accepted: 9 July 2026



Conflicts of Interest

The authors declare that they have no conflicts of interest to this work.



Data Availability Statement

The data that support the findings of this study are openly available in Mendeley Data at https://doi.org/10.17632/ssycxg2snx.1, reference number [21].



Author Contribution Statement

Baghavathi Priya Sankaralingam: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Resources, Writing – review & editing, Supervision, Project administration. Krithikha Sanju Saravanan: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Resources, Writing – review & editing, Supervision, Project administration. Dhanush Kumar Vijayakumar: Methodology, Software, Formal analysis, Resources, Data curation, Writing – original draft, Writing– review & editing, Visualization. Veda Chatiyode: Formal analysis, Resources, Data curation, Writing – original draft, Writing – review & editing, Visualization.

Downloads

Published

2026-08-21

Issue

Section

Research Articles

How to Cite

Sankaralingam, B. P., Saravanan, K. S., Vijayakumar, D. K., & Chatiyode, V. (2026). Tamil Palm-Leaf Manuscripts Restoration Using a Transformer–CNN Model. Journal of Computational and Cognitive Engineering. https://doi.org/10.47852/bonviewJCCE62028671