Handwritten Text Extraction and Digitization
DOI:
https://doi.org/10.47392/10.47392/IRJAEH.2026.0640Keywords:
Handwritten Text Extraction, Structured Information Extraction (SIE), Vision-Language Model, Transformer Architecture, Masked Detection Modelling (MDM), Optical Character Recognition (OCR), Model Performance Analysis, Multi-Modal LearningAbstract
Extracting structured information from visually rich documents remains a complex task due to variations in layout, text alignment, and reading order. Traditional methods based on IOB tagging or graph decoding often struggle with irregular text sequences and the computational burden of large relational graphs. This paper introduces a novel anchor-based approach that redefines entity representation and association for structured information extraction. The proposed model, named Hwte, integrates visual and linguistic features through a multi-modal transformer architecture that jointly detects entities and their relationships. A new pre-training objective, Masked Detection Modelling (MDM), is introduced to enhance the model’s ability to predict both textual and spatial information simultaneously. Experimental evaluations on benchmark datasets demonstrate that the proposed method achieves superior accuracy and robustness compared to existing solutions, highlighting its effectiveness for real-world document understanding tasks.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Research Journal on Advanced Engineering Hub (IRJAEH)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
.