Multim. Tools Appl. | 2021

An adaptive text-line extraction algorithm for printed Arabic documents with diacritics

Abstract

The performance of document text recognition depends on text line segmentation algorithms, which heavily relies on the type of language, author’s writing style, pen type, and document quality. In this paper, we present a novel unsupervised text-line segmentation algorithm for printed Arabic documents with and without diacritics. The presented approach employs a projection profile along with connected components in an iterative manner to detect text-lines. The primary benefits of the presented algorithm are (i) it is not threshold dependent, (ii) it is not required a training phase for threshold selection, and (iii) it is robust towards page rotation, font type, size, and style variation for both with and without diacritics documents. The extensive computational simulations on manually collected dataset prove the efficiency of the proposed scheme compared with several baseline and states of the art methods, including, Voronoi, X-Y Cut, Docstrum, Smearing and Seam-carving methods. Computational time analysis also presented.

Volume 80

Multim. Tools Appl. | 2021

An adaptive text-line extraction algorithm for printed Arabic documents with diacritics

Abstract

Volume 80

Pages 2177-2204

DOI 10.1007/s11042-020-09737-1

Language English

Journal Multim. Tools Appl.

Full Text