摘要
An approach to searching user-specified words/phrases in Chinese document images, without the requirements of layout analysis, is proposed in this paper. Bounding boxes of Chinese character images are first determined using connected component analysis. Next, a suitable character from the user-specified word/phrase is chosen as the initial character to search for a matching candidate in the document. Once a matched candidate is found, its adjacent characters in the horizontal and vertical directions are examined for matching with other corresponding characters in the user-specified word/phrase, subject to the constraints of positional relation and size similarity The character matching is done in two stages. The coarse matching is carried out based on the stroke density features. A weighted Hausdorff distance(WHD) is proposed for the second matching phase. Experimental results show that the proposed method can effectively search the user-specified Chinese word/phrase from horizontal or vertical text lines of document images.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 57-60 |
| 页数 | 4 |
| 期刊 | Proceedings - International Conference on Pattern Recognition |
| 卷 | 16 |
| 期 | 3 |
| 出版状态 | 已出版 - 2002 |
| 已对外发布 | 是 |
学术指纹
探究 'Word spotting in Chinese document images without layout analysis' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver