Get 20M+ Full-Text Papers For Less Than $1.50/day. Start a 14-Day Trial for You or Your Team.

Learn More →

Heuristics for identification of bibliographic elements from title pages

Heuristics for identification of bibliographic elements from title pages This paper presents a methodology for automatic identification of bibliographic data elements from the title pages of books. Also enumerates the various steps like scanning the title pages, running optical character recognition (OCR) software, generating HTML files out of title pages and applying heuristics to identify the bibliographic data elements. Much of the paper deals with the surveys undertaken to analyze the characteristics of various bibliographic descriptive elements like title, author, publisher and other elements. The first survey deals with the sequence of the bibliographic data in the title pages. The second survey deals with the font size, font type and the proximity of each bibliographic element on the title pages. The survey results are then used to develop heuristics, in order to develop a rule‐based expert system which can identify the bibliographic elements on the title pages. The results of the system are presented, along with problems encountered. http://www.deepdyve.com/assets/images/DeepDyve-Logo-lg.png Library Hi Tech Emerald Publishing

Heuristics for identification of bibliographic elements from title pages

Library Hi Tech , Volume 22 (4): 8 – Dec 1, 2004

Loading next page...
 
/lp/emerald-publishing/heuristics-for-identification-of-bibliographic-elements-from-title-XwRmjElR7o
Publisher
Emerald Publishing
Copyright
Copyright © 2004 Emerald Group Publishing Limited. All rights reserved.
ISSN
0737-8831
DOI
10.1108/07378830410570494
Publisher site
See Article on Publisher Site

Abstract

This paper presents a methodology for automatic identification of bibliographic data elements from the title pages of books. Also enumerates the various steps like scanning the title pages, running optical character recognition (OCR) software, generating HTML files out of title pages and applying heuristics to identify the bibliographic data elements. Much of the paper deals with the surveys undertaken to analyze the characteristics of various bibliographic descriptive elements like title, author, publisher and other elements. The first survey deals with the sequence of the bibliographic data in the title pages. The second survey deals with the font size, font type and the proximity of each bibliographic element on the title pages. The survey results are then used to develop heuristics, in order to develop a rule‐based expert system which can identify the bibliographic elements on the title pages. The results of the system are presented, along with problems encountered.

Journal

Library Hi TechEmerald Publishing

Published: Dec 1, 2004

Keywords: Bibliographic systems; Data handling; Cataloguing; Classification schemes; Information operations

References