Skip to main content
Go to the home page of the European Commission (opens in new window)
English English
CORDIS - EU research results
CORDIS

Digital Bridge: Optical Character Recognition for Early Printed Books in Latin

Objective

This project aims to provide the first viable and accurate solution for digitising early printed books in Latin using Optical Character Recognition. Our basic OCR package will be free and open-source, in order to ensure affordability, longevity, and openness for improvement (three failures of our commercial competitors). Our Company Limited by Guarantee will market costumisation, training, support, and further development tailored to specific collections of books (the standard failure of open-source solutions). Customisation services are essential in our market. Early printed Latin cannot be successfully digitised using standard OCR packages (whether open-source or commercial): these currently have an accuracy of no more than 15%. We plan to modify the open-source Tesseract engine, by training it to account for Latin grammar and early typography: this will increase its accuracy of recognition to about 80%. Customisation tailored to specific collections of books will further improve accuracy to about 95% to 98%.

Our company will address the needs of libraries, digital publishers, researchers, learned societies, and private collectors of early books. Our commercialisation plan is modelled on that of other successful businesses based on open-source software.

The demand for Latin OCR is strong, as publishers and libraries switch to digital publication and storage. From the invention of printing in the Renaissance until well into the 19th century, Latin was the European language of every intellectual discourse: the natural sciences, mathematics, philosophy, theology, law, literary criticism, geography, archaeology, music, medicine. The subsequent shift to using the vernacular languages was a seismic event. We are now experiencing a revolution of similar proportions: the advent of digital publication is bringing opportunities and risks whose outlines are still unclear. This project aims to offer a solid technical bridge between the digital future and the Latin past.

Fields of science (EuroSciVoc)

CORDIS classifies projects with EuroSciVoc, a multilingual taxonomy of fields of science, through a semi-automatic process based on NLP techniques. See: The European Science Vocabulary.

You need to log in or register to use this function

Programme(s)

Multi-annual funding programmes that define the EU’s priorities for research and innovation.

Topic(s)

Calls for proposals are divided into topics. A topic defines a specific subject or area for which applicants can submit proposals. The description of a topic comprises its specific scope and the expected impact of the funded project.

Funding Scheme

Funding scheme (or “Type of Action”) inside a programme with common features. It specifies: the scope of what is funded; the reimbursement rate; specific evaluation criteria to qualify for funding; and the use of simplified forms of costs like lump sums.

ERC-POC - Proof of Concept Grant

See all projects funded under this funding scheme

Call for proposal

Procedure for inviting applicants to submit project proposals, with the aim of receiving EU funding.

(opens in new window) ERC-2014-PoC

See all projects funded under this call

Host institution

UNIVERSITY OF DURHAM
Net EU contribution

Net EU financial contribution. The sum of money that the participant receives, deducted by the EU contribution to its linked third party. It considers the distribution of the EU financial contribution between direct beneficiaries of the project and other types of participants, like third-party participants.

€ 148 178,00
Address
STOCKTON ROAD THE PALATINE CENTRE
DH1 3LE DURHAM
United Kingdom

See on map

Region
North East (England) Tees Valley and Durham Durham CC
Activity type
Higher or Secondary Education Establishments
Links
Total cost

The total costs incurred by this organisation to participate in the project, including direct and indirect costs. This amount is a subset of the overall project budget.

€ 148 178,00

Beneficiaries (1)

My booklet 0 0