Index based Statistical Analysis of Large Text Corpora

Project Information

ISALTEC

Grant agreement ID: 625160

Project closed

Start date 1 July 2014

End date 30 June 2016

Funded under

Specific programme "People" implementing the Seventh Framework Programme of the European Community for research, technological development and demonstration activities (2007 to 2013)

Total cost

€ 161 968,80

EU contribution

€ 161 968,80

161 968,80

Coordinated by

LUDWIG-MAXIMILIANS-UNIVERSITAET MUENCHEN
Germany

Objective

The statistical analysis of large text corpora is a fundamental method for gaining insights into the structure of language, e.g. for grammar development, machine translation, terminology and named entity extraction, text correction, semantic text analysis, and others. Progress in these fields helps to improve related applications in information science (search engine technology) and many other text oriented disciplines.
The core contribution of this project is a new methodology aimed at fundamentally improving statistical analysis of large text corpora. A weakness of current methods in corpus analysis is insufficient use of contextual information. Properly understanding the role, function and meaning of a phrase or word (which is important for many applications, e.g. for translation, search, etc.) is often only possible when taking sentence/paragraph contexts into account. We want to develop and study a new representation of corpora which is superior to present formats in three respects. Most importantly, it offers a much better use of contextual information. At the same time it helps to better distinguish between arbitrary and meaningful parts of text and gives hints on how to compose/decompose phrases. With these properties, the new representation gives a basis for fundamentally improving statistical analysis of corpora. The new representation is derived from a special text index structure which gives immediate access to contexts of any size. The index imposes a natural graph structure on the the phrases in the corpus, which implies that interesting graph-based statistical methods can be applied. Further more it can be efficiently constructed and updated in practice.
To practically demonstrate the large potential of the new methodology in NLP we will concentrate on the machine translation where we expect to achieve improved translation methods for words and phrases.

Fields of science (EuroSciVoc)

CORDIS classifies projects with EuroSciVoc, a multilingual taxonomy of fields of science, through a semi-automatic process based on NLP techniques. See: The European Science Vocabulary.

This project has not yet been classified with EuroSciVoc.
Be the first one to suggest relevant scientific fields and help us improve our classification service

Programme(s)

Multi-annual funding programmes that define the EU’s priorities for research and innovation.

FP7-PEOPLE - Specific programme "People" implementing the Seventh Framework Programme of the European Community for research, technological development and demonstration activities (2007 to 2013)

Topic(s)

Calls for proposals are divided into topics. A topic defines a specific subject or area for which applicants can submit proposals. The description of a topic comprises its specific scope and the expected impact of the funded project.

FP7-PEOPLE-2013-IEF - Marie-Curie Action: "Intra-European fellowships for career development"

Call for proposal

Procedure for inviting applicants to submit project proposals, with the aim of receiving EU funding.

FP7-PEOPLE-2013-IEF
See other projects for this call

Funding Scheme

Funding scheme (or “Type of Action”) inside a programme with common features. It specifies: the scope of what is funded; the reimbursement rate; specific evaluation criteria to qualify for funding; and the use of simplified forms of costs like lump sums.

MC-IEF - Intra-European Fellowships (IEF)

Coordinator

LUDWIG-MAXIMILIANS-UNIVERSITAET MUENCHEN

EU contribution

€ 161 968,80

Address

GESCHWISTER SCHOLL PLATZ 1
80539 MUNCHEN
Germany

Region

Bayern Oberbayern München, Kreisfreie Stadt

Activity type

Higher or Secondary Education Establishments

Links

Contact the organisation

Website

Participation in EU R&I programmes

HORIZON collaboration network

Total cost

No data

Objective

Fields of science (EuroSciVoc) CORDIS classifies projects with EuroSciVoc, a multilingual taxonomy of fields of science, through a semi-automatic process based on NLP techniques. See: The European Science Vocabulary.

Programme(s) Multi-annual funding programmes that define the EU’s priorities for research and innovation.

Topic(s) Calls for proposals are divided into topics. A topic defines a specific subject or area for which applicants can submit proposals. The description of a topic comprises its specific scope and the expected impact of the funded project.

Call for proposal Procedure for inviting applicants to submit project proposals, with the aim of receiving EU funding.

Funding Scheme Funding scheme (or “Type of Action”) inside a programme with common features. It specifies: the scope of what is funded; the reimbursement rate; specific evaluation criteria to qualify for funding; and the use of simplified forms of costs like lump sums.

Coordinator

Download Download the content of the page

Fields of science (EuroSciVoc)

CORDIS classifies projects with EuroSciVoc, a multilingual taxonomy of fields of science, through a semi-automatic process based on NLP techniques. See: The European Science Vocabulary.

Programme(s)

Multi-annual funding programmes that define the EU’s priorities for research and innovation.

Topic(s)

Calls for proposals are divided into topics. A topic defines a specific subject or area for which applicants can submit proposals. The description of a topic comprises its specific scope and the expected impact of the funded project.

Call for proposal

Procedure for inviting applicants to submit project proposals, with the aim of receiving EU funding.

Funding Scheme

Funding scheme (or “Type of Action”) inside a programme with common features. It specifies: the scope of what is funded; the reimbursement rate; specific evaluation criteria to qualify for funding; and the use of simplified forms of costs like lump sums.