During the first eighteen months we focussed on presenting the concepts of TDM, the legal and technological barriers, and the idea of the OpenMinTeD infrastructure and platform.
One of the main goals is to deliver supporting services to various stakeholders facilitating the adoption of the OMTD infrastructure, especially from the perspective of training and supporting the community of interest. To this the progress of this period summarizes in the following achievements:
• Establishment and deployment of the community training and support services Knowledge Base (KB).
• Development of 3 supportive taxonomies integrated with the KB, defining the scope of the community training and serving to classify training materials. These include the “Text and Data Mining”, “TDM Methods” and “Research Workflow” taxonomies.
• The identification and the upload of seed general TDM training material to the KB.
• The development of a plan/schedule for the delivery of project specific self-learning as well as moderated (e.g. webinar, f2f training) training activities utilising the KB.
• The conducting of a survey of publishers and the subsequent analysis of information on technical issues around machine access to research publications.
• Working on the development of the first draft of the support KB collecting information on machine access to research publications from different publishers.
• Work in progress on the preparation of an initial FAQ style consulting service, to be served out of the KB, addressing common questions regarding legal limitations for TDM.
Community driven requirements and evaluation
• The methodology for collecting requirements relevant to TDM from research communities that have been identified as potential end users of the OpenMinTeD project services was defined.
• Analysis and harmonization of requirements collected from the different research communities by identifying commonalities and differences. Moreover, the most representative stakeholder personas were identified, namely the “Text-mining researcher”, the “Researcher”, the “Data Curator” and the “Technical Manager”. Finally, concrete requirements were collected for both for the OpenMinTeD platform and the foreseen use case applications.
• The functional specifications of the OpenMinTed Platform were defined in order to satisfy the user’s needs, and therefore make the OpenMinTed platform more appealing in various scenarios and applications.
Interoperability
• One of the first activities was the compilation of a report on the landscape of tools and standards in the domain of text and data mining and natural language processing.
• A methodology to create interoperability specification consisting of requirements and how to address them. The first version of the OMTD-SHARE metadata schema was produced by the metadata working group in collaboration with all the other working groups.
• The work on the data interoperability toolkit has focused on defining a modular architecture of connectors to non-standard content provider systems, including ingestion components, harvesting components for each system, converters to the OpenMinTeD schema.
Platform Design and Implementation
• The major layers and components of the platform were identified in the overall platform architecture.
• During the first period of the project the main functionalities of the registry were defined, and the service was designed and implemented.
• The user interface of the platform was designed by taking into account the functional specifications, as well as the results of the OMTD-Share internal data model.
Platform Integration
• The basic software engineering infrastructure required by the project and the platform testing methodology have been established.