Malínek, Vojtěch
Institute of Czech Literature, Czech Academy of Sciences, Czech Republic
malinek@ucl.cas.cz
Umerle, Tomasz
Institute of Literary Research, Polish Academy of Sciences, Poland
tomasz.umerle@ibl.waw.pl
Literary studies, because of their inherent interest in the analysis of texts, deeply depend on the existence of high-quality subject bibliographies. Subject bibliographies assess resources’ relevance for a literary domain – regardless of their accessibility – and as such they can provide invaluable information for users (especially scholars), and serve as necessary tools for critical analysis of available sources for any literary research, in particular for the data-driven studies that aim to perform empirical research on representative datasets.
From this perspective, bibliography is no longer solely perceived only as a listing of references about a specific subject (Marco 1989) aimed to help in discovering and organizing literature for research, but a fundamental tool for research on book history, “sociology of texts” (McKenzie 1999), bibliometrics, cultural analytics, and other data-driven approaches to literary research (Moretti 2005, Jockers 2013, Bode 2018). Within the last few years the systematic research on bibliographical data is accelerating and the term “bibliographical data science” appears in the current scientific discourse (Lahti et al. 2019). Bibliographical data serves as a basis for conducting research e.g. on diachronic analysis of the development of various cultural phenomena such as history of book-printing in Scandinavian countries in early modern era (Tolonen et al. 2019), analysis of the transformation of Polish literary system after 1989 (Maryl 2021), or analysis of the global circulation of the translations of literary works (Vimr 2019). Another research orientation is connected with the creation of various tools for bibliographical data analysis (Péter et al. 2020), or the assessment of data quality (Király 2019).
Our poster presents the Literary Bibliography Research Infrastructure (LiBRI) initiative which aims to become an international literary bibliography that would support such research. LiBRI is a joint initiative of Czech Literary Bibliography (Česká literární bibliografie – CLB) and Polish Literary Bibliography (Polska bibliografia literacka – PBL), both being long-term existing bibliographical infrastructures operating within the national Academies of Sciences.
Our aim is to sum-up and present the results of the first phase of the LiBRI implementation. Its main goal was to integrate both databases and provide one access point to these two robust and heterogeneous datasets of bibliographical information. Both of them have different methodological backgrounds, and description rules, in particular with respect to the subject headings systems (proprietary ones, national ones, MARC21 based standards). To build a joint service for both resources it was first and foremost necessary to implement a joint data model. It was a challenge, as PBL has been adhering to its own proprietary data model systematically developed since the 1950´s, while CLB has undergone various changes of its data structure since the 1990´s and lately is using the system of Czech national authorities for better interoperability of its datasets.
The main goal for the second phase of LiBRI implementation (2021 –2022) is data unification and data enrichment. The LiBRI team has to analyse how to effectively interconnect two large datasets differing in language and how to foster their interoperability and interconnectedness. The main challenges here are: (a) how to unify the two different subject description systems in two different languages; (b) which language to choose as a primary one (English or vernacular one) and to what extent it is possible and effective to offer multilingual metadata; (c) which internationally recognized persistent identifiers system for interconnection of subject headings to choose.
For these purposes, the internationally recognized system of Library of Congress subject headings was chosen which will be further enriched with the specific literary science terms where needed. The data should be presented defaultly in English. Simultaneously it will contain VIAF identifiers as a primary system of persistent identifiers.
By the end of 2020 a beta version of LiBRI service has been made publicly available at libri.ucl.cas.cz with the open source VuFind discovery system which has been chosen as a technological platform for joint interface. The official version should be launched at literarybibliography.eu during the autumn 2021. It offers an unified access to approximately 4,000,000 bibliographical records on literature in Czech lands and in Poland, chronologically covering the period since the end of 18th century up to the latest days.