User Assistance for Resource Discovery and Contextualisation in Close Reading for the Open Library: An API-Driven Distributed Data Research Model

Sugimoto, Go
Donau University Krems (work is done during Austrian Academy of Sciences, Austrian Centre for Digital Humanities)
go.sugimoto@donau-uni.ac.at

Table of contents

1. Introduction

Over the years resource discovery on the web has become a challenge for research users. At the same time, a vast amount of interdisciplinary knowledge may be required to understand and analyse a wide range of digital resources acquired by the discovery process. Digital Humanities (DH) indeed faces those two challenges. In exchange of the advantages of the cross-domain research promoted and exercised in DH (Terras 2016, Isemonger 2018), it becomes eminent that the researchers need more and more assistance to consume information outside their expertise and to efficiently execute research on the web. For example, they may need to read documents in other languages or from other domains requiring new background knowledge, and analyse unfamiliar data obtained from different disciplines.

2. Web Application for Open Library

To address those issues, CAROL (Cross Assistant Reading for Open Library) is developed to help the users to explore a broad spectrum of rich resources publicly available in the Open Library (Figure 1). The Open Library is an open online catalogue of over 20 million edition record and provides access to 1.7 million scanned versions of books [1] . It gives the humanities researchers excellent opportunities, because there are thousands of primary and secondary literatures about any subjects.

Figure 1. CAROL Home page

CAROL is a search engine combining two search functions: a) metadata search to find resources in the library and b) full-text search to look inside a selected resource [2] (Figure 2). It offers a seamless investigation experience to support the research process.

Figure 2. Search results page

In addition, the full-text search triggers an array of API calls to offer user friendly functionalities for the search results. Within each result, CAROL 1) presents a text snippet with the search keywords highlighted, 2) provides a book viewer to show the corresponding page, 3) displays all named entities identified in the snippet with explanatory information (including maps if locations are found), 4) creates a consolidated map to plot all locations found, and 5) translates the snippet into English (in case the language is not English) (Figure 3, 4, 5, 6).

Figure 3. A snippet and book viewer

Figure 4. Recognised entities

Figure 5. Consolidated map

Figure 6. Translation from Italian to English

CAROL opens up more possibilities for them to bravely explore thousands of underexplored valuable resources, without worrying about the languages and domain knowledge. Contextualised and background information enable them to concentrate on their close reading (Jänicke et al. 2015). New insights may be found during their research process.

3. Scopes and Challenges

CAROL has been developed, taking a full benefit of API (Cohen 2005, Tasovac et al. 2016). Two important scopes are considered: data and service independence. Firstly, it does not own and store any data. All datasets are obtained and processed on the fly and simply displayed to the users. There is no maintenance cost and new data is constantly added by the Open Library. Secondly, it is service neutral. It can easily add, delete, and update its services, because APIs are plugged in on demand. For instance, new full-text repositories and data processing functionalities can be added. The latter may include tokenisation, lemmatisation, PoS tagging, as well as dependency parsing and sentiment analysis. As such, CAROL serves as an example of a flexible and portable solution for the de-centralised data integration for DH. On the other hand, the downside is control and slow performance. It is not possible to ensure a perfect service due to the lack of control. The latter is neither a design problem, nor a code optimisation problem. Rather the project proved that the web infrastructure was currently not satisfactory to offer a robust distributed data research, using a series of APIs (Sugimoto 2017).

4. Conclusions

As quantitative methods and distant reading (Moretti 2013) have dominated the headlines of DH research, the close reading is overshadowed to some extent. In the CAROL project, the author attempts to re-emphasise its significance, without compromising the intake of new methodologies for the DH research paradigm. However, CAROL is not meant for technical innovation. It aims to demonstrate a constructing fusion of conventional humanities research and well-established technology, by offering a multi-dimensional support for interdisciplinary studies. Due to its simplicity, the target users are not only the DH researchers, but also the general public. In short this article highlights the potential and dilemma of the distributed data-driven research in the present landscape of web-based DH applications.

5. Notes

[1] https://openlibrary.org/help/faq/about (accessed March 12, 2020).

[2] The Open Library does not offer full-text search for all resources.

Appendix A

Bibliography
  1. Cohen, Dan (2005): Do APIs Have a Place in the Digital Humanities? <http://www.dancohen.org/2005/11/21/do-apis-have-a-place-in-the-digital-humanities/> [19.05.2021].
  2. Internet Archive (ed.) (2021): About Open Library <https://openlibrary.org/help/faq/about#what> [12.03.2020].
  3. Isemonger, Ian (2018): "Digital Humanities and Transdisciplinary Practice: Towards a Rigorous Conversation", in: Transdisciplinary Journal of Engineering & Science 9: 116-138 DOI: https://doi.org/10.22545/2018/00105.
  4. Jänicke, Stefan / Franzini, Greta / Cheema, Muhammad Faisal / Scheuermann, Gerik (2015): On Close and Distant Reading in Digital Humanities: A Survey and Future Challenges DOI: https://doi.org/10.2312/eurovisstar.20151113.
  5. Moretti, Franco (2013): Distant Reading. London: Verso.
  6. Sugimoto, Go (2017): "Who is open data for and why could it be hard to use it in the digital humanities? Federated application programming interfaces for interdisciplinary research", in: International Journal of Metadata, Semantics and Ontologies 12, 4: 204-218 DOI: https://doi.org/10.1504/IJMSO.2017.10014806.
  7. Tasovac, Toma /Barbaresi, Adrien / Clérice, Thibault / Edmond, Jennifer / Ermolaev, Natalia / Garnett, Vicky / Wulfman, Clifford (2016): "APIs in Digital Humanities: The Infrastructural Turn", in: ADHO (ed.): Digital Humanities 2016: Conference Abstracts. Jagiellonian University & Pedagogical University, Kraków 93–96 <http://dh2016.adho.org/abstracts/191> [19.05.2021].
  8. Terras, Melissa (2016): "Being the Other: Interdisciplinary Work in Computational Science and the Humanities", in: Deegan, Marilyn / McCarthy, Willard (eds.): Collaborative Research in the Digital Humanities. London: Routledge 213-230.