Frank, Ingo
Leibniz Institute for East and Southeast European Studies, Germany
frank@ios-regensburg.de
Interdisciplinarity “is typically characterized by integration of information, data, methods, tools, concepts, and/or theories from two or more disciplines or bodies of specialized knowledge. Proactive focusing, blending, and linking of disciplinary inputs foster a more holistic understanding of a question, topic, theme, or problem by individuals or teams” (Klein 2014: 15). In order to support researchers in finding relevant research output to be (re)used in interdisciplinary research, information systems have to provide interdisciplinary perspectives on data and the tools and methods which were used to create the data. Unfortunately, research data repositories are not very well suited to the information needs of interdisciplinary researchers. Their metadata does often not provide methodological information about the creation of research data and the theoretical and disciplinary context in which the data were collected.
Our paper shows how knowledge organization can support interdisciplinary research by enabling detailed descriptions of what methods were applied to create research data. For example, a survey is created by using a questionnaire as a tool or instrument, a coding scheme is used for content analysis, etc. Coding schemes or classification systems used as tools for the method of content analysis may be influenced by theoretical and disciplinary perspectives (see for more detailed examples Franzosi 2009: 33-34).1 Thus, the research question is: How can different disciplinary perspectives on research data be described by fine-grained metadata?
We employ a kind of phenomenon-based knowledge organization system (see Szostak et al. 2016) applied to research data. We designed a DCAT (Data Catalog Vocabulary)2 application profile (Heery / Patel 2000) where the phenomenon is described by subject classification and metadata about the spatial and temporal coverage of the dataset. Different disciplinary classification systems and thesauri can be used in order to describe the object of research: e. g. Iconclass for art history and JEL for economics. The knowledge organization systems are provided in SKOS format by an Apache Jena Fuseki SPARQL triple store and can be discovered through our Skosmos SKOS browser (Suominen et al. 2015) (see fig. 1).

Fig. 1. Screenshot of our Skosmos browser providing disciplinary knowledge organization systems for subject classification and methodological information
The application profile is extended with elements from the DDI-RDF Discovery Vocabulary (Disco) (Bosch et al. 2013). In principle we obtain a faceted classification system for interdisciplinary knowledge organization through metadata fields representing the facets phenomenon, method, theory, and discipline. Our institutional research data repository provides only the DDI Controlled Vocabulary for Mode Of Collection (Jaaskelainen et al. 2010) for the classification of data collection methods at the moment (see fig. 2), but the Taxonomy of Digital Research Activities in the Humanities (TaDiRAH)3 could also be used as classification system for the method facet.

Fig. 2. Screenshot of our research data repository based on the DKAN open data platform
We add provenance information by following the best practice for modeling data lineage4. The Dublin Core provenance property is used to document the data lineage (possibly in narrative form) and the Dublin Core source property is used to provide the bibliographic or archival description of the source(s) used to create the data. This way of modeling provenance information is limited, because there is no explicit representation of activities undertaken to collect data from sources and/or applying specific methods and tools to create data.
Activities could be modeled with the Activity class from the PROVenance Interchange Ontology (PROV-O) (Lebo et al. 2013), but is too general by default. CRMdig5 would be suited to model digitization processes in digital humanities projects, but seems to be too complex and after all too specific for our purposes. The Scholarly Ontology (SO) (Pertsas / Constantopoulos 2017) combines aspects of our faceted classification for methodological information with event-based modeling. Therefore, we apply the activity perspective of SO (see fig. 3), because it provides a convenient point of view to document the research activities taken to create research data in a digital humanities project.

Fig. 3. Diagram of the activity perspective of the Scholarly Ontology (from Pertsas / Constantopoulos 2017)
Allowing metadata curators to use different classification systems or thesauri to describe objects of research and data collection methods enables researchers to retrieve research data from their disciplinary perspectives. Furthermore, it enables interdisciplinary researchers to find and integrate data from other disciplinary contexts into their interdisciplinary research.
The paper also explores and discusses the limits of the different knowledge organization approaches and therefore enables the discussion of further development of knowledge organization for interdisciplinary research data management.
There remain questions of the adequate level of granularity of metadata descriptions: e. g. should the activity of georeferencing a collection of digitized historical maps be documented for the whole collection or for each particular map?
We propose an event-based modeling approach as provided by SO to describe fine-grained methodological and provenance information about the activities carried out to create research data.
Further experiments6 have to show how the approach could be extended to the detailed description of the creation of historical sources (e. g. historical statistics used as research data).7 Even more relevant for research data management would be an extension to model also planned activities.8