Modelling Historical Information with Structured Assertion Records

Baillie, James
University of Vienna, Austria
james.baillie@univie.ac.at

Andrews, Tara
University of Vienna, Austria
tara.andrews@univie.ac.at

Romanov, Maxim
University of Vienna, Austria
maxim.romanov@univie.ac.at

Knox, Daniel
University of Vienna, Austria
daniel.knox@univie.ac.at

Vargha, Maria
Charles University, Prague
maria.vargha@univie.ac.at

This paper will present an experimental approach to representing historical data in digital format, in particular in situations where there are multiple and contested readings of that information, via use of a five-point ‘pentile’ data model based on historical assertions.

One of the difficulties inherent in representations of historical data, especially in prosopographical or event modelling contexts, has been the lack of capacity to effectively capture historical argumentation. This tends to result in approaches that acknowledge the construction of a dataset itself as an argument but cannot, in so doing, easily present alternative arguments (cf. Arguing with Digital History working group 2017) or alternatively, approaches that try to reduce to a minimum the presentation of historical argumentation within the dataset, without disambiguation and clarification of the linked sections of primary source material. The factoid model commonly used for prosopographical databases is a key example of this type, with its explicit assumption that historians will, if they wish, provide further layers of analysis themselves (Bradley et al. 2015: 30).

However, reduced models of this sort have proven dissatisfying for historians. For one thing, it is already widely accepted that there is no such thing as a “neutral” dataset (cf. Drucker 2011; Posner 2016). The inherent needs of categorisation require historical judgement which may be open to discussion and challenge, especially in fraught areas like the meanings of historical identities. The models used often have limited capacity for showcasing and presenting the argumentation and historical thought involved in these necessary decisions, which at times weakens their utility for subsequent users.

Furthermore, a dataset of the kind discussed above has limitations on its use cases. If conflicting source data is not disambiguated, it cannot effectively be mapped, or looked at through comparative visualisations: the results of doing so would be nonsensical, as there would be nothing to stop individuals being in multiple places at once, or both dead and alive simultaneously, where sources conflict on those points. It would take a bold scholar to try and argue, for example, that the Georgian Kartlis Tskhovreba’s hazy chronology of the late 12th century Byzantine collapse is as good a representation of events as the chronicle of Niketas Choniates, who was an eyewitness to many relevant events. But in an event model using the minimum-argument rule, this would be the implied logic at the data level and thus in any analyses of that data. All this should not be taken to mean that there are no use cases for minimum-argument models: far from it, as these systems have the strong advantages of being comparatively quick to compile and of providing good close text linkages, which makes them useful for indexing functions. There is nonetheless a clear gap in their effective usage, which historians have struggled to adequately fill through adaptations and workarounds within this model.

We are therefore adopting an alternative approach to the problem in the context of the recently-granted ERC project RELEVEN (Re-evaluating the Eleventh Century through Linked Events and Entities). This involves an adaptation of the linked open data triple data approach, in which data is stored with a subject, predicate, and object. In the project we expand this to a five point system: subject, predicate, object, asserter, and source. This turns a specific triple into a referenced assertion, with an explicit distinction between the documentary source of information and the historian who interpreted it. The proposed system thereby provides a simple data structure that attempts to capture historical information, with its inherent sets of disagreements and conflicts, rather than either representing a single abstracted proposition of historical fact or representing a textual reading that makes no attempt to differentiate the respective merits of source propositions. With this pentile expansion of the triple structure, we are no longer constrained by a minimal-argumentation approach to reading the texts. If we are building a model of what historians consider to have happened, we can represent the majority of this information, including the source evidence that exists but is commonly disregarded, in this flexible but simple data format.

It should be noted that the source reference in the pentile need not be in relation to a singular element of a primary source, as for a factoid – for example, the natural reference point for archaeological data might be a report or paper. Indeed there is no reason why secondary argumentation should not provide the source element, especially when particularly close argumentation or discussion of lacunae are necessary to provide the full basis for an assertion. 

Furthermore, by taking into account the nuances of scholarly interpretation, the pentile structure allows data from wide ranging scholarly approaches to be considered side by side. For example, information gathered from numismatic and literary evidence can co-exist in the same data-set and model and inform the reading of each other. One frequent issue in historical argument and models is the extent to which historians and members of allied disciplines with different methodological specialisms develop different perspectives on the same topic which are hard to compare and reconcile; with the pentile structure these assertions can be placed side by side within the same data collection whilst retaining the importance of the provenance of those viewpoints. As such, the pentile structure allows for a more interdisciplinary approach to collation and modelling of data.

As to the potential scope of information that the model can cover, our pilot project suggests that it is well suited to modelling narratives or prosopographical data, but can also be used with a range of other reference points. The presence of discrete and discernible entities is one of the largest requirements, though that does not necessarily mean consistent identification – if a person named Alexios might be Alexios I or his grandson who died in 1142, for example, then this could be modelled as a third Alexios with IsSameAs or a similar predicate connecting him to both the aforementioned prosopons (each assertion connected, naturally, to the scholar who is making this identification.) This approach to the data, both highly granular and linked effectively to its asserters, allows the effective comparison of different historical and prosopographical viewpoints within one dataset in a way that is not provided by other models.

However, less consistently defined entities, and arguments over entity definitions (e.g. what constitutes “paganism” or “heterodoxy”, or “being Roman”) may be harder to capture with this system; these are areas that could be interesting for ongoing research. One historian’s assertion that a certain figure was “Christian”, for example, may rest on different premises to another’s. However, the act of making these comparisons may help crystallise some of these arguments by providing a framework that requires such discussions about comparability - and where the assertion of an identity term or other attribute may not be comparable, the evidence bases for such assertions could potentially be modelled with appropriate expansions of our system, for example by allowing source nodes to be objects of pentiles in their own right in order to build more complex arguments, something we are keen to discuss and develop further.

By formalising these elements of historical readings and interpretations, we ultimately hope to provide a new basis for their examination and comparison, tightening arguments into systems where the precise conflicts of conclusion between different historians can be made visible at a granular level. The system may additionally act as a potential tool for scholars who wish to present large scale historical datasets in a consistent and interchangeable format, and with appropriate predicates will allow for compatible exports in accordance with ontologies such as CIDOC-CRM and Symogih (Le Boeuf et al. 2019; Beretta 2017). It may further allow discussions of current historical analyses of particular periods to be better differentiated from presentations of underlying sources, and the presentation of these in ways that preserve the fundamentals of historical argumentation and allow further debate to take place most fruitfully.

Appendix A

Bibliography
  1. Arguing with Digital History working group (13.11.2017): Digital History and Argument. White paper. Roy Rosenzweig Center for History and New Media <https://rrchnm.org/argument-white-paper/> [15.06.2021].
  2. Bradley, John / Pasin, Michele (2015): ”Factoid-based prosopography and computer ontologies: towards an integrated approach”, in: Digital Scholarship in the Humanities 30, 1: 86-97.
  3. Beretta, Francesco (2017): L’interopérabilité des données historiques et la question du modèle: l’ontologie du projet SyMoGIH. Nanterre: Presses universitaires de Paris Nanterre <https://halshs.archives-ouvertes.fr/halshs-01559816> [15.06.2021].
  4. Drucker, Johanna (2011): “Humanities Approaches to Graphical Display”, in: Digital Humanities Quarterly 5, 1 <http://www.digitalhumanities.org/dhq/vol/5/1/000091/000091.html> [15.06.2021].
  5. Aalberg, Trond / Balzer, Detlev / Bekiari, Chryssoula / Bountouri, Lina / Crofts, Nick / Doerr, Martin / Dunsire, Gordon / Görz, Günther / Gill, Tony / Hagedorn-Saupe, Monika / Hiebel, Gerald / Holmen, Jon / Inkari, Juha / Iorizzo, Dolores / Kotipelto, Juha / Krause, Siegfried / Lampe, Karl Heinz / Lamsfus, Carlos / Le Bœuf, Patrick / Lindenthal, Jutta / Nyman, Mika / Ore, Christian Emil / Rold, Lene / Riva, Pat / Smiraglia, Richard / Stein, Regine / Stiff, Matthew / Žumer, Maja (2019): Definition of the CIDOC Conceptual Reference Model version 6.2.7 <http://www.cidoc-crm.org/Version/version-6.2.7> [15.06.2021].
  6. Posner, Miriam (2016): “What’s Next: The Radical, Unrealized Potential of Digital Humanities”, in: Gold, Matthew K. / Klein, Lauren F. (eds.): Debates in the Digital Humanities. Minneapolis: University of Minnesota Press 32-41 <https://dhdebates.gc.cuny.edu/read/untitled/section/a22aca14-0eb0-4cc6-a622-6fee9428a357> [15.06.2021].