Automating Artl@s - extracting data from exhibition catalogues

Topalov, Barbara
Ecole Du Louvre (France)
topalov.barbara@gmail.com

Gabay, Simon
UniNE/UniGE (Switzerland)
simon.gabay@unine.ch

Joyeux-Prunel, Béatrice
UniGE (Switzerland)
Beatrice.Joyeux-Prunel@unige.ch

Romary, Laurent
INRIA (France), BBAW (Germany
laurent.romary@inria.fr

Rondeau du Noyer, Lucie
independant scholar
lucie.rondeau.du.noyer@chartes.psl.eu

Table of contents

Databases for art history usually focus on images, and are primarily made by or for museums to display collections or inventories. 1 Other databases do not focus on the work of art itself, but on a corpus about the work of art – such as databases used for provenance research 2 or by collectors monitoring the market. 3 The BasArt database of the Artl@s project 4 belongs to the second category as it aims at recording exhibitions all over the world since the invention of exhibition catalogues in the 17th century (Paris, Salon de l’Académie, 1673). 5 BasArt is essential for art history, not only because it makes available one of the most widely used sources in the discipline, at a global scale and over several centuries, but also because it makes it possible to globalise the horizons of research and to move on to unusual quantitative, spatial, and transnational analyses, without specific knowledge in digital technologies (Joyeux-Prunel / Marcel 2015).

If data has long been entered manually, by simply copying the content that in the exhibition catalogues describes the works exhibited, we are now designing a new workflow, based on GROBID dictionaries technology 6 , to semi-automate the task, that is to say to retrieve the content from pdf digitisation of the catalogues, describe it semantically, in order to transform the XML-TEI output of this new workflow into the structured BasArt postGIS database. As our database is structured to deal with catalogues from very different periods of time, and from very diverse locations and languages, we intend to take advantage of the tasks involved in designing a semi-automated retrieving chain – in particular, the description necessary to train recognition algorithms –, to produce research on the history of the form of exhibition catalogues over time and space, and according to the language. The project’s ambition is also to enable broader developments that will be useful for other projects dealing with comparable resources and corpora – such as auction catalogues, collection directories, or catalogues raisonnés.

1. Exhibition catalogues

In order to semantically describe catalogues, we need to analyse their form. This form has a history, which meets the history of the inventory, whose description is a challenge for anyone who wishes to model their structure, but also try to generalise this modeling to any type of similar documents. The inventorium (inventory) can be considered as an early form and the origin of the catalogue. This list of a set of objects forming a collection was originally built by recording collection items in the very order in which they were stored or displayed (Barbier 2015). That is why in France, as early as the Middle Ages, inventories were primarily used to associate a given object to the place where it was preserved. The inventory is first and foremost a practical tool: it is meant to find and identify an object. Inventories also have a scientific/epistemic interest insofar as the classification of objects echoes the classifications of the natural sciences, especially in the case of curiosities cabinets (Findlen 1994). The practice of classifying allows comparison: "Unlike the inventory, the catalogue organises a set of data and arranges them by object, so as to make them comparable with each other.” (Recht 1996).

With the advent of printing, and the ability to publish inventories for a wider audience, the form of the inventory became semi-industrialised for a few particular types: the catalogue raisonné of an artist’s work, the inventory of a private collection or a museum, the dealer's catalogue, the auction catalogue, the exhibition catalogue. It is with this type of printed data that we believe we can develop a semi-automated pipeline for content retrieval. This process begins with the example of exhibition catalogues.

Generally the content of an exhibition catalogue is quite regular: information about the event exhibition, its title, date, supposed location, its organisers; a list of the exhibitors – first and last name, sometimes completed by biographical detail, a year and place of birth and an address. Most exhibition catalogues complete these lists of artists with a list, for each exhibitor, of the works exhibited (Joyeux-Prunel 2015). Whereas this content is not always represented in the same way in the book-catalogue, it is actually organised according to a limited number of possibilities.

2. A Stable Layout, Over Time and Place

In the documentation we have gathered (c. 1000 catalogues since the 18th c., mainly from Western countries), the most recurrent format displays the following information: name of the artist (here, highlighted in red), information about the artist (blue) and a list of the works exhibited (yellow) with potentially additional information (green).

Catalogue de l'exposition annuelle du musée de Rouen, 1860, p. 19

2.1. Type 1

Salons, yearly municipal exhibitions and academy exhibitions which were held regularly since the 18th century, usually published catalogues with a very stable layout over time:

Catalogue de l’exposition... de Rouen, 1853, p.4

Catalogue de l’exposition... de Rouen, 1862, p. 26

This kind of layout can be found throughout the 19th century, whatever the continent, at least until the 1940s in most cases. Data is structured the same way in most of the exhibition catalogues that have been printed – no matter the time or the place – over this period of time. It is therefore easy to build a single model for several catalogues. If we look at exhibition catalogues published in Nancy in the 1840’s, Paris in the 1920’s, Venice in the 1910’s or São Paulo in the 1950’s, we can observe that they follow a similar pattern, which should facilitate the automation of data extraction.

Catalogue… exposés à Nancy, 1843, p. 3

Société des artistes indépendants, 1921, p. 17

Esposizione Internazionale… di Venezia, 1910

Bienal de São Paulo, 1951, p. 54.

Tentoonstelling van schilder- en andere werken, 1852, p. 12

2.2. Type 2

A secondary structure coexists with the latter one – a structure that first presents the number and the title of the work, followed by the name of its author. The order is relatively opposite to that described above. This second type of layout, following the exhibition order, has been used for the Exhibition of the Royal Academy in London since 1780, but it can be found in other countries ( e.g. Canada). It has been extremely stable over time too.

The Exhibition of the Royal Academy, 1785, p. 1

The Exhibition of the Royal Academy, 1831, p. 6

The Exhibition of the Royal Academy, 1907, p. 8

The Exhibition of the Royal Academy, 1975, p. 10

We assume for the moment that these two types may have corresponded to different areas of cultural influence - on the one hand the influence of the Parisian Salon catalogue model, on the other that of a more Anglo-Saxon model. It is likely that another factor explaining these differences is simply that catalogues by work order assume a more commercial character, where it is the work for sale that is presented, and not the artist. This hypothesis is supported by the presence of Type 2 in many gallery catalogues in the early 20th century.

3. Beyond exhibition catalogues

According to the two types presented supra, entries can be grouped under super-entries, which potentially transmit general properties to entries that they include such as the location of the work in the exhibition (gallery 1, red wing…), or more importantly artistic forms (paintings, sculpture, architecture…).

These two types can be found in similar documents, the structure of which is very similar to exhibition catalogues: bibliographies (Lindemann et al. 2018), dictionaries (Khemakhem et al. 2018b), auction catalogues (Gabay et al. 2020), and phone directories (Khemakhem et al. 2018c).

Catalogue de feu M. de Bruyères Chalabre, 1833, p. 102.

Annuaire des imprimeurs et des libraires, 1841, p. 135 .

Such data has been recently identified as “encyclopedic-like” (Khemakhem et al. 2018a), i.e. entry-based data for which a dedicated retrieval tool has been developed: GROBID dictionaries (Khemakhem 2020). Thanks to the latter, PDF files are automatically structured and annotated in XML-TEI files, which is used as a pivot format to ease encoding refinements via regexes.

The current TEI guidelines 7 propose elements allowing research to encode various kinds of lists (<list>, <listPerson>, <listPlaces>, <listObject>…) and their content (<item>, <person>, <place>, <object>…). However, some scholars have noted that encoding a catalogue as a “mere list” was not sufficient if some of its entries feature analytical or descriptive content (Nelson 2016). Therefore, following the work on dictionaries with TEI-lex 0 (Romary 2018) on the <entry> element, we would like to propose to the TEI Consortium the creation of a <catalogueEntry> element. The behaviour of this TEI markup would be similar to <entry>, transforming <form> and <sense> elements into <catalogueDesc> and <catalogueItem>. A possible encoding would therefore be the one presented infra.

The stakes of this type of classification are both to allow a clearer description of the content, and the possibility of spreadsheet-type retrievals, where the dependencies of a sub-entry on its entry would be maintained (in particular in the repetition of the entry for each line of the sub-entry in the case of an export in csv). This type of processing is ideal for a pivot to databases, more easily handled by users interested in quantitative and cartographic visualisation as is the case for the Artl@s project.

Appendix A

Bibliography
  1. Barbier, Frédéric / Dubois, Thierry /Sordet, Yann (2015): De l’argile au nuage, une archéologie des catalogues : IIe millénaire av. J. C. - XXIe siècle. Paris / Genève: Bibliothèque Mazarine / Bibliothèque de Genève / Éditions des Cendres.
  2. Findlen, Paula (1994): Possessing Nature: Museums, Collecting, and Scientific Culture in Early Modern Italy. Berkeley / Los Angeles: University of California Press.
  3. Gabay, Simon / Rondeau Du Noyer, Lucie / Khemakhem, Mohamed (2020): “Selling autograph manuscripts in 19th c. Paris: digitising the Revue des Autographes”, in: IX Convegno AIUCD, AIUCD, Milan, Italy, Jan 2020.
  4. Giffrey, Jules (1869-1872): Collection des livrets des anciennes expositions depuis 1673 jusqu'en 1800. Paris: Liepmann Sohn & Dufour.
  5. Joyeux-Prunel, Béatrice / Marcel, Olivier (2015): “Exhibition Catalogues in the Globalization of Art. A Source for Social and Spatial Art History”, in: Artl@s Bulletin 4, 2: Article 8.
  6. Khemakhem, Mohamen (2020): Standard-based Lexical Models for Automatically Structured Dictionaries. PhD thesis, INRIA.
  7. Khemakhem, Mohamed / Romary, Laurent /Gabay, Simon / Bohbot, Hervé / Frontini, Francesca (2018a): “Automatically Encoding Encyclopedic-like Resources in TEI”, in: The annual TEI Conference and Members Meeting, Tokyo, Japan , Sep 2018.
  8. Khemakhem, Mohamed / Herold, Axel / Romary, Laurent (2018b): “Enhancing Usability for Automatically Structuring Digitised Dictionaries”, in: GLOBALEX workshop at LREC 2018, Miyazaki, Japan, May 2018.
  9. Khemakhem, Mohamed / Brando, Carmen / Romary, Laurent / Mélanie-Becquet, Frédérique / Pinol, Jean-Luc (2018c): “Fueling Time Machine: Information Extraction from Retro-Digitised Address Directories”, in: JADH2018 "Leveraging Open Data"8, Tokyo, Japan, Sep 2018.
  10. Lindemann, David / Khemakhem, Mohamed (2018): “Retro-digitizing and Automatically Structuring a Large Bibliography Collection”, in: European Association for Digital Humanities (EADH) Conference, Galway, Ireland, December 2018.
  11. Nelson, Brent (2017): “Curating Object-Oriented Collections Using the TEI”, in: Journal of the Text Encoding Initiative 9.
  12. Recht, Roland (1996): “La Mise en ordre : note sur l’histoire du catalogue”, in: Les Cahiers du Musée National d’Art Moderne 56/57: 20-35.
  13. Romary, Laurent / Tasovac, Toma (2018): “TEI Lex-0: A Target Format for TEI-Encoded Dictionaries and Lexical Resources”, in: The annual TEI Conference and Members Meeting, Tokyo, Japan, Sep 2018.
Notes
1.
2.
3.
4.
5.
According to Giffrey (1869).
6.
7.