The Corpus for Idiolectal Research (CIDRE)

Seminck, Olga
Centre national de la recherche scientifique; Lattice (Langues, Textes, Traitements informatiques, Cognition) - UMR 8094
olga.seminck@cri-paris.org

Gambette, Philippe
Université Gustave Eiffel; Lattice (Langues, Textes, Traitements informatiques, Cognition) - UMR 8094
philippe.gambette@univ-eiffel.fr

Legallois, Dominique
Université Sorbonne nouvelle; Lattice (Langues, Textes, Traitements informatiques, Cognition) - UMR 8094
dominique.legallois@sorbonne-nouvelle.fr

Poibeau, Thierry
Centre national de la recherche scientifique; Lattice (Langues, Textes, Traitements informatiques, Cognition) - UMR 8094
thierry.poibeau@ens.psl.eu

Table of contents

1. Abstract

It is well known that the idiolect (the language of an individual) evolves over time. However, there is a lack of quantitative studies on this topic, due to the lack of large corpora (but see Barlow 2013; Mollin 2009; Petré et al. 2019 for a few examples). To study what is specific in an idiolect and how it evolves over a lifetime, we assembled, cleaned and dated the fiction works of 11 very prolific 19th and early 20th century French writers. This resulted in the CIDRE corpus counting 37 million words and over 400 books.

2. Motivation

We want to assemble a longitudinal corpus for stylistics studies, that is:

3. Corpus Assembling

Criteria to select relevant authors:

Programming Scripts:

4. Data

5. Dating of Works

Important contribution of ours : Annotation of Books with year of writing

Table 1: Some examples from the Gréville Corpus

6. Dates of Works

7. Availability and Licenses

8. Acknowledgements

This work was funded in part by the French government under the management of the Agence Nationale de la Recherche as part of the "Investissements d’avenir" program, reference ANR-19-P3IA-0001 (PRAIRIE 3IA Institute).

Appendix A

Bibliography
  1. Barlow, Michael (2013): "Individual differences and usage-based grammar", in: International Journal of Corpus Linguistics 18, 4: 443–478.
  2. Mollin, Sandra (2009): "'I entirely understand' is a Blairism: The methodology of identifying idiolectal collocations", in: International Journal of Corpus Linguistics 14, 3: 367–392.
  3. Petré, Peter /Anthonissen, Lynn / Budts, Sara / Manjavacas, Enrique / Silva, Emma-Louise / Standing, William / Strik, Odile A.O. (2019): "Early Modern Multiloquent Authors (EMMA): Designing a large-scale corpus of individuals’ languages", in: ICAME Journal 43, 1: 83–122.