The Corpus of German Song Lyrics: Recent Developments and Interdisciplinary Potential

Schneider, Roman
Leibniz-Institut für Deutsche Sprache (IDS), Deutschland
schneider@ids-mannheim.de

Popular music and its lyrics have established themselves as integral parts of modern everyday culture. Given the varied situations in which we are surrounded by them - and the mass media presence of its prominent actors - lyrics certainly have an impact on the use of language that should not be underestimated. For example, the hip-hop genre with its largely melody-free rap chants has evolved from an originally subcultural phenomenon to a global form of articulation for (mostly) younger people (Androutsopoulos 2003), and therefore has become a legitimate subject of contemporary digital humanities studies, linguistics, and even language education (Werner / Tegge 2020). For demonstrable reasons related to the limited quantitative data available, extensive empirical investigations that rely on large textual databases have so far focused primarily on Anglo-American characteristics.

The multilayer-annotated corpus of German lyrics (Schneider 2020) contributes to closing this data gap in the elusive continuum between standard and non-standard varieties as well as between written and spoken language. It allows multivariate, statistically based, and reproducible analyses of lexical, pragmatic or morphological phenomena, and comprises both thematic and author-specific archives. These include collected works of artists, covering a wide range from hip-hop to rock-pop to political singer-songwriters (Schneider et al. 2021), as well as the most successful German-language chart songs over the past five decades. In compliance with legal requirements, selected XML-TEI-annotated archives are freely available for academic use. They offer machine-driven annotations of constituent structures and manual annotations of lemmata, part-of-speech tags, named entities, neologisms, and rhyme types. A dedicated website ( www.songkorpus.de) features a corpus front end with a wide range of search functions, exploration functions on various description levels, and live visualizations; cf. Figure 1.

Figure 1: The Songkorpus Website

We introduce substantial new corpus features, e.g. additional archives - increasing the number of corpus tokens to more than three million - and the calculation of word embeddings, using the GloVe unsupervised learning algorithm (Pennington et al. 2014). A downloadable comprehensive collocation dataset now contains all bi-, tri-, tetra-, penta- and hexagrams of the corpus lyrics, with nine measures of association strength each, including widespread parameters like pointwise mutual information or log-likelihood as well as innovative approaches like lexical gravity or cost reduction (cf. Gries 2015). This is intended as a starting point for some methodical evaluation of the reliability and meaningfulness of statistical measures with respect to rather rare collocations, that are nevertheless worth investigating because of their unusal or creative usage. Moreover, a measure of context similarity, based on five-word context windows and multi-dimensional vector space models, allows for the examination of multiword expressions (MWE) that do not quite fit into their contexts. Typical examples are idioms or metaphors; both seem prominent in song lyrics (cf. Werner 2012). This new measure comes in two variants: one variant (CO_VEC_LEX) only takes content words (nouns, main verbs, adjectives and adverbs) into account, the second variant (CO_VEC) also counts in function words. All measures can be fruitfully applied for empirical tasks, e.g. the data-driven identification of idiomatic multiword-expressions (cf. Amin et al. 2021).

We expect the corpus to provide a wide range of additional starting points for future transdisciplinary projects. Besides (computational) linguistics and literary studies, benefiting research areas can be located in the broad spectrum of cultural studies, language didactics, or media sciences.

Appendix A

Bibliographie
  1. Amin, Miriam / Fankhauser, Peter / Kupietz, Marc / Schneider, Roman (2021): "Data-driven Identification of Idioms in Song Lyrics", in: Association for Computational Linguistics (ed.): Proceedings of the 17th Workshop on Multiword Expressions (MWE 2021). Bangkok, Thailand (online), August 6, 2021 13-22 <https://aclanthology.org/2021.mwe-1.3.pdf> [21.08.2021].
  2. Androutsopoulos, Jannis (ed.) (2003): HipHop: Globale Kultur – lokale Praktiken. Bielefeld: transcript.
  3. Gries, Stefan Th. (2015): "Quantitative designs and statistical techniques", in: Biber, Douglas / Reppen, Randi (eds.): The Cambridge Handbook of English Corpus Linguistics. Cambridge: University Press 50-71.
  4. Pennington, Jeffrey / Socher, Richard / Manning, Christopher D. (2014): "GloVe: Global Vectors for Word Representation", in: Association for Computational Linguistics (ed.): Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, October 2014 1532-1543 <https://aclanthology.org/D14-1162.pdf> [21.08.2021].
  5. Schneider, Roman (2020): "A Corpus Linguistic Perspective on Contemporary German Pop Lyrics with the Multi-Layer Annotated Songkorpus", in: ELRA (ed.): Proceedings of The 12th Language Resources and Evaluation Conference (LREC). Marseille, France, May 2020 835-841 <https://aclanthology.org/2020.lrec-1.105.pdf> [21.08.2021].
  6. Schneider, Roman / Hansen, Sandra / Lang, Christian (2021): "Das Vokabular von Songtexten im gesellschaftlichen Kontext – ein diachron-empirischer Beitrag", in: Sprache in Politik und Gesellschaft: Perspektiven und Zugänge. Jahrbuch 2021 des Instituts für Deutsche Sprache. Berlin / Boston: De Gruyter 295-304.
  7. Werner, Valentin (2012): "Love is all around: A corpus-based study of pop lyrics", in: Corpora 7, 1: 19-50.
  8. Werner, Valentin / Tegge, Friederike (eds.) (2020): Pop Culture in Language Education. Theory, Research, Practice. London: Routledge.