Exploring Computer Vision for Film Analysis: A Case Study for Five Canonical Movies

Schmidt, Thomas
Media Informatics Group, University of Regensburg, Germany
thomas.schmidt@ur.de

El-Keilany, Alina
Media Informatics Group, University of Regensburg, Germany
Alina.El-Keilany@stud.uni-regensburg.de

Eger, Johannes
Media Informatics Group, University of Regensburg, Germany
johannes.eger@stud.uni-regensburg.de

Kurek, Sarah
Media Informatics Group, University of Regensburg, Germany
sarah.kurek@stud.uni-regensburg.de

Table of contents

1. Introduction

Quantitative methods have a long tradition in film analysis going back to the predigital era (Salt 1974; Vonderau 2020). Nowadays, multiple projects explore movies via computational methods to investigate colors (Burghardt et al. 2016, 2018; Flueckiger 2017; Kurzhals et al. 2016; Masson et al. 2020; Pause / Walkowski 2018), shot lengths (Baxter et al. 2017; DeLong 2015) or annotation possibilities (Halter et al. 2019; Kuhn et al. 2015; Schmidt / Halbhuber 2020; Schmidt et al. 2020a). Recent research has also led to the definition of the term Distant Viewing (Arnold / Tilton 2019) to describe large-scale digital movie analysis. A lot of the current research is focused on the analysis of text via scripts or subtitles (Byszuk 2020; Holobut et al. 2016; Holubut / Rybicki 2020; Hoyt et al. 2014). However, developments in computer vision have led to novel methods for the image channel of movies and are already applied in computer science to develop recommender systems (Deldioo et al. 2016; Wei et al. 2004) but also in Digital Humanities (DH) to analyze movies (Howanitz et al. 2019; Pustu-Iren et al. 2020; Zaharieva et al. 2012) and other visual media (Schmidt et al. 2020e). We argue that these methods are beneficial for digital film studies and give new perspectives.

We present an exploratory study for the methods: Object detection, emotion recognition, gender- and age-prediction. We apply state-of-the-art models on a subset of frames of five different movies of varied decades and genres. We apply the exploratory research approach defined by Wulff (1998) for traditional film analysis in this study for computational approaches. Our goals are (1) to inspect the benefits and problems of the methods, (2) explore if the methods uncover specific characteristics of the movies and (3) what research questions seem promising to follow in further large-scale studies.

2. Material

We limited the analysis on five movies. Table 1 presents the movies and metadata. For all movies except Avengers, we use a digitally restored version. All movies have a 720x576 resolution, 25 frames per second and 32 bits per sample. We focus on canonical work and Hollywood productions.

Table 1. Movies and metadata.

3. Methods

All analysis was performed in Python 3. We extracted the frames of every movie since all of the applied methods are image-based. However, we take one frame per second of a movie and regard this as the sample of a movie. We decided to employ this approach because using all frames makes the data processing very performance/resource-intensive and we argue that one frame per second offers sufficient information for our first explorations.

To perform the object detection, we use Detectron2 (Wu et al. 2019) which offers state-of-the-art object detection models by Facebook AI Research1. We use a pretrained masked RCCN-model trained on the well-known COCO-Dataset (Lin et al. 2015), which can predict 80 object classes including vehicles, animals, and sports objects. Applying this prediction model on an image, we receive the number of predicted objects, the locations, and the prediction confidence (0-100%). As threshold for the detection, we select 50% which is usually very low but fits our exploratory approach.

Emotion recognition is a sub-field of affective computing (cf. Halbhuber et al. 2019; Hartl et al. 2019; Ortloff et al. 2019; Schmidt et al. 2020c) and is often applied in DH to predict sentiment and emotions from written text (Moßburger et al. 2020; Schmidt / Burghardt 2018; Schmidt, 2019; Schmidt et al. 2019a; Schmidt et al. 2020b). We focus on the image channel of movies and for the emotion prediction we use the Python module FER2 (Goodfellow et al. 2013). The module first performs face detection via a MTCNN Face Detector3 (Zhang et al. 2016) and then predicts the emotion via a convolutional neural network (CNN) trained on over 35,000 images. The model predicts the seven classes anger,disgust,fear,happiness,sadness,surprise and neutral on a scale from 0 to 1. All values sum up to 1 for one face.

We perform gender- and age-prediction via the module py-agender4 which is also a CNN trained on the IMDB-Wiki dataset (Rothe et al. 2018) consisting of over 500,000 faces. The model achieves a mean average error of 4.08 on standardized datasets (Agustsson et al. 2017). For the gender prediction the model produces a value between 0 and 1, with values below 0.5 being male and above being female faces.

4. Results

4.1. Object detection

We summarize the results of the object detection by looking at the 10 most frequent objects overall and per movie. Table 2 and 3 show the objects starting with the most frequent per unit. Freq is the absolute number of detected instances while % is the percentage of frames at least one of the specific objects was detected.

Table 2. Detected objects per movie and overall (part 1).

Table 3. Detected objects per movie and overall (part 2).

Persons are the most frequently detected “objects” (figure 1). Other frequent objects are mostly furniture (book, chair), clothes (tie, handbag) and drinking objects (cup, wine glass).

Figure 1. Frame with the most detected persons ( Metropolis).

Comparing the movies, we identified that movies below 90% of frames with persons are indeed the more action-oriented movies ( Avengers,Metropolis) or include fantasy/animal-like characters ( Wizard of Oz). Many modern objects (e.g cell phones and airplanes) are more frequent in the contemporary movie Avengers (figure 2). One outlier we identified is the clock-object in Metropolis, which is not a frequent object in the other movies but represents a well-studied reoccurring motif of this specific movie (figure 3; cf. Cowan 2007).

Figure 2. Detected airplanes in Avengers.

Figure 3. Clocks as a reoccurring motif in Metropolis.

While we did not perform a systematic evaluation, but we identified a lot of mistakes in the prediction e.g. guns were predicted as handbags or the character “Cowardly Lion” in Wizard of Oz was oftentimes predicted as dog (figure 4).

Figure 4. The “Cowardly Lion” in Wizard of Oz detected as „dog“.

Nevertheless, we see potential in the method of object detection to explore specifics of the mise-en-scène as well as motif-like reoccurring objects in movies (Zaharieva / Breiteneder 2012). Furthermore, as object classes of the COCO dataset are not necessarily fitting for movies, we recommend exploring the possibilities of post-training via Detectron to analyze objects that are not part of the pretrained models.

4.2. Emotion recognition

For the emotion recognition we decided to create an average for a frame if multiple faces are detected. If no face is detected, we mark the frame with missing values. Table 4 summarizes the results. Maximums and minimums are marked in bold.

Table 4. Emotion values per movie and overall (M=mean, Max=maximum, Sd=standard deviation).

Overall, highest averages for emotions are the neutral (M=0.24) and the sad class (M=0.29). Surprise (M=0.11) and disgust (M=0.00) are rather rare among the movies. The two comedies in the movie corpus (Wizard of Oz, Some Like it Hot) do indeed have the highest happy-averages (M=0.13) (figure 5).

Figure 5. Frame with maximum happy value (Some Like it Hot).

However, the results are rather inconsistent since Wizard of Oz has also the highest sad- and angry-averages and therefore is the movie with generally the strongest emotional expressions. Breakfast at Tiffany’s on the contrast is the most neutral movie (M=0.37; figure 6).

Figure 6. Frame with highest neutrality value in the corpus (Breakfast at Tiffany’s).

Additionally, we performed a Welch-ANOVA to investigate if the movies differ to each other significantly (all requirements for the test are met according to Field (2009)). Indeed, we do find significant differences (p<0.05) for all emotion categories but rather small effects according to Cohen (1988) defining η²<0.01 as weak, <0.06 as moderate and <.14 as strong effect. We report the p-, F- and η²-value (table 5).

Table 5. Results of Welch-ANOVA-Tests for all emotion categories

The strongest effect can be seen for neutral. Performing post-hoc tests and inspecting a box-plots graph (figure 7) we identified Breakfast at Tiffany’s as interesting outlier. This might be due to the fact that the main characters of the movie try to stay rather “unaffected” up until the ending of the movie while Wizard of Oz, as a musical, consist of strong emotional outbursts.

Figure 7. Box-plots graph for the emotion class neutral

4.3. Gender- and age-recognition

Table 6 illustrates the descriptive statistics for the gender- and age-detection.

Table 6. Descriptive statistics for age and average gender.

The average age is for most movies is around 40 which is a rather consistent over-estimation since most leading actors in the selected movies are around 30. Performing a Welch-ANOVA shows that the difference between the movies is significant (p<0.001, F=336.07, η²=0.09) with a moderate effect. The strongest outlier movie, as shown with post hoc tests, is Wizard of Oz with a child/teenager as leading actor that gets correctly detected as around 14-16 years old (figure 8).

Figure 8. Lowest age in the corpus (Wizard of Oz).

An average score for gender below 0.5 points to more male detections and it is striking that all movies point below 0.5, thus a more frequent representation of males which is in line with the reality of the movies. There is a significant difference considering gender but with a smaller effect compared to age (p<0.001, F=251.36, η²=0.06) and with the strongest differences concerning Wizard of Oz. The differences become apparent regarding the distribution of gender-classes (table 7). We assigned every frame with male if average gender > 0.6 and female if <0.4. We decided to include a class androgynous for in-between-values pointing to either multiple genders on one screen or uncertainty by the model.

Table 7. Frequency distributions of gender classes.

Wizard of Oz has the most frames classified as androgynous. In general, this means that female and male characters are equally on the frame but in this case the classification is due to the high number of human-like fantasy creatures for which the model is unsure to pick a gender (figure 9).

Figure 9. An “androgynous” face (Wizard of Oz).

5. Discussion

While this study was rather small and exploratory in the approach, we did gain important first insights for our future research. Overall, we find it promising that we were able to find significant results, even for this small set of movies. For object detection we see the most potential in adjusting pretrained models to objects that are of interest for a specific research question. We see a lot of potential for interesting diachronic but also genre-based emotion and gender analysis with larger corpora. For this case study, we did not find striking differences of method performance considering technical differences between the movies. We are planning systematic evaluations on a cross section of movies of different decades to get a better understanding on the performance of the methods before we move on to explore more concrete research questions. Modern cultural artefacts have shown to be of interest for gender studies in the DH context (Schmidt et al. 2020d). We see potential concerning research on the intercourse of gender and film studies. We plan to explore the relationship of gender representations with expressed emotions throughout the time to explore how the representation of gender roles developed. Furthermore, we want to also explore multimodal approaches combining the various modality channels of movies (similar to Schmidt et al. 2019b).

Appendix A

Bibliography
  1. Agustsson, Eirikur / Timofte, Radu / Escalera, Sergio / Baro, Xavier / Guyon, Isabelle / Rothe, Rasmus (2017): “Apparent and Real Age Estimation in Still Images with Deep Residual Regressors on Appa-Real Database”, in: 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017) 87–94 DOI: 10.1109/FG.2017.20.
  2. Arnold, Taylor / Tilton, Lauren (2019): “Distant viewing: Analyzing large visual corpora”, in: Digital Scholarship in the Humanities DOI: 10.1093/digitalsh/fqz013.
  3. Baxter, Mike / Khitrova, Daria / Tsivian, Yuri (2017): “Exploring cutting structure in film, with applications to the films of D. W. Griffith, Mack Sennett, and Charlie Chaplin”, in: Digital Scholarship in the Humanities, 32, 1: 1–16 DOI: 10.1093/llc/fqv035 .
  4. Burghardt, Manuel / Kao, Michael / Walkowski, Niels-Oliver (2018): “Scalable MovieBarcodes – An Exploratory Interface for the Analysis of Movies”, in: IEEE VIS Workshop on Visualization for the Digital Humanities 2.
  5. Burghardt, Manuel / Kao, Michael / Wolff, Christian (2016): “Beyond Shot Lengths – Using Language Data and Color Information as Additional Parameters for Quantitative Movie Analysis”, in: Digital Humanities 2016: Conference Abstracts. Jagiellonian University & Pedagogical University, Kraków 753-755.
  6. Byszuk, Joanna (2020): “The Voices of Doctor Who – How Stylometry Can be Useful in Revealing New Information About TV Series”, in: Digital Humanities Quarterly 014, 4.
  7. Cohen, Jacob (1988): Statistical power analysis for the behavioral sciences. Academic press.
  8. Cowan, Michael (2007): “The Heart Machine: 'Rhythm' and Body in Weimar Film and Fritz Lang’s Metropolis”, in: Modernism / Modernity 14, 2: 225–248 DOI: 10.1353/mod.2007.0030.
  9. Deldjoo, Yashar / Elahi, Mehdi / Cremonesi, Paolo / Garzotto, Franca / Piazzolla, Pietro (2016): “Recommending Movies Based on Mise-en-Scene Design”, in: Proceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems 1540–1547 DOI: 10.1145/2851581.2892551.
  10. DeLong, Jordan (2015): “Horseshoes, handgrenades, and model fitting: The lognormal distribution is a pretty good model for shot-length distribution of Hollywood films”, in: Literary and Linguistic Computing 30, 1: 129–136 DOI: 10.1093/llc/fqt030.
  11. Field, Andy P. (³2009): Discovering statistics using SPSS: And sex, drugs and rock „n“ roll. SAGE Publications.
  12. Flueckiger, Barabara (2017): “A Digital Humanities Approach to Film Colors”, in: The Moving Image: The Journal of the Association of Moving Image Archivists 17, 2: 71–94. JSTOR DOI: 10.5749/movingimage.17.2.0071.
  13. Goodfellow, Ian J. et al. (2013): Challenges in Representation Learning: A report on three machine learning contests. arXiv:1307.0414 [cs, stat] <http://arxiv.org/abs/1307.0414> [14.06.2021].
  14. Halbhuber, David / Fehle, Jakob / Kalus, Alexander / Seitz, Konstantin / Kocur, Martin / Schmidt, Thomas / Wolff, Christian (2019): “The Mood Game - How to use the player’s affective state in a shoot’em up avoiding frustration and boredom”, in: Alt, Florian / Bulling, Andreas / Döring, Tanja (eds.): Mensch und Computer 2019 - Tagungsband. New York: ACM DOI: 10.1145/3340764.3345369.
  15. Halter, Gaudenz / Ballester-Ripoll, Rafael / Flueckiger, Barabara / Pajarola, Renato (2019): “VIAN: A Visual Annotation Tool for Film Analysis”, in: Computer Graphics Forum 38, 3: 119–129 DOI: 10.1111/cgf.13676.
  16. Hartl, Philipp / Fischer, Thomas / Hilzenthaler, Andreas / Kocur, Martin / Schmidt, Thomas (2019): “AudienceAR - Utilising Augmented Reality and Emotion Tracking to Address Fear of Speech”, in: Alt, Florian / Bulling, Andreas / Döring, Tanja (eds.): Mensch und Computer 2019 - Tagungsband. New York: ACM DOI: 10.1145/3340764.3345380.
  17. Hołobut, Agata / Rybicki, Jan / Woźniak, Monika (2016): “Stylometry on the Silver Screen: Authorial and Translatorial Signals in Film Dialogue", in: Book of Abstracts of the International Digital Humanities Conference (DH) (2016).
  18. Hołobut, Agata / Rybicki, Jan (2020): “The Stylometry of Film Dialogue: Pros and Pitfalls", in: Digital Humanities Quarterly 014, 4.
  19. Howanitz, Gernot / Bermeitinger, Bernhard / Radisch, Erik / Sebastian Gassner / Rehbein, Malte / Handschuh, Siegfried (2019): “Deep Watching - Towards New Methods of Analyzing Visual Media in Cultural Studies", in: Book of Abstracts of the International Digital Humanities Conference (DH) (2019).
  20. Hoyt, Eric / Ponto, Kevin / Roy, Carrie (2014): “Visualizing and Analyzing the Hollywood Screenplay with ScripThreads", in: Digital Humanities Quarterly 008, 4.
  21. Kuhn, Virginia / Craig, Alan / Simeone, Michael / Satheesan, Simeone P. / Marini, Luigi (2015): “The VAT: Enhanced video analysis", in: Proceedings of the 2015 XSEDE Conference: Scientific Advancements Enabled by Enhanced Cyberinfrastructure 1–4 DOI: 10.1145/2792745.2792756.
  22. Kurzhals, Kuno / John, Markus / Heimerl, Florian / Kuznecov, Paul / Weiskopf, Daniel (2016): „Visual Movie Analytics", in: IEEE Transactions on Multimedia 18, 11: 2149–2160 DOI: 10.1109/TMM.2016.2614184.
  23. Lin, Tsung-Yi / Maire, Michael / Belongie, Serge / Bourdev, Lubomir / Girshick, Ross / Hays, James / Perona, Pietro / Ramanan, Deva / Zitnick, C. Lawrence / Dollár, Piotr (2015): Microsoft COCO: Common Objects in Context. arXiv:1405.0312 [cs] <http://arxiv.org/abs/1405.0312> [14.06.2021].
  24. Masson, Eef / Olesen, Christian G. / Noord, Nanne van / Fossati, Giovanna (2020): “Exploring Digitised Moving Image Collections: The SEMIA Project, Visual Analysis and the Turn to Abstraction", in: Digital Humanities Quarterly 014, 4.
  25. Moßburger, Luis / Wende, Felix / Brinkmann, Kay / Schmidt, Thomas (2020): “Exploring Online Depression Forums via Text Mining: A Comparison of Reddit and a Curated Online Forum", in: Proceedings of the Fifth Social Media Mining for Health Applications Workshop & Shared Task 70-81.
  26. Ortloff, Anna-Marie / Güntner, Lydia / Windl, Maximiliane / Schmidt, Thomas / Kocur, Martin / Wolff, Christian (2019): “SentiBooks: Enhancing Audiobooks via Affective Computing and Smart Light Bulbs", in: Alt, Florian / Bulling, Andreas / Döring, Tanja (eds.): Mensch und Computer 2019 - Tagungsband. New York: ACM DOI: 10.1145/3340764.3345368.
  27. Pause, Johannes / Walkowski, Niels-Oliver (2018): “Everything is illuminated. Zur numerischen Analyse von Farbigkeit in Filmen", in: Zeitschrift für digitale Geisteswissenschaften.
  28. Pustu-Iren, Kader / Sittel, Julian / Mauer, Roman / Bulgakowa, Oksana / Ewerth, Ralph (2020): “Automated Visual Content Analysis for Film Studies: Current Status and Challenges", in: Digital Humanities Quarterly 014, 4.
  29. Rothe, Rasmus / Timofte, Radu / Van Gool, Luc (2018): “Deep Expectation of Real and Apparent Age from a Single Image Without Facial Landmarks", in: International Journal of Computer Vision 126, 2: 144–157 DOI: 10.1007/s11263-016-0940-3.
  30. Salt, Barry (1974): “Statistical style analysis of motion pictures", in: Film Quarterly 28, 1: 13-22.
  31. Schmidt, Thomas (2019): “Distant Reading Sentiments and Emotions in Historic German Plays", in: Abstract Booklet, DH_Budapest_2019. Budapest, Hungary 57-60.
  32. Schmidt, Thomas / Burghardt, Manuel (2018): “An Evaluation of Lexicon-based Sentiment Analysis Techniques for the Plays of Gotthold Ephraim Lessing", in: Association for Computational Linguistics (ed.): Proceedings of the Second Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature. Santa Fe, New Mexico 139-149.
  33. Schmidt, Thomas / Burghardt, Manuel / Dennerlein, Katrin / Wolff, Christian (2019a): “Sentiment Annotation in Lessing’s Plays: Towards a Language Resource for Sentiment Analysis on German Literary Texts", in: 2nd Conference on Language, Data and Knowledge (LDK 2019). LDK Posters. Leipzig, Germany.
  34. Schmidt, Thomas / Burghardt, Manuel / Wolff, Christian (2019b): “Towards Multimodal Sentiment Analysis of Historic Plays: A Case Study with Text and Audio for Lessing’s Emilia Galotti", in: Proceedings of the DHN (DH in the Nordic Countries) Conference. Copenhagen, Denmark 405-414.
  35. Schmidt, Thomas / Engl, Isabella / Halbhuber, David / Wolff, Christian (2020a): “Comparing Live Sentiment Annotation of Movies via Arduino and a Slider with Textual Annotation of Subtitles", in: DHN Post-Proceedings 212-223.
  36. Schmidt, Thomas / Engl, Isabella / Herzog, Juliane / Judisch, Lisa (2020d): “Towards an Analysis of Gender in Video Game Culture: Exploring Gender-specific Vocabulary in Video Game Magazines", in: Proceedings of the Digital Humanities in the Nordic Countries 5th Conference (DHN 2020). Riga, Latvia.
  37. Schmidt, Thomas / Halbhuber, David (2020): “Live Sentiment Annotation of Movies via Arduino and a Slider", in: Digital Humanities in the Nordic Countries 5th Conference (DHN 2020). Late Breaking Poster.
  38. Schmidt, Thomas / Kaindl, Florian / Wolff, Christian (2020b): “Distant Reading of Religious Online Communities: A Case Study for Three Religious Forums on Reddit", in: Proceedings of the Digital Humanities in the Nordic Countries 5th Conference (DHN 2020). Riga, Latvia.
  39. Schmidt, Thomas / Mosiienko, Anastasiia / Faber, Raffaela / Herzog, Juliane / Wolff, Christian (2020e): “Utilizing HTML-analysis and computer vision on a corpus of website screenshots to investigate design developments on the web", in: Proceedings of the Association for Information Science and Technology 57, 1: e392. DOI: 10.1002/pra2.392.
  40. Schmidt, Thomas / Schlindwein, Miriam / Lichtner, Katharina / Wolff, Christian (2020c): “Investigating the Relationship Between Emotion Recognition Software and Usability Metrics", in: i-com 19, 2: 139-151 DOI: 10.1515/icom-2020-0009.
  41. Vonderau, Patrick (2020): “Quantitative Werkzeuge”, in: Hagener, Malte / Pantenburg, Volker (eds.): Handbuch Filmanalyse. Springer Fachmedien 399–413 DOI: 10.1007/978-3-658-13339-9_28.
  42. Wei, Cheng-Yu / Dimitrova, Nevenka / Chang, Shih-Fu (2004): “Color-mood analysis of films based on syntactic and psychological models", in: 2004 IEEE international conference on multimedia and expo (ICME) (IEEE Cat. No. 04TH8763) 2: 831-834.
  43. Wu, Yuxin / Kirillov, Alexander / Massa, Francisco / Lo, Wan-Yem / Girshick, Ross (2019): “ Detectron2.” <https://github.com/facebookresearch/detectron2> [14.06.2021].
  44. Wulff, Hans J. (1998): “Semiotik der Filmanalyse: Ein Beitrag zur Methodologie und Kritik filmischer Werkanalyse", in: Kodikas/Code 21, 1-2: 19-36.
  45. Zaharieva, Maia / Breiteneder, Christian (2012): “Recurring Element Detection in Movies”, in: Schoeffmann, Klaus et al. (eds.): Advances in Multimedia Modeling. Springer 222–232 DOI: 10.1007/978-3-642-27355-1_22.
  46. Zhang, Kaipeng / Zhang, Zhanpeng / Li, Zhifeng / Qiao, Yu (2016): “Joint Face Detection and Alignment Using Multitask Cascaded Convolutional Networks", in: IEEE Signal Processing Letters 23, 10: 1499–1503 DOI: 10.1109/LSP.2016.2603342.
Notes
1.
2.
3.
4.