Segal, Zef
The Open University of Israel, Israel
zefsegal@gmail.com
The second half of the nineteenth century saw the establishment of numerous Hebrew periodicals that would play a major role in constituting a modern Hebrew “Republic of letters” (see Figure 1). They provided a platform for a lively discourse reflecting diverse ideological, political, and cultural approaches. In the multilingual context of Jewish communities, Hebrew was not an obvious choice for a journal, but it had the advantage of bridging the geographical and cultural distances between individuals and communities spread all over the Jewish Diaspora (Bartal 1994; Blondheim 1997; Soffer 2004a; Bartal 2007; Soffer 2009; Beer-Marx 2017). For Jewish communities—far from their homeland, lacking a central political and economic leadership, and spread throughout the world—the Hebrew press functioned as a public-sphere (Penslar 2000). Issues of these journals found their way to Jewish communities all around the world, thousands of miles far from their place of publication.

Figure 1: A three-dimensional map of the editorial locations of late nineteenth century Hebrew periodicals and their movement through time and space. The vertical axis represents the time of publication, and each vertical box represents a single periodical. Blue arrows reflect movement of the editorial locations of a certain periodical. Relocated periodicals are represented with a similar colored box in their new locations. The map was created in QGIS.
Hebrew journals were varied in nature and degree of institutionalization but most of them were limited in human and financial resources. As such, the publication of letters from private writers in near and far Jewish communities, made up a major part of the early Hebrew weeklies. The special characteristics of the Hebrew journals—their perception of themselves as communal enterprises; the geographical and cultural distance between communities that imagined themselves part of the same entity; and the relatively small group of modern Hebrew writers at the time—take to the extreme the "network authorship" journalistic model (Cordell 2015). This model assumes textual circulation and composition that is communal rather than individual. This communal authorship finds its expression, among other things, in the reuse and re-printing of texts. This reuse of texts occurs for different reasons: it could be a legitimate acknowledged citation of a Hebrew news item or a citation of the same translated news article from foreign newspapers, but it could also be a result of intentional plagiarism (Segal 2019). The anonymity of many authors contributed to this re-use of text.
While previous studies of Jewish journalistic networks used qualitative research methods (Bartal 1994; Soffer 2007; Kouts 2013, Beer-Marx 2017) such as discourse analysis, this study uses computational tools to provide a wider perspective on the phenomenon of Hebrew periodical networks. The corpus consists of five major Hebrew journals (HaTzfira, HaMagid, HaMelitz, HaLebanon, and Havazelet) published between 1874 and 1883. Recent digital approaches to the study of historical periodicals have shown such tools can provide a “distant reading”, which offers a new generalized approach to otherwise untraceable periodical networks (Murphy 2014).
The growing academic field of periodical studies is a direct result of advances in digital technology in the last two decades (Latham / Scholes 2006; DiCenzo 2015). Keyword-searchable digital archives and algorithmic mining tools have become accessible to an increasing number of researchers, and have transformed our view of journals from mere containers of discrete bits of information to autonomous objects of analysis. The most important of these tools is network analysis. The periodical’s complex and composite form "embodies the concept of the network on both a material level (in the juxtapositions and interconnections it generates between different texts) and on an institutional level (in the collaboration between authors, editors, illustrators, publishers, and readers, which goes into producing it)" (Fagg et al. 2013). Accordingly, newspapers and journals offer researchers a perspective to engage with the question of how social and intellectual connections are forged, furthered, and diffused within a public sphere. The interest in periodical networks is seen in special issues dedicated to the topic in Victorian Periodicals Review (2011), American Periodicals (2013) and The Journal of Modern Periodical Studies (2014).
The network metaphor is primarily used as a model of two aspects of periodical culture. First, a single periodical forms an intertextual network of individual texts, authors, and titles (Murphy 2014; Segal forthcoming). Second, periodicals are rarely isolated and exist within a larger print culture. Contributors, editors, publishers, and readers form an important, though sometimes barely visible, social and institutional network connecting different periodicals (Colavizza et al. 2014; Cordell 2015). At the same time, this line drawn between a journal’s internal and external networks is extremely permeable and porous, since the contacts between internal and external "nodes" established in previous editions always form the basis for ongoing editorial decisions. Consequently, external and internal factors are only artificially separable over the course of a journal project.
The main problem is in detecting the existence and structure of the networks. In some cases, we can find a network of authors by recognizing recurring names (Haberman 2008; Ehrlicher / Herzgsell 2016). However, during the nineteenth century many articles lacked clear authorship, due to, among other things, a culture of reprinting (Cordell 2015) or the use of pseudonyms (Unsworth / Morton 1981; Mintz 1995). Although this excludes effectively using name-recognition methodology to uncover periodical networks, hints remain within the text. "In some cases," state Smith et al. (2013), "we cannot directly observe network links, or even get a census of network nodes, and yet we can still observe text that provides evidence for social interactions." Gennete’s (1997) notion of transtextuality – "all that sets the text in relationship, whether obvious or concealed, with other texts" – is useful in understanding Smith's remark. Instead of approaching networks through people and working outwards, we can identify and analyze networks through published texts and their similarities in exact phrasing and styles.
One direct method of tracing transtextuality is to identify copied and quoted texts within other texts. However, recycled text is often only a small portion of a complete document and may also be significantly rephrased. As a result, finding these reused fragments "has always been subject to the limitations of human reading and recollection" (Olsen et. al 2011). However, as Olsen et al. continue, "one tantalizing promise of emerging digital libraries is that computer technology may augment the scholarly functions of reading and recollection by identifying related passages in very large collections." This vision is supported by developments in the field of plagiarism detection (Brin et al. 1995; Lovepreet / Kumar 2020), as well as tailor-made computer codes for specific digital humanities projects (Lee 2007; Smith et al. 2013; Ganascia et al. 2014).
Accordingly, this study utilizes a commercial software, Originality, used by Israeli universities to check the originality of academic work. The software analyzes documents based on a previous corpus, identifying copied passages and sentences, and measuring the level of originality. It provides a detailed report for every document, which marks the reused text and its source. The results are astonishing and prove the existence of extensive and varied forms of textual reuse. The corpus consists of 27,348 journalistic items, including advertisements. Some form of textual overlap was identified in 12,102 of these items (see Figure 2 for network visualization), and the average time gap between an original text and its reuse was a little less than seven months.

Figure 2: Textual reuse, visualized as a network. In this figure, each node represents a single journalistic item, and each link represents an overlap of sentences between a pair of items. The graph was created by Gephi.
Text reuse indicates news flows and direct intellectual links. But a much more nuanced connection between texts can be revealed by similarity in style. Stylometry, the statistical analysis of a literary style, is used to prove the authenticity of documents or to settle questions of authorial identity in anonymous or disputed texts (Mosteller / Wallace 1963; Chance 2003). Much like our ability to recognize individual voices, stylometrists attempt to evaluate the lexical richness of texts and establish similarities and differences between authors.
This study uses stylo, a flexible R package for high-level analysis of writing style (Eder et al. 2016). Instead of using stylometric analysis for authorship detection, it is used as a classification method, differentiating between literary circles and journalistic genres (Stamatatos et. al 2000). This approach is based on the distribution of the most frequent words (MFW) within each textual item (articles or advertisements). In order to reduce the noise of the results, all non-Hebrew letters were omitted from the texts and all the textual files smaller than 500 bytes were omitted from the corpus. This approach was tested on a limited corpus of 6,175 journalistic items published in the five journals between 1882 and 1883. Our results show stylistic similarity between 49,236 pairs of items, 1,654 of which are meaningful similarities between journals. These similarities range from identical items or authors to similar tones and subjects.
The evaluation of the results was conducted in a mixed-methods approach: edge weights and delta-similarity were used to identify proximity between journalistic items, network analysis to identify clusters, and close reading of a smaller sample to develop a typology of similarities.
The results prove the existence of a journalistic network of writers, information, and readers, among journals published in distant locations and different ideological orientations. They reflect an increasing number of inter-journal writers, announcements, advertisements, as well as on-going discussions. As it seems, publishers, writers, and advertisers were aware of the existence of these transnational connections. The rapid development of a transnational Jewish public sphere in the second half of the 19 th century is revealed in this analysis. These results shed light on the multiple affordances of applying textual reuse and stylometric analysis to a large corpus of Hebrew journals. Those relate to the significance of Hebrew journals in the production and maintenance of a Jewish national imagination; to the characteristics of the emerging Hebrew print culture; and to the general understanding of journalistic networks.
This paper presents the research, its integration of multiple computational tools, and suggests a categorization of textual similarities and reuse: from linguistic conventions of the Hebrew language, such as reuse of biblical and rabbinic phrases, to outright plagiarism of articles.