Abstract
This study investigates the dynamics of anti-Muslim hate speech within Norwegian social media during the period between 2010 and 2021. Using a dataset of more than one million comments from Twitter and Facebook, we developed a custom hate speech classifier trained on an annotated corpus of 3,277 comments in Norwegian language. We identify that despite representing a small share of the total comments, hate speech content has increased over time. In an effort to understand the social network characteristics of hate speech content, we delve deeper into Twitter conversations as we can more easily identify how this content is spread. We develop network metrics to assess the prevalence, distribution, and diffusion of hateful content. The findings reveal that regardless of the number of users or tweets in a conversation, the volume of hateful content tends to remain constant. Furthermore, a small fraction of users contribute disproportionately to the dissemination of hate speech, with most conversations being limited in participant diversity. These results contribute to the growing field of computational social science by offering a novel methodology for studying hate speech in under-resourced languages and suggesting that mitigating hate speech may be possible through targeted network interventions rather than content removal alone.
Introduction
Hate speech against minorities has been a widespread phenomenon on social media platforms that harms both the individuals whose messages are targeted and the quality of the virtual public debate (Chen, 2017; Saha et al., 2019). Therefore, platform administrators, academics, civil society organizations, and public authorities have been allocating resources to create ways to detect and deal with hate speech. However, much of the progress towards developing tools to detect hate speech is made in the English language, leaving social media ecosystems using less spoken languages in knowledge blind spots (Poletto et al., 2021).
Due to the linguistic nature of hate speech, attacks and offenses acquire many potential forms that are specific to a socio-cultural community and how it treats individuals and groups from particular ethnicities, religions, gender, or sexual orientation (Baider, 2020). Moreover, the dynamic nature of language leads to temporal variation in how hate speech is expressed: derogatory terms that are common today may differ from those used a decade ago, and specific events, such as a terrorist attack, can trigger debates unique to a country. Consequently, many translation strategies that rely on the extensive resources available for the English language are limited in their ability to identify equivalent expressions and terms in other languages. Similarly, most hate speech detection models are trained on data covering a relatively narrow time span, raising concerns about their applicability to unlabeled data from periods not included in the training data. Despite recent remarkable advances in hate speech detection, there is still a pressing need for high-quality labeled datasets to support model training. In addition to the general lack of resources for less widely spoken languages, few studies to date have deployed hate speech detection models to quantitatively map the evolution of this phenomenon over time or to leverage the networked structure of online interactions to analyze patterns of hate speech diffusion.
To address some of these gaps, this article examines hate speech directed at Muslims on Norwegian-language social media. Muslim minorities have been frequent targets of hate speech in many Western societies, particularly on social media platforms (Awan & Zempi, 2017; Burke et al., 2020; Evolvi, 2019). Although hate speech against Muslims on Norwegian social media has been the subject of several qualitative studies (Berntzen & Sandberg, 2014; Fangen, 2020; Strømmen & Stormark, 2015), few efforts have sought to quantify its volume or patterns of diffusion. Moreover, the relatively limited amount of content in Norwegian language produced on social media has hindered the development of automated classification resources for detecting hate speech in this context. To our knowledge, this paper presents one of the few publicly available annotated training corpora designed to identify hate speech targeting Muslim minorities in Norwegian and aims to encourage further progress in this area.
Our study is also inspired by the growing body of research employing deep learning techniques to train models capable of detecting hate speech. Among these, neural network architectures have shown particularly promising results, achieving strong performance in automatic hate speech detection even for languages with limited annotated data (Hüsünbeyi et al., 2022; Jahan & Oussalah, 2023; Mittal, 2023; Ramos et al., 2024; Toktarova et al., 2023). Building on this literature, we find that a classifier based on an artificial neural network achieves state-of-the-art performance while requiring relatively modest computational resources.
Applying our classifier to comments collected from two major social media platforms—Twitter and Facebook—covering the period from 2010 to 2021, we find that the overall share of hateful comments is relatively low, although it has increased significantly over time. We also observe that a small subset of users is responsible for a disproportionately large share of the hateful comments detected. These findings align with previous research on English-speaking contexts (Vidgen et al., 2019) and with studies of other online phenomena such as misinformation (Frisli, 2025).
An important innovation of our study is the application of recent developments in network analysis to go beyond the simple quantification of hate speech (Schroeder et al., 2022). We develop new metrics to measure the diffusion of hate speech within social media networks and apply these to a subset of Twitter conversations. Interestingly, our results indicate that the volume of hateful content within a conversation remains relatively constant and does not vary with either an increase in participants or message volume on the platform. These findings suggest that the level of hate speech within a network tends to stagnate as discussions progress.
Beyond this introduction, the article is organized into four sections. Section 2 discusses the main challenges of empirically studying hate speech, presents our definition of the concept, and explains how it is operationalized. Section 3 details our data collection strategy, describes the construction of our classifier, and reports its performance on the collected data. Section 4 introduces the metrics used to assess the hatefulness of Twitter conversations and analyzes the characteristics of these discussions. Section 5 concludes the article.
Defining Hate Speech Against Muslims in the Norwegian Language
One fundamental challenge regarding the study of hate speech is its definition. There is no simple or universally accepted definition of hate speech. One reason being that hate speech manifests itself through different forms, from subtle constructions to more straightforward and explicit ones. In addition, the perception of what is considered hate speech is a direct result of the history and socio-cultural development of a country, leading to different legal definitions across nations. Although most national laws that criminalize hate speech circumscribe it to negative or derogatory comments about elements of someone’s identity (e.g., race, gender, religion, and sexual orientation), broader definitions also include collective attacks targeting all individuals or groups sharing that element (Baider et al., 2020). Furthermore, there are important ethical and legal discussions about to what extent hate speech should be allowed or not under the broader right of free speech (Howard, 2019).
Taking this context into account, most empirical studies on hate speech tend to adopt definitions that go beyond the legal ones, as they tend to be very narrow and specific. The adoption of broader definitions is grounded on the fact that, even if one form of hate speech cannot be characterized by legal definitions, it can still offend its target and deteriorate the virtual public spaces created by social media platforms. However, achieving a consensus definition is not a sufficient condition to identify hate speech. For example, different interpretations of the definition, lack of a complete context, or ethnicity of the annotators themselves contribute to the disagreement (Larimore et al., 2021). As a result, empirical studies that try to measure the occurrence and development of hate speech tend to produce divergent results.
Despite these conceptual challenges, the increasing number of studies trying to measure the phenomena contributes to a gradual refining of methods, the development of resources, and the identification of trends. One persistent challenge, however, is the bias of these efforts towards comments and discussions in the English language or other more popular spoken ones. This bias towards text availability makes many of the recently developed resources for the automatic detection of hate speech not directly applicable to less spoken languages (Röttger et al., 2022). Therefore, researchers interested in quantifying hate speech in social media environments based on less spoken languages must often “build the wheel again” and develop annotated datasets in the target language.
However, the accumulated experience with automated hate speech detection provides us with valuable lessons. A common limitation of many models for automated hate speech detection is the difficulty in identifying changes over time in the discourse targeting specific groups. As particular events can spark a cascade of hate comments on social media or specific derogatory terms can be adopted to refer to minority groups, the models are constrained by the historical and contextual characteristics of the training data. In other words, the shorter the time span of the training data in a classification model, the less robust it will be in detecting hate speech related to other periods or contexts (Florio et al., 2020). This means that for the purposes of having models that can quantify historical patterns of hate speech, we need labeled data covering our period of interest. By expanding the reach of automated classification models to other languages, we not only to identify patterns typical of that community of speakers but can also identify similarities between different linguistic communities. Therefore, developing automated classifiers for hate speech in less spoken languages can help us generalize findings about hate speech dynamics previously identified by studies looking at more populous community of speakers.
One fundamental challenge in studying hate speech lies in its definition. There is no single, universally accepted understanding of the term. This is partly because hate speech manifests in diverse forms, ranging from subtle insinuations to overtly explicit attacks. Moreover, perceptions of what constitutes hate speech are deeply shaped by a country’s historical and socio-cultural development, resulting in varying legal definitions across nations. While most national laws criminalizing hate speech restrict it to negative or derogatory comments directed at aspects of personal identity (e.g., race, gender, religion, or sexual orientation), broader interpretations also include collective attacks targeting all individuals or groups who share those characteristics (Baider et al., 2020). Additionally, important ethical and legal debates persist over the extent to which hate speech should be tolerated under the broader right to freedom of expression (Howard, 2019).
Given this complexity, most empirical studies of hate speech adopt definitions that extend beyond legal frameworks, which tend to be narrow and specific. The rationale for using broader definitions is that—even when an instance of hate speech does not meet the legal threshold—it can still cause harm to its targets and erode the quality of public discourse in online spaces. However, even with a working definition, identifying hate speech remains challenging. Differences in interpretation, incomplete contextual information, and even the demographic characteristics of annotators (e.g., ethnicity) can contribute to disagreements in labeling (Larimore et al., 2021). Consequently, empirical studies measuring the prevalence and evolution of hate speech often produce divergent results.
Despite these conceptual challenges, the growing number of studies measuring hate speech has contributed to a gradual refinement of methods, the development of new resources, and the identification of emerging trends. A persistent issue, however, is the bias toward studying English-language data or other widely spoken languages. This bias, rooted in differences in data availability, limits the direct applicability of many existing automated detection tools to less widely spoken languages (Röttger et al., 2022). As a result, researchers studying hate speech in such linguistic contexts often have to “reinvent the wheel” by developing annotated datasets tailored to their target language.
Nevertheless, the accumulated experience from automated hate speech detection research provides valuable lessons. A common limitation of existing models is their difficulty in capturing temporal changes in hate speech directed toward specific groups. Particular events may trigger spikes in hateful discourse, or new derogatory terms may emerge to describe minority groups, yet models trained on static data struggle to recognize these shifts. In other words, the shorter the temporal coverage of the training data, the less robust the model becomes when applied to other time periods or contexts (Florio et al., 2020). Consequently, models designed to track historical patterns of hate speech require labeled data spanning the full period of interest. Expanding automated classification models to new languages allows researchers not only to identify patterns specific to those linguistic communities but also to compare cross-linguistic similarities in hate speech dynamics. In this way, developing classifiers for less widely spoken languages contributes to a more general understanding of how hate speech evolves across cultural and linguistic boundaries.
Although some progress has been made in developing hate speech classifiers for other Scandinavian languages, such as Swedish and Danish (Fernquist et al., 2019; Sigurbergsson & Derczynski, 2020), very little has been done for Norwegian. Despite Norway’s relatively small population of 5.4 million in 2023, it has one of the highest rates of social media use per capita in the world. According to Statistics Norway, 1 88% of the adult population uses social media. Not surprisingly, the Norwegian authorities have shown increasing concern about the quality of public debate on these platforms and the prevalence of hate speech. Muslim minorities, in particular, have been frequent targets of online hate and a motivating factor in violent hate crimes in Norway (Bjørgo & Gjelsvik, 2017; Fangen & Nilsen, 2021). Although numerous studies have examined the discursive and sociological aspects of anti-Muslim hate speech, only a few have attempted to quantify it, and these typically rely on manual coding of small samples over limited time spans (Brekke et al., 2019; Burkal & Veledar, 2018).
In this study, we adopt a broad definition of hate speech, encompassing any derogatory or negative comments directed at Muslims as individuals or as a collective. Our conceptualization aligns closely with the notion of Islamophobia used by Vidgen and Yasseri (2020) and Bleich (2011). However, in operationalizing this definition, we apply a more detailed set of criteria to determine whether a comment should be classified as hateful, as outlined below.
2
1. Comments that hold Muslims collectively responsible for violence, terrorism, or political situations in other countries. 2. Comments portraying Muslims as inherently violent or dangerous. 3. Comments depicting Muslims as agents of the “Islamification” of Norwegian and/or Western cultures. 4. Comments asserting that Muslims aim to conquer Europe. 5. Comments claiming that Muslims possess a mentality inherently incompatible with Western values or lifestyles. 6. Comments demanding the banning of Muslim symbols or clothing. 7. Comments calling for the removal or exclusion of Muslims from Norway or Europe. 8. Comments dehumanizing Muslims.
Measuring Hate Speech
Data Collection and Hate Speech Classification
The data used in this study consist of public comments written in Norwegian by Facebook and Twitter users about Muslims between January 2010 and September 2021. To identify relevant comments, we employed three complementary strategies: (1) we collected posts from the official profiles of selected Norwegian media outlets 3 on both platforms using a keyword list developed in collaboration with three experts on hate speech and extremism in Norway. They were members of the larger project that initiated this study 4 ; (2) we applied the same keyword list to identify relevant comments on Twitter mentioning at least one of these terms; and (3) we collected comments from three public Facebook groups and profiles known for their critical stance toward Muslims. 5
This process resulted in five datasets:
For all datasets, we retained the full text of each comment, the date it was posted, and a unique identifier for the profile that authored it.
We applied the definition of hate speech discussed above to a training corpus of 3,277 Norwegian-language comments. These comments comprise a sample drawn from Datasets 1–4, supplemented with material collected from blogs, websites known for anti-Islamic content, and online discussion forums referencing Muslims. 7 These additional sources were included to increase the number of hateful comments in the corpus.
We employed a stratified sampling strategy based on the temporal distribution of comments. Weighting our sample according to the distribution of comments over time allowed us to better capture the variation in terminology and themes associated with discussions about Muslims in Norway during the study period. 8
For annotation, two members of the research team independently coded each comment. In cases of disagreement, the annotators discussed their rationale until consensus was reached. If no agreement could be achieved, the comment was discarded and replaced with another. The final corpus comprised 2,371 non-hateful and 806 hateful comments. 9
We tested several classification algorithms, including logistic regression, support vector machines, random forest, and a BERT-based model trained by the Norwegian National Library. 10 Among these, a simple artificial neural network with four layers achieved the best results. The annotated dataset was divided into three subsets: a training set (70%), a validation set (15%), and a test set (15%), all stratified according to the proportion of hateful and non-hateful comments. Prior to training, comments were cleaned by removing punctuation, stop words, URLs, and emojis. The remaining text was tokenized at the word level and used as the input layer for the neural network. 11
Model performance on the test and on the unlabeled datasets
In addition, we evaluated the model on a random sample of unlabeled data to assess its external performance. As shown in Table 1, the overall accuracy of the model decreased slightly compared to the test set. While the classification of hateful comments improved, the accuracy for non-hateful comments decreased substantially. This suggests that our results may slightly overestimate the prevalence of hate speech, and therefore, results considering small samples should be interpreted with caution. 12
Despite these limitations, the performance of the model aligns with the state of the art in the detection of hate speech. In a recent review of 22 hate speech classifiers, Chhabra and Vishwakarma (2023) report F1-scores ranging from 0.47 to 0.93. However, comparisons across studies remain limited due to differences in corpus composition and annotation strategies. Furthermore, the models that perform best usually target English-language data, while classifiers for other languages achieve F1-scores between 0.7 and 0.8.
Applying the Classifier
When applying the classifier to the comments in Datasets 1 and 2, we found that 3.7% of Facebook comments related to news articles were hateful, compared to 6.5% of Twitter comments. Although the total number of comments from Twitter (Dataset 2) is relatively small and does not allow robust temporal analysis, the Facebook data (Dataset 1) reveal a clear increase in hateful comments over time. Between 2010 and 2015, the monthly average number of hateful comments related to news stories was 35, increasing to 132 during the period 2016–2021. Figure 1 displays the distribution of hateful comments over time and across media outlets. This increase is likely due to the effects of the refugee crisis in Europe and its aftershocks in Norway. The peak in hateful comments across most outlets occurs around 2015 and 2016, coinciding with the years when the number of asylum seekers in Norway—primarily from Syria and Afghanistan—reached a historic high of approximately 35,000 in a period of 12 months. (Left) Number of hateful comments in Facebook posts of 12 different Norwegian media outlets. (Right) Number of hateful tweets based on general search. (Inset) Hateful comments in selected Facebook groups (%)
Applying the model to Twitter comments obtained through the general search (Dataset 3) reveals a similar pattern. Although the average monthly share of hateful comments remains low in general, it increases markedly over time (Figure 1, right panel). From 2010 to 2015, the model identified an average of 0.4% hateful comments per month, compared to 2.2% between 2016 and 2021. In particular, in certain periods (i.e., September to December 2020) the monthly share reached as high as 6%. Examining the content of these comments reveals that the surge in hateful discourse was linked to public debates in Norway at the time, which centered on reports that COVID-19 transmission rates were higher among ethnic minority groups predominantly composed of Muslims, such as Somalis and Pakistanis.
For public groups on Facebook critical to Muslims, the model finds a considerably higher share of hateful comments. For all comments collected from these groups, the model classified 4.8% as hateful. However, we find periods in which hateful comments can reach 20% of all comments written in a single month such (Figure 1, inset). Despite these substantial differences, the higher share of hateful comments in these groups is not surprising, as they are subject to far less moderation by group administrators and other users compared to the Facebook pages of media outlets.
When we analyze the profiles who write hateful comments, we identify a similar pattern in both platforms: a minority of profiles that are very active. In general, for both platforms, the median number of hateful comments by profile is 1. However, when we look at the top 1% of the most active profiles in the three Facebook groups (about 38 profiles), they are responsible for 16% of all hateful comments. For Twitter, the top 1% of the most active profiles (n = 14) is responsible for 28% of all hateful comments identified in dataset 3. Interestingly, there is noticeable variation over time among the most active profiles. After we inspected these profiles, we identified that while some previously active users appear to have left the platform altogether, others seem to have withdrawn from discussions related to Muslims.
Network Dynamics of Hate Speech: Twitter as a Case
Metrics to Assess Evolving Networks of Hate Speech
Beyond tracking the temporal evolution of hate speech, our classifier results also enable exploration of the network characteristics of Twitter comments included in dataset 5. Specifically, we draw on a recent framework for extracting contact networks among Twitter users based solely on non-verbal information derived from tweet and retweet activity (Schroeder et al., 2022). This approach allows us to reconstruct a graph of social connections between pairs of users by tracking both the frequency of shared content and the time elapsed between publication and resharing. Analyzing the topology of this derived network enables us to examine whether the volume of hate speech within a conversation correlates with either the number of participants or the overall volume of messages.
For the purposes of our analysis, we represent dataset 5 by
With such terminology, we now introduce some metrics to quantitatively assess hateful speech in social networks. We first introduce the hate score
Furthermore, we construct the network
Important for our analysis in the next section is that the hate score can be analyzed complementary to the corresponding “overall” topological distance between users
This means that the closer
From both the hate score
Second, as we have the time stamps of these tweets, we can map to what extent conversations are becoming more or less hateful overtime and with more or less participants. To this end, we compute the values of
Analysis of the Evolving Networks of Hate Speech
We start our analysis by inspecting the quantities characterizing the users, namely, the user’s hatefulness (Top-Left) Distribution of users hatefulness 
An interesting observation from Figure 2 is that approximately one percent of hateful tweets seem to be a typical fraction in the community we analyzed. From the bottom plot, showing
To address the features of conversations, we consider the hate score (Top-Left) Distribution of the hate score 
As for the distribution of hate scores (top-left), the distribution is much more uniform within a broad range of values. It seems that there are two regions, one approximately at
The scatter plots in Figure 3 (bottom) reveal further features of the hateful tweets. On the right, we observe that the number of users in a conversation and the corresponding hate score are negatively correlated. Notice that the conversations that align along the vertical line on the very left of the plot correspond to the conversations having exactly 1 user, that is, one single participant tweeting. This line is followed by other vertical lines that correspond to the conversations that have exactly 2 users, 3 users, etc.
These vertical lines having the least number of participants correspond to oblique lines in the plot on the right, showing the scattering of the size
This scaling relation can be written as
To end this section, we also discuss the features characterizing pairs of users, namely, the topological distance 
Conclusion
In this article, we analyze the statistical and topological features of hate speech diffusion on social networks. To do so, we provide a precise definition of hate speech in the specific context of anti-Muslim discourse in Norway and outline how this concept is operationalized and detected on the Norwegian Twitter network. By processing Twitter data from 2010 to 2021, we constructed a digital network in which the information exchanged among users was classified as either hateful or non-hateful. We then introduced a set of quantitative metrics designed to assess and characterize users, tweets, and conversations according to their degree of hatefulness.
Our analysis of hate speech targeting Muslims in the Norwegian Twitter network reveals two main findings. First, approximately one percent of all tweets appear to be hateful—a proportion that remains stable both in terms of users’ average level of hatefulness per se and in the hate scores assigned to exchanges between pairs of users.
Second, and perhaps more strikingly, we find quantitative evidence that the amount of hateful content within a conversation tends to remain constant, regardless of the number of tweets or participants involved. Unexpectedly, this pattern echoes findings by King et al. (2017), who analyzed social media control strategies in China. According to their study, rather than removing critical content outright, government actors often dilute it by flooding conversations with unrelated material. While we did not examine the content of non-hateful posts in detail, our findings suggest that a similar dilution effect may emerge organically—driven by ordinary users rather than coordinated moderation efforts. Another possible explanation for this pattern relates to reporting dynamics: as conversations expand, hateful content may be flagged more frequently by participants. Since our data were collected during a period when Twitter maintained relatively active content moderation policies, increased user reporting could have contributed to keeping hateful content at stable levels. Future research could further investigate the content of non-hateful posts within hateful conversations and explore how patterns of moderation have evolved following changes in Twitter’s (now X) ownership and policies.
This finding also opens new avenues for mitigating hate speech on social networks. In particular, network moderators and policymakers could aim to increase participation in conversations dominated by a small number of highly active users disseminating hate speech. If the total volume of hateful content does not scale with conversation size, then broadening participation may reduce the relative impact of hate sources within the network. On the other hand, exposing more users to hateful content most likely has undesired consequences that are beyond the scope of this study.
Future research could extend this work in several directions. First, having introduced metrics to characterize users and their interactions in terms of hatefulness, it is now possible to examine how these quantities evolve over time. Second, while the present analysis focused on individual and pairwise statistics, further research could explore higher-order network properties—such as degree-degree correlations based on hatefulness to assess whether users with high levels of hatefulness tend to interact with similarly hateful or less hateful counterparts.
Footnotes
Acknowledgments
We would like to thank all colleagues who contributed with inputs and comments for this article. Special thanks to Anne Birgitta Nielsen, Tone Liodden, Stian Lid and Henrik Wiig.
Author Note
Pedro G. Lind: School of Economics, Innovation and Technology, Kristiania University of Applied Sciences, Oslo, Norway.
Ethical Considerations
All data collected for this article followed the Data Protection Impact Assessment plan approved by Oslo Metropolitan University and the Norwegian Centre for Research Data (number 377704).
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The data collection and part of the analysis were conducted with the financial support of the Norwegian Directorate of the Children, Youth and Family Affairs (BufDir).
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
Notes
Author Biographies
