Abstract
This study examines how elite figures shape polarisation on Twitter/X through the interplay of content, structure, and engagement strategy. Drawing on data from nine globally influential users (2010–2021), the research integrates natural language processing, network analysis, and causal modelling to test five hypotheses grounded in social identity, agenda-setting, and two-step flow theories. Entity co-occurrence networks reveal that polarised discourse forms denser, more clustered networks than non-polarised content, indicating tighter semantic cohesion around socially and politically charged entities. Thematic and sentiment analyses show that posts addressing non-core topics – particularly those concerning social justice, environmental sustainability, philanthropy, and global welfare – are nearly five times more likely to be polarised than core professional themes. Negative emotional tone further amplifies this effect, while higher tweet-to-retweet ratios reduce polarisation, underscoring the moderating role of original content production. A user-level mediation analysis tested whether topical diversity transmits the effect of follower scale on polarisation but found no significant indirect pathway, suggesting that larger audiences do not necessarily foster communicative moderation. The findings advance understanding of elite discourse by linking structural and thematic polarisation to behavioural mechanisms of engagement. Theoretically, the study bridges social identity, agenda-setting, and two-step flow frameworks to explain how elites balance audience alignment and expressive risk. Practically, it highlights how emphasising original content, inclusive framing, and professional identity consistency can mitigate divisive online dynamics and foster more cohesive digital publics. To support transparency and reproducibility, the dataset and analytical code are made publicly available.
Introduction
Social media platforms such as Twitter/X have evolved into central arenas of public communication and identity performance, where elite figures shape not only opinion but also the structural dynamics of discourse (Murthy, 2024). Politicians, business leaders, athletes, and entertainers collectively reach vast global audiences – numbering in the hundreds of millions – leveraging their visibility to influence agendas, mobilise publics, and construct narratives that transcend traditional media boundaries (Hunter & Biglaiser, 2025; Peter & Muth, 2023). Yet these same dynamics also amplify a growing concern: the polarisation of online discourse – the emergence of emotionally charged, ideologically divided, and structurally fragmented publics (Druckman et al., 2021; Katalinić et al., 2023; Settle, 2018).
Polarisation manifests in both affective and ideological forms. Affective polarisation reflects emotional hostility toward opposing groups, whereas ideological polarisation captures the divergence of beliefs and issue positions (Küçük & Can, 2020; Wang et al., 2023). The architecture of social platforms reinforces these tendencies through algorithmic curation, selective exposure, and feedback loops that privilege controversy and emotional intensity (Stroud, 2014; Van Bavel et al., 2021). Elites, in this context, act not merely as participants but as structural amplifiers of polarisation – shaping what topics gain traction, how issues are framed, and how audiences align around them (Colliander et al., 2017; Weismueller et al., 2024).
Importantly, polarisation should be analytically distinguished from related but non-identical phenomena such as controversy and advocacy. Controversy refers to the presence of disagreement or contestation around an issue, while advocacy denotes the expression of a clear normative position or evaluative stance. Polarisation, as conceptualised in this study, emerges when these elements interact to structure discourse in ways that intensify group differentiation, moral boundary drawing, or affective alignment among audiences. Accordingly, polarisation does not require overt ideological antagonism or negative tone; positively framed advocacy may still function as polarising when it reinforces in-group identity, activates moralised distinctions, or concentrates attention around contested issues. This perspective situates the study closer to affective and discursive accounts of polarisation, while remaining agnostic about fixed ideological positioning.
Elite Influence and Polarised Discourse
The communicative power of elite users stems from their dual role as content producers and agenda-setters. With vast audiences, their posts can rapidly diffuse through networks, affecting how issues and entities are perceived (Majic et al., 2020; Wies et al., 2023). However, elite discourse is not uniformly polarising. It varies according to professional identity, audience scale, and engagement strategies (Eslami et al., 2024). For example, politicians often articulate partisan messages that reinforce ideological boundaries, while entertainers or entrepreneurs may alternate between professional and socio-political commentary (Tsaliki, 2015). Empirical research supports these distinctions, illustrating how elite interactions across professional domains contribute to polarisation. A recent empirical study further demonstrates this dynamic: Kommiya Mothilal et al. (2022) found that celebrities frequently engage with politicians along partisan lines on Twitter/X, reinforcing ideological alignment and illustrating how entertainment and political spheres converge in shaping polarised discourse. Such findings highlight the need to integrate thematic, behavioural, and structural perspectives to fully understand how elite communication drives polarisation.
Previous studies have explored polarisation among general users (Garimella, 2018; Kearney, 2019) and investigated sentiment (Reiter-Haas et al., 2023) or topic polarisation (Falkenberg et al., 2022; Xing et al., 2024). Yet, few have examined how elite discourse structures polarisation through entity networks, where people, organisations, and issues co-occur within posts. Such semantic networks capture not only what elites talk about but also how their discourse links diverse actors and ideas, following approaches that employ semantic graph analysis to reveal relational structures in social media data (Bodaghi & Oliveira, 2024). Earlier work by Bodaghi and Oliveira (2020) likewise highlighted the value of structural perspectives, showing that the characteristics of rumour spreaders – such as connectivity, activity frequency, and interaction centrality – shape how information cascades evolve, thereby linking network structure to behavioural diffusion in polarised contexts. Furthermore, behavioural features – such as the tweet-to-retweet ratio – reflect elites’ balance between producing original frames and amplifying others’ messages, an overlooked mechanism in understanding polarisation dynamics (Bodaghi & Zhu, 2024; Katz & Lazarsfeld, 1955).
Indeed, despite extensive research on user-level polarisation, several important gaps remain in understanding elite communication – gaps that this study seeks to address through the following research questions (1) How do structural patterns in discourse (e.g. entity co-occurrence networks) differ between polarised and non-polarised content? (2) How do professional identity and topic alignment (core vs. non-core issues) shape polarisation among elites? and (3) How do audience size, topical diversity, and engagement style (tweet-to-retweet ratio) interact to influence polarisation? Addressing these questions requires integrating network analysis, topic modelling (Batool & Byun, 2024), and behavioural metrics to capture how elite communication strategies reinforce or mitigate polarisation.
Theoretical Framework and Research Hypotheses
This study builds an integrative framework combining social identity theory (Tajfel & Turner, 1979), agenda-setting and framing (Entman, 1993; McCombs & Shaw, 1972), and the two-step flow of communication (Katz & Lazarsfeld, 1955). From a social identity perspective, audiences interpret elite communication through shared group affiliations and identity cues. Consistent with Tsaliki (2015), engagement with familiar, profession-linked themes reinforces parasocial bonds and collective identification, stabilising audience responses and maintaining alignment with follower expectations. However, when elites extend their communication into moralised or cross-domain topics – such as social justice, environmental sustainability, or humanitarian issues – they often encounter heightened audience contestation and credibility tension (Ehrenfeld, 2021). Such cross-domain engagement can activate intergroup boundaries and defensive reactions (Smyczynski, 2021), creating conditions conducive to polarised discourse. This process reflects the broader psychological function of polarisation as a foundation for collective engagement, wherein shared opposition and group differentiation energise participation and identity expression (Smith et al., 2024).
From a network perspective, polarisation manifests in the structure of communication, consistent with network and homophily theories suggesting that social and semantic cohesion arise through repeated association and selective connection (Granovetter, 1973; McPherson et al., 2001). Prior research demonstrates that polarised discussions on social media produce dense, internally cohesive interaction networks centred on a few highly influential users (Garimella, 2018), and that emotionally charged or partisan content further intensifies engagement and attention concentration (Weismueller et al., 2024). Extending this logic to the semantic level, the present study conceptualises polarisation as reflected in denser, more centralised entity co-occurrence networks, where attention converges around a limited set of emotionally or ideologically salient named entities – not necessarily humans. Conversely, non-polarised discourse disperses attention across a broader semantic field, yielding sparser networks indicative of more distributed engagement and less concentrated influence.
Finally, from a communication-behavioural perspective rooted in two-step flow theory, audience scale and engagement style likely moderate polarisation potential. Prior work shows that follower scale shapes how messages perform: indegree (follower count) exhibits nonlinear effects on audience engagement, and allowing influencers creative freedom (more original content) can attenuate engagement deficits associated with very large followers (Wies et al., 2023). Building on these insights, we propose that elites with large, heterogeneous audiences may choose to diversify topical focus to maintain broad appeal – a strategic response that could reduce reliance on divisive content and lower the likelihood of polarised responses. We further argue that elites’ production strategy (e.g. relative rates of original posts vs. retweets) will shape framing control and exposure to external polarisation: greater original content should increase framing control and thus reduce polarisation risk, whereas heavy retweeting may import and amplify contentious narratives. These propositions remain empirically unsettled; this study tests them by integrating network, thematic, and behavioural measures.
Building on this theoretical foundation, the study advances five hypotheses linking content, structure, and behaviour in elite communication. First, consistent with the idea of entity-centralised framing, polarised posts are expected to exhibit denser and more clustered entity co-occurrence networks than non-polarised ones, reflecting heightened semantic cohesion around emotionally or ideologically salient entities (H1). Second, with respect to topical alignment, posts addressing non-core issues – particularly those beyond an elite’s professional domain – are anticipated to display greater polarisation than content focused on core themes (H2). Third, in line with engagement strategy theory, a higher tweet-to-retweet ratio, indicating greater reliance on original content, is expected to correlate with lower polarisation by allowing elites greater framing control (H3). Fourth, audience scale is theorised to shape topical diversity: elites with larger and more heterogeneous follower bases are predicted to engage with a broader range of subjects, thereby reducing dependence on divisive content (H4). Finally, it is proposed that this relationship between follower count and polarisation is mediated by topical diversity, such that broader thematic repertoires attenuate polarisation tendencies (H5).
Methodology
Twitter/X activity data were collected from nine globally influential public figures between January 2010 and December 2021. The sample was constructed using a follower-based elite selection strategy, grounded in established research on opinion leadership, agenda-setting, and digital influence (Katz & Lazarsfeld, 1955; McCombs & Shaw, 1972; Wies et al., 2023). Follower count provides a transparent and widely used proxy for potential audience reach, visibility, and agenda-setting capacity on Twitter/X, where message diffusion and discursive influence are shaped by network scale. At the time of data collection, seven individuals worldwide had surpassed 100 million followers. Five of these were included in the study: Elon Musk, Barack Obama, Cristiano Ronaldo, Katy Perry, and Narendra Modi. Two additional figures in this category (Rihanna and Justin Bieber) were excluded to avoid professional redundancy within the entertainment domain, which was already represented by Katy Perry. Donald Trump, who later entered the 100-million-follower group, was excluded because his account was suspended during part of the study period, resulting in discontinuous data that would undermine longitudinal comparability. To enhance professional and topical diversity, four additional elite figures with smaller but still substantial follower counts were included: Bill Gates, Melinda Gates, Marc Benioff, and John Collison. This stratified design enables comparison between extreme-reach elites and mid-tier elites, while spanning politics, entertainment, sports, technology, philanthropy, and entrepreneurship. The resulting sample balances maximum audience visibility with cross-domain heterogeneity, reducing the likelihood that observed patterns are driven by a single profession or celebrity category.
Alternative selection criteria – such as engagement metrics (likes, retweets, replies) or prior content characteristics – could have been employed. However, follower count was prioritised because it directly captures the breadth of potential exposure, which is central to the study’s theoretical focus on elite-driven polarisation through content diffusion, framing, and network structuring. While engagement metrics could highlight interaction quality, they risk endogeneity in polarisation studies, as polarised content inherently drives higher engagement (Falkenberg et al., 2022). Similar follower-based or stratified approaches have been adopted in prior research examining elite or influencer effects in political and digital communication contexts (Kommiya Mothilal et al., 2022; Wies et al., 2023). We acknowledge that elite case selection involves trade-offs. While follower-based sampling emphasises highly visible actors, it may underrepresent niche or emerging elites with smaller but more tightly engaged audiences. Nevertheless, this approach is theoretically appropriate for studying polarisation potential, as polarisation effects are most socially consequential when produced by actors with large and heterogeneous audiences. Accordingly, the study adopts an analytical generalisation logic, using theoretically informed elite cases to identify structural, thematic, and behavioural mechanisms linking elite communication to polarisation dynamics.
Twitter/X Activity and Engagement Overview of Selected Public Figures. Follower counts represent values as of 2024, reported to illustrate audience scale
Annotation and Polarisation
Traditional approaches to studying polarisation often rely on ideological scaling, mapping texts onto a continuous spectrum of political alignment (Németh, 2023). While effective in strictly political contexts, such approaches are less suited to the present study, which examines both political and non-political discourse across diverse domains. Detecting polarisation solely through ideological markers or keyword-based methods risks overlooking emotionally charged, value-laden, or controversial content that does not map cleanly onto partisan dimensions. To address this limitation, the study adopted a classification-based approach that operationalises polarisation through two theoretically grounded criteria: controversy and stance-taking. These criteria are not treated as synonymous with polarisation itself, but as necessary discursive components through which polarisation becomes observable in heterogeneous elite communication. Controversy captures the presence of potential audience disagreement or value conflict, while stance-taking reflects the elite actor’s active positioning within that contested space. Polarisation, in this framework, arises when an elite message simultaneously engages a contested issue and advances an evaluative position that can structure affective alignment or opposition among audiences. This distinction allows the study to capture both negatively and positively valenced polarised content, including advocacy-oriented messages that polarise through moral affirmation or identity signalling rather than overt antagonism.
This approach draws on prior work that conceptualises polarisation as a combination of affective engagement, moral evaluation, and discursive contestation rather than ideology alone (Kearney, 2019; Wang et al., 2023). Using this framework, tweets and retweets were classified as polarised or non-polarised using large language models (LLMs), which are well suited for capturing contextual nuance across heterogeneous topics. Specifically, GPT-3.5-Turbo (OpenAI) was used to annotate each post based on two criteria: (1) controversy, defined as content likely to provoke disagreement, debate, or value conflict among audiences and (2) stance-taking, defined as the expression of a clear evaluative position or opinion on an issue. Posts meeting both criteria – for example, advocacy for social causes, criticism of policies, or morally framed commentary – were classified as polarised. Posts meeting neither criterion, such as neutral announcements or routine fan engagement, were classified as non-polarised. To ensure conceptual clarity and reduce ambiguity, posts meeting only one criterion were excluded from analysis.
While theoretically grounded, this operationalisation – like all content-analytic approaches – has inherent limitations. As with most content-analytic methods, especially those applied to short and context-dependent texts, coding decisions may not always fully capture authorial intent or the nuanced interpretations of diverse audiences. Tweets may employ irony, humour, strategic ambiguity, or implicit framing that complicates the clear identification of controversy or stance-taking. Moreover, disagreement between inferred controversy and the author’s intended meaning is an unavoidable challenge in large-scale discourse analysis. These limitations are not unique to LLM-based annotation and have been widely documented in studies of polarisation, controversy detection, and stance analysis using both human and automated methods (Reiter-Haas et al., 2023; Wang et al., 2023). Importantly, the goal of this study is not to recover subjective intent, but to identify discursive features that are likely to function as polarising in public communication, consistent with how audiences encounter and interpret elite messages in practice.
In addition to these general interpretive challenges, the use of LLMs for annotation introduces model-specific limitations that warrant explicit acknowledgement. LLMs such as GPT-3.5-Turbo are trained on large, heterogeneous corpora that may encode normative assumptions, cultural biases, or dominant discourse frames, which can influence how controversy and stance-taking are inferred. As a result, classifications may reflect prevailing linguistic norms rather than universally shared interpretations, particularly for culturally specific, ironic, or context-dependent expressions. Moreover, LLM-based annotation is inherently model- and version-dependent: changes in training data, model architecture, or prompting strategies may yield variation in classification outcomes over time. While standardised prompts were used to ensure consistency within this study, full reproducibility across models or future versions cannot be guaranteed. These limitations are increasingly recognised in computational social science and should be understood as trade-offs associated with scalable, context-sensitive annotation rather than as flaws unique to this study.
To assess reliability, a subset of posts was manually reviewed by a human annotator, supported by Grok 4 (xAI), with final adjudication performed by the human reviewer. This validation demonstrated high agreement with GPT-3.5 classifications. In a random sample of 558 posts, agreement on surface topic assignment was 94.4% (Wilson 95% CI = 92.2–96.1), and agreement on polarisation status reached 99.5% (Wilson 95% CI = 98.4–99.8). These results provide strong evidence that, despite acknowledged interpretive limits, the annotation procedure yields a consistent and reliable operationalisation of polarisation at scale.
Topic and Sentiment Extraction
To examine thematic patterns, a two-stage topic analysis was applied to the annotated dataset. Recent methodological work shows the usefulness of transformer-based topic pipelines combined with classifier models for extracting coherent, domain-specific topics and predicting downstream labels (e.g. Gottumukkala et al., 2025). In the first stage, surface-level topics were identified using GPT-3.5-Turbo. In the second stage, these topics were clustered into broader thematic categories with Grok 3. For example, the surface-level topic assigned to the tweet, ‘Aiming for first flight of Falcon Heavy on Feb 6 from Apollo launchpad 39A at Cape Kennedy. Easy viewing from the public causeway’, was ‘Space X Falcon Heavy First Launch’, while the corresponding thematic topic was ‘Space and Technology’. To ensure reliability, a subset of posts was manually reviewed by a human annotator, aided by Grok 4, with the human judgement serving as final. In a random sample of 270 posts, agreement between Grok 3’s thematic assignments and the adjudicated human coding reached 77.8% (Wilson 95% CI = 72.4–82.3), indicating substantial consistency between automated clustering and human interpretation. To further characterise content dynamics, sentiment analysis was conducted using the Valence Aware Dictionary and sEntiment Reasoner (VADER), a lexicon and rule-based tool from the Natural Language Toolkit (NLTK). VADER provides positive, negative, and compound sentiment scores for each tweet and retweet, with the compound score – normalised to a range of −1 to 1, capturing overall sentiment. Formally, the compound score is calculated by equation (1), where
For each public figure and content type (tweets and retweets), a set of features was computed for each thematic topic: (1) frequency rate (proportion of posts with the topic), (2) compound mean (average sentiment score), (3) compound variance (variance in sentiment), and (4) polarisation frequency (proportion of polarised posts). These features were calculated separately for tweets and retweets to capture differences in engagement.
Network Construction and Structural Analysis (H1)
To examine H1: Entity-centralised framing, entity co-occurrence networks were constructed for each user, separately for polarised and non-polarised posts. For each figure, four distinct entity co-occurrence networks were generated – based on entities extracted from (1) polarised tweets, (2) non-polarised tweets, (3) polarised retweets, and (4) non-polarised retweets. These networks were designed to map relationships between entities mentioned within each category, thereby revealing how polarisation influences connectivity and influence patterns. This approach aligns with Bodaghi and Oliveira’s (2025) macro-level analysis of Twitter/X news entity graphs, which demonstrated that global information ecosystems are organised around a limited set of semantically central entities exerting disproportionate influence across discourse networks.
Named entity recognition (NER) was performed using the SpaCy library (en_core_web_lg model) to extract entities labelled as PERSON (individuals), ORG (organisations), and GPE (geopolitical entities, such as countries or cities). Prior to extraction, posts were pre-processed to remove emojis, URLs, and other non-text elements. In the resulting networks, nodes represent unique entities, and undirected edges connect pairs of entities co-occurring within the same post. Edge weights reflect the frequency of these co-occurrences across samples, with repeated mentions of the same entity in a single post collapsed into a single occurrence to avoid inflation through self-repetition. Networks were constructed using the NetworkX package in Python, with separate networks generated for each of the four content categories per figure. For example, Figure 1 illustrates the polarised and non-polarised retweet networks created for Barack Obama. To illustrate how these structural differences manifest in practice, consider Barack Obama’s retweet behaviour. In polarised contexts, his retweet networks concentrate tightly around a small set of politically and socially salient entities – such as ‘Congress’, ‘justice’, ‘healthcare’, and prominent civil rights organisations – producing dense, highly clustered semantic structures. These configurations reflect moments where discourse converges around contested institutional actors and moralised policy domains, effectively creating semantic echo chambers. By contrast, Obama’s non-polarised retweet activity spans a broader and more heterogeneous set of entities, including scientific initiatives, cultural events, educational programmes, and nonpartisan civic organisations, resulting in more diffuse and weakly clustered networks. This contrast highlights how polarisation is not merely a matter of topic choice, but of how attention becomes structurally concentrated around particular symbolic anchors in elite discourse. Retweet Entity Networks of Barack Obama. (a) Non-polarised network; (b) polarised network. Polarised retweet networks exhibit tighter clustering around politically and socially salient entities, whereas non-polarised networks display broader semantic dispersion. To reduce visual clutter, only a subset of nodes are labelled
To characterise structural properties, a set of standard network metrics was computed, including the number of nodes and edges (network size and connectivity), average weighted degree (entity connectivity), degree centrality (influential entities), closeness centrality (proximity within the network), clustering coefficient (local clustering), betweenness centrality (bridging roles), eigenvector centrality (entity influence), and Shannon entropy (diversity). Additionally, network density was computed as a measure of overall interconnectedness, accounting for both nodes and edges. Higher density indicates a more tightly knit network, while lower density reflects more fragmented discourse. Formally, the density (
Multilevel Logistic Regression (H2–H3)
To examine H2 (Topical Alignment) and H3 (Engagement Strategy), a tweet-level logistic regression analysis was conducted to model the likelihood that a post was polarised versus non-polarised. The analysis tested whether posts addressing non-core topics were more likely to be polarised, and whether a higher tweet-to-retweet ratio – indicating greater reliance on original content – was associated with reduced polarisation. Sentiment valence was included as a covariate to account for the emotional tone of posts, given prior evidence linking affective intensity with polarisation (Weismueller et al., 2024). Because the dataset was imbalanced, with non-polarised posts exceeding polarised ones (45,036 vs. 12,332), the analysis employed random undersampling of the majority class to achieve a balanced dataset of 24,664 observations (12,332 polarised and 12,332 non-polarised). This approach ensured model stability and interpretability by preventing bias toward the dominant class while preserving an equal number of examples across categories. The model was estimated using maximum likelihood estimation in the Statsmodels package in Python. Equation (3) presents the regression specification.
Mediation Analysis (H4–H5)
To evaluate H4 (audience scale and topical diversity) and H5 (mediation through diversity) we conducted a user-level causal mediation analysis following the counterfactual framework of Imai et al. (2010). Topic diversity for each user was operationalised as Shannon entropy of the user’s topic distribution:
where
Results
Consistent with H1, the entity co-occurrence networks derived from combined tweet and retweet content revealed systematic structural differences between polarised and non-polarised discourse across elite users. Non-polarised networks were substantially larger but more diffuse, typically containing four to six times as many entities (nodes) as their polarised counterparts (e.g. Elon Musk: 1,946 vs. 417 nodes; Barack Obama: 5,679 vs. 2,688). In contrast, polarised networks were markedly denser, reflecting tighter clustering of discussion around a smaller set of ideologically salient entities. This cross-user pattern is visualised in Figure 2(A), which displays individual-level differences in network density (polarised – non-polarised) along with bootstrap confidence intervals and the overall group mean. Structural Differences Between Polarised and Non-Polarised Discourse Networks (H1). (A) Per-user differences in network density (polarised – non-polarised); red dots indicate observed effects, horizontal grey lines show bootstrap 95% CIs, and the blue shaded band marks the group-level mean. (B) Log-scale comparison of tweet and retweet network densities across all users. (C)–(E) Distribution of user-level effects for density, global clustering, and degree centralisation; boxplots show interquartile ranges, while red dots denote individual users and outliers
Proceeding from the network density formulation in equation (2), permutation and bootstrap tests were applied to compare polarised and non-polarised networks across users. Permutation and bootstrap analyses further confirmed that these structural distinctions were statistically significant. As shown in Figure 2(C)–(E), polarised networks exhibited significantly higher density (mean difference = 0.0146, 95% CI [0.0069, 0.0249], Wilcoxon p = 0.0039) and global clustering (mean difference = 0.1408, 95% CI [0.0507, 0.2512], p = 0.0195), indicating stronger local cohesion and more modular discourse communities than non-polarised networks. Although degree centralisation differences were positive on average (mean = 0.0152, 95% CI [–0.0040, 0.0360], p = 0.25), the effect was not statistically significant, suggesting variation across users in how polarised discussions are structured around dominant entities.
At the individual level, these effects were most pronounced among political and philanthropic figures. For instance, Barack Obama’s polarised network displayed a density of 0.0082 compared to 0.0034 for his non-polarised posts (Δ = 0.0048, p = 0.0002), while Narendra Modi showed a comparable increase (Δ = 0.0013, p = 0.0002). Entertainment figures such as Katy Perry and Cristiano Ronaldo exhibited smaller yet directionally consistent differences, reflecting their lower engagement in contentious discourse. Figure 2(B) illustrates these comparative patterns across tweet and retweet networks on a log scale, showing consistently higher density values for polarised content.
Taken as a whole, the five panels in Figure 2(A)–(E) demonstrate that polarisation is structurally manifested through increased network cohesion and clustering, signifying an entity-centralised framing of discourse in which elite communication orbits around socially or politically charged topics.
Thematic Patterns and Sentiment
Consistent with H2, the thematic and sentiment analyses revealed that the likelihood of polarisation varied systematically with topic alignment and emotional tone. Descriptive results confirmed a strong alignment between professional identity and dominant thematic focus. Musicians (e.g. Katy Perry) primarily engaged in entertainment and personal expression; business figures (Elon Musk, Bill Gates, Marc Benioff) focused on technology and entrepreneurship; political leaders (Barack Obama, Narendra Modi) emphasised governance and public affairs; philanthropists (Melinda Gates) concentrated on global health and social justice; and athletes (Cristiano Ronaldo) centred on sports and leisure. Hashtag usage further reinforced this alignment (e.g. #AmericanIdol for Perry and #MannKiBaat for Modi).
Building on this descriptive foundation, the analysis applied the multilevel logistic framework defined in equation (3) to test whether emotional tone and topic type predicted the likelihood of polarisation. The model achieved good fit (McFadden’s pseudo-R2 = 0.0896; LLR p < 0.001), indicating that these factors explained meaningful variance in polarisation likelihood. Results showed that non-core topics – those extending beyond a user’s primary professional domain – were strongly associated with higher polarisation (β = 1.57, SE = 0.04, z = 39.97, p < 0.001), corresponding to an odds ratio of 4.82. This means that posts on non-core or socially charged topics – spanning domains such as social justice and equality, environmental sustainability, philanthropy and humanitarian action, and community and cultural engagement – were nearly five times more likely to be polarised than those addressing core professional themes. Figure 3(A) visualises these findings, comparing average polarisation rates across core and non-core thematic categories. Across all professional domains, non-core topics consistently exhibited the highest levels of polarisation among the topics discussed by the users, with the notable exception of Narendra Modi, whose most polarised discourse centred on military affairs. However, since military affairs do not stand for his core professional domain of governance, this case does not break the broader pattern – polarisation remains higher when elites engage with non-core topics. An additional noteworthy pattern in Figure 3(A) appears for Melinda Gates, whose ‘core’ interests – gender equality and social justice – substantially overlap with non-core, high-polarisation domains for others, resulting in nearly identical polarisation rates across topic categories. Topic Alignment, Engagement Strategy, and Sentiment Effects on Polarisation Across Elite Users. (A) Average polarisation rates for core versus non-core thematic topics per user. (B) Tweet-level polarisation as a function of the tweet-to-retweet ratio. (C) Retweet-level polarisation as a function of the tweet-to-retweet ratio. In scatterplots (b) and (C), bubble colour represents average sentiment (darker shades indicate more positive tone)
As hypothesised in H3, emotional tone also played a significant role: the sentiment coefficient was negative and highly significant (β = −0.83, p < 0.001), indicating that less positive or more negative posts were substantially more likely to be polarised. In substantive terms, each one-unit increase in sentiment positivity reduced the odds of polarisation by roughly 56% (OR = 0.44). This pattern was consistent across professional domains – for example, Bill Gates’ non-polarised tweets averaged a sentiment of 0.38, compared with 0.13 for polarised ones, while Narendra Modi showed a similar drop (0.44 → 0.22). However, exceptions emerged in the entertainment sector: Cristiano Ronaldo’s polarised tweets often retained high positivity (mean = 0.45), reflecting celebratory or advocacy-driven expressions rather than antagonistic tone.
The model also detected a modest but significant effect of the tweet-to-retweet ratio (β = −0.011, p < 0.001, OR = 0.99), implying that greater emphasis on original tweeting – relative to retweeting – was associated with a lower likelihood of polarised expression. Figure 3(B) and (C) visualise this relationship: polarisation rates decline as the tweet-to-retweet ratio increases, suggesting that retweet-heavy behaviour amplifies polarisation, whereas original posting corresponds to more balanced discourse. Additionally, bubble colour intensity in both plots indicates sentiment tone.
Overall, these results provide strong support for H2 and H3. Polarisation is not uniformly distributed across all communication but is amplified when elites move beyond their core professional themes or adopt more negative emotional tone. Thematically, non-core engagement and affective intensity emerge as key drivers of divisive discourse, linking the thematic structure of elite communication to the emotional and network-level manifestations of polarisation observed in earlier analyses.
Engagement and Influence Dynamics
The interplay among posting behaviour, audience scale, and topical strategy further illuminates how elite communication shapes polarisation. The tweet-to-retweet ratio emerged as a critical behavioural dimension, distinguishing between content production and amplification. Users with higher ratios – those who post primarily original tweets – tended to exhibit broader topical repertoires and less polarised discourse, whereas those relying heavily on retweets displayed denser, more amplification-driven networks. This pattern reinforces the idea that original posting allows greater framing control, while retweeting amplifies externally generated – and often divisive—narratives (see also Figure 3).
To assess whether audience scale indirectly influences polarisation through topical diversity (H4–H5), a user-level mediation analysis was conducted based on the framework specified in equations (4)–(6), using regression and bootstrap estimation to derive the direct and indirect paths. As shown in Figure 4(A), the path from follower size to topic diversity was negative and nonsignificant (a = −0.035, p = 0.39), indicating that larger audiences did not correspond to broader topical variety. In contrast, topic diversity significantly predicted the proportion of polarised posts (b = 0.334, p = 0.007), while the direct effect of follower size on polarisation was small and nonsignificant (c′ = 0.006, p = 0.50). This significant association indicates that, independent of audience size, users with broader topical repertoires tended to express higher proportions of polarised content – a finding that underscores the role of issue diversity as an independent driver of polarisation rather than a mediating buffer. This result highlights that topic diversity functions as a direct correlate of polarisation, not merely as an intervening variable, suggesting that elites who traverse multiple issue arenas may encounter heightened discursive tension. Bootstrap mediation tests confirmed that neither the indirect (ACME = −0.015, 95% CI [−0.065, 0.014]) nor direct (ADE = 0.007, 95% CI [−0.021, 0.042]) effects differed significantly from zero, providing no evidence of mediation. Mediation Analysis of Follower Scale, Topical Diversity, and Polarisation. (A) Schematic mediation diagram showing point estimates for paths a, b, and c′; stars indicate conventional significance levels from OLS regressions, with bootstrap intervals for indirect, direct, and total effects shown below. (B)–(D) Scatterplots and OLS linear fits with 95% confidence regions for: (B) path a (log-transformed follower count vs. topic diversity); (C) path b (topic diversity vs. proportion polarised); and (D) path c′ (log-transformed follower count vs. proportion polarised). Red lines indicate OLS fits; shaded bands represent 95% confidence regions
The estimated paths derived from the mediation equations (equations (5) and (6)) are visualised by Figure 4(B)–(D). Path a (Figure 4(B)) shows a weak decline in topical diversity as follower counts increase; Path b (Figure 4(C)) reveals a positive association between diversity and polarisation; and Path c′ (Figure 4(D)) illustrates no meaningful direct relationship between follower scale and polarisation after accounting for diversity. The combined pattern suggests an inconsistent or suppressor-type relationship, where greater topical diversity slightly offsets the weak direct association between audience size and polarisation, though neither pathway reaches statistical significance.
Taken together, these findings indicate that engagement strategy – particularly the balance between original content and amplification – plays a more decisive role in shaping polarisation than audience scale or topical breadth. While follower size alone does not systematically predict polarisation, how elites use their platforms – whether to frame issues or amplify others – emerges as the key behavioural mechanism influencing divisive discourse.
Discussion
This study examined how elite figures across professional domains shape polarised discourse on Twitter/X through their content choices, engagement strategies, and audience scale. By integrating entity co-occurrence networks, topic modelling, and behavioural metrics, the analysis provides new evidence linking structural, thematic, and psychological dimensions of elite communication to polarisation dynamics. Overall, the results offer partial yet theoretically coherent support for the proposed hypotheses (H1–H5), indicating that polarisation arises not solely from ideology but from how elites navigate professional identity, emotional tone, and communicative control.
Structural Manifestations of Polarisation (H1)
Consistent with H1, polarised content produced denser and more clustered entity co-occurrence networks than non-polarised content. These structural differences reveal how polarised discourse gravitates toward a narrower set of emotionally or ideologically salient entities – such as ‘Congress’ or ‘COVID-19’ – that anchor discussion around divisive themes. Non-polarised discourse, by contrast, displayed broader and more diffuse networks, reflecting inclusive communication that spans diverse entities and audiences. This aligns with network and framing theory, suggesting that polarisation operates as a form of entity-centralised framing, where attention condenses around symbolic anchors that intensify shared identity and conflict. Such findings extend prior research that focused on user-based echo chambers, showing that elite-driven polarisation is embedded in the semantic architecture of communication itself. This structural concentration resonates with recent evidence that ideological polarisation in the U.S. electorate unfolds along multiple ideological and demographic dimensions (Ojer et al., 2025). It also echoes Bodaghi and Oliveira’s (2022) observation that fake-news diffusion on Twitter/X follows a theatrical division of roles – super spreaders, normal spreaders, and unwelcome spreaders – revealing how information visibility depends on a small subset of highly connected actors. This role concentration conceptually aligns with our finding that polarised discourse similarly centres around a few symbolic anchors and elite communicators.
Thematic and Affective Drivers (H2–H3)
The results for H2 and H3 highlight the central role of topical alignment and emotional tone in shaping polarised discourse. Posts addressing non-core topics – issues outside an elite’s professional domain – were nearly five times more likely to be polarised. These included themes such as social justice, equality, environmental sustainability, philanthropy, and humanitarian action, which inherently invoke moral or ideological positions. From a social identity theory perspective, these topics challenge audience expectations and activate out-group perceptions, thereby intensifying affective divides. Conversely, core topics (e.g., sports for athletes and technology for entrepreneurs) reinforce in-group norms and reduce contention by maintaining alignment with followers’ identity-based expectations.
Sentiment patterns further support H3, showing that negative or less positive emotional tone is a reliable predictor of polarisation. This aligns with prior evidence that emotional arousal and negativity bias increase the diffusion of divisive content. However, notable exceptions – such as Cristiano Ronaldo’s positively framed but polarised tweets – suggest that positivity can also accompany polarisation when used for advocacy or self-promotion, consistent with the logic of strategic framing rather than antagonism. Together, these findings synthesise social identity and framing perspectives, indicating that polarisation intensifies when elites depart from professional identity anchors and when affective cues heighten moral or ideological salience.
Engagement and Audience Effects (H4–H5)
Behaviourally, the tweet-to-retweet ratio proved to be a meaningful indicator of communicative control. Elites who produced more original tweets (high ratios) exhibited less polarised discourse, whereas those relying on retweets (low ratios) showed greater polarisation – supporting H3 and the two-step flow of communication model (Katz & Lazarsfeld, 1955). Retweeting amplifies third-party voices, often importing external frames and controversy, while tweeting original content allows elites to maintain narrative ownership. This pattern also aligns with the spiral of silence and social proof theories. As Noelle-Neumann (1974) proposed, individuals monitor prevailing opinion climates and often temper or conceal their views to avoid social isolation. Retweeting serves this function for elites: it allows them to engage with contentious or divisive topics indirectly – endorsing a position through amplification rather than explicit authorship. In parallel, Cialdini’s (2001) principle of social proof suggests that visible engagement by others legitimises particular narratives, making such indirect endorsement appear socially acceptable. Through retweets, elites thus navigate reputational risk while signalling alignment with dominant or morally resonant viewpoints.
Contrary to H4 and H5, follower count did not significantly predict either greater topical diversity or reduced polarisation. The mediation analysis revealed a weak, nonsignificant negative association between follower size and topic diversity, and no reliable indirect effect via diversity. Notably, topic diversity itself was positively associated with polarisation – users covering a wider range of themes tended to express more polarised content. This direct effect, although not embedded within a significant mediation pathway, highlights topical diversity as an independent correlate of polarisation – suggesting that diversification may expose elites to contested domains or moral arenas, thereby heightening rather than diffusing divisiveness. This suggests that broader topical repertoires do not attenuate polarisation; instead, they may expose elites to a wider set of contentious issues or audiences with conflicting expectations. The absence of a significant indirect pathway therefore reflects an inconsistent or suppressor-type pattern, where diversity and audience scale exert opposing but weak influences on polarisation.
Theoretical Integration
Synthesising across H1–H5, the findings illustrate how elite communication integrates structural, thematic, and behavioural dimensions of polarisation. From a social identity perspective, core-domain communication reinforces in-group cohesion, whereas cross-domain engagement activates intergroup distinctions and moral contestation. The observed network centralisation (H1) reflects this process structurally – polarised discourse coalesces around a limited set of emotionally charged entities, consistent with framing and agenda-setting theory. Behaviourally, results align with the two-step flow of communication and psychological signalling models: original tweets preserve framing autonomy, while retweets enable low-risk amplification consistent with the spiral of silence and social proof mechanisms. Finally, the positive association between topical diversity and polarisation indicates that discourse complexity among elites may heighten rather than diffuse ideological tension. Together, these results bridge micro-level message strategies with macro-level patterns of polarised structure, extending theoretical accounts of elite influence in digital communication.
Importantly, the patterns identified in this study should be understood as establishing a structural and behavioural baseline for elite-driven polarisation within the Twitter/X ecosystem during the pre-ownership period (2010–2021). During this period, platform governance, algorithmic ranking, and amplification norms were relatively stable compared to the subsequent transformations associated with the platform’s rebranding as X, including changes to content moderation policies, verification regimes, and engagement incentives. By isolating polarisation dynamics under these earlier conditions, the present findings provide a critical baseline for analysing how subsequent algorithmic and ownership shifts on X may intensify or reshape elite-driven discourse. Future research can build on this baseline to assess whether recent platform evolutions intensify entity-centralised framing, alter amplification strategies, or recalibrate elites’ trade-offs between audience loyalty and expressive autonomy.
Practical Implications
The findings yield several practical insights for platform governance, media strategy, and public communication. Encouraging elites to prioritise original content creation may help mitigate polarisation by maintaining narrative control and reducing the diffusion of externally polarised material. When addressing socially charged or cross-domain issues, elites can lessen divisive reactions by framing messages through inclusive, identity-consistent narratives that align with their professional persona and audience expectations. To illustrate how such reframing may operate in practice, consider a hypothetical elite post addressing climate policy. A polarising formulation such as ‘Climate inaction is a moral failure by political opponents who ignore the science’ directly activates ideological boundaries and moralised blame, increasing the likelihood of affective polarisation. The same issue could, however, be reframed as ‘Advances in clean energy innovation present major economic and environmental opportunities that merit cross-sector collaboration’, which foregrounds expertise and shared benefits rather than partisan attribution. While both messages communicate a clear stance, the latter emphasises problem-solving and inclusive framing, thereby reducing divisive potential while maintaining audience reach and engagement.
At the platform level, these findings suggest that recommender and ranking systems play a consequential role in shaping elite-driven polarisation. Algorithmic interventions that reward originality, contextual framing, and sustained authorship – while discouraging unreflective amplification such as indiscriminate retweeting of contentious material – may help limit the spread of affectively polarised discourse. From a platform design perspective, this implies that polarisation mitigation may benefit less from content removal and more from adjusting visibility incentives that privilege amplification over authorship. Designing systems that surface original, context-rich elite communication, rather than engagement-maximising signals alone, may reduce the structural concentration of attention observed in polarised discourse networks. The results also carry implications for media literacy initiatives. Because elite messages can polarise even when positively framed or advocacy-oriented, audiences may benefit from greater awareness of how stance-taking, amplification cues, and professional identity signalling shape affective alignment online. Media literacy efforts that emphasise how polarisation operates structurally and emotionally – rather than only ideologically – may improve public interpretation of elite discourse across political and non-political domains. Finally, collaboration between policymakers, communicators, and high-reach opinion leaders should focus on fostering empathetic, multi-perspective framing within core domains, rather than indiscriminate topical diversification. Such strategies emphasise depth over breadth, promoting understanding and trust within audiences while curbing the structural and emotional drivers of polarisation.
Limitations and Future Research
While the study offers novel insights into how elite figures shape polarisation through network and thematic structures, several considerations warrant further investigation. The dataset spans the period from 2010 to 2021, encompassing substantial changes to the Twitter/X platform architecture, including expansions of character limits, the transition from chronological to algorithmic timelines, and evolving norms of engagement and content visibility. Importantly, the core communicative mechanisms examined in this study – elite authorship, amplification via retweeting, audience scale, and thematic framing – remained structurally consistent throughout the observation period. As a result, the theoretical relationships identified here are not tied to any single interface configuration but reflect more general dynamics of elite-driven polarisation in networked public communication. Nevertheless, it is plausible that the strength or expression of these relationships varies across platform phases, as affordances governing visibility, virality, and user attention have evolved over time. The present analysis aggregates across years to identify stable structural and behavioural patterns. Future research could explicitly examine temporal heterogeneity by comparing periods before and after major platform changes, such as the introduction of algorithmic timelines or character-limit expansions. Such analyses would help clarify whether observed polarisation mechanisms intensify, attenuate, or transform under different technological regimes. Extending this framework to post-2021 developments on Twitter/X and to other social media platforms would further test the temporal robustness and generalisability of elite polarisation dynamics.
Methodologically, this study emphasises entity co-occurrence networks as a lens for mapping discursive structure, rather than direct user-interaction networks (Madraki et al., 2025). Integrating semantic and social network perspectives would provide a more comprehensive account of how elite framing strategies translate into audience engagement and behavioural polarisation. A related methodological sensitivity concerns the use of LLMs for polarisation annotation. While GPT-3.5-Turbo enables scalable and context-aware classification across heterogeneous content, such models may reflect cultural or normative biases embedded in their training data. In particular, controversy and stance-taking may be over- or under-identified in ways that privilege Western-centric political norms, moral vocabularies, or discursive styles, potentially affecting the classification of culturally specific or implicitly framed content. Although human validation mitigates this risk within the present dataset, future research could further strengthen robustness by employing cross-model validation (e.g. comparing outputs across different LLMs or versions), multilingual or culturally adaptive prompting strategies, and targeted sensitivity analyses to assess the stability of polarisation classifications across cultural contexts.
Another limitation concerns the study’s exclusive focus on elite actors. While this elite-centred design is a deliberate analytical strength – allowing clear identification of how high-visibility communicators structure discourse and amplify polarisation – it necessarily abstracts away from downstream audience dynamics. Polarisation may evolve differently in reply networks, quote-tweet cascades, or among mid-tier influencers who mediate between elites and general users. Future research could extend this framework by integrating audience reply structures, examining interactional polarisation within comment threads, or comparing elite discourse patterns with those of mid-level influencers and issue publics. Such extensions would enable multi-layered analyses of how elite framing strategies interact with grassroots engagement to shape polarisation across the broader communication ecosystem. Finally, ongoing advances in LLMs and multimodal analytics offer promising opportunities to capture narrative framing, emotional nuance, and causal dynamics with greater precision, deepening understanding of elite communication in an evolving digital public sphere.
Conclusion
This study advances understanding of online polarisation by demonstrating that elite-driven polarisation on social media is not reducible to ideological position-taking or audience size alone, but emerges from how elites strategically frame content, manage engagement, and structure discourse across domains. By integrating semantic network analysis, topical classification, and behavioural metrics, the findings show that polarisation is a multidimensional communicative outcome – produced through the interaction of content choices, emotional tone, and modes of amplification.
Across cases, polarisation was consistently lower when elites communicated within their professional core domains, employed more positive affect, and relied on original content rather than amplification. Conversely, polarisation intensified when elites engaged in non-core issues that carry moral or ideological salience and when they amplified external voices via retweets. These patterns align with and extend prior research on elite influence and polarisation, which emphasises agenda-setting power, affective signalling, and indirect endorsement as key mechanisms through which elites shape public opinion (Falkenberg et al., 2022; Katz & Lazarsfeld, 1955; Van Bavel et al., 2021). Our findings add a structural dimension to this literature, showing that polarised elite discourse is embedded in denser and more centralised semantic networks that concentrate attention around symbolic and emotionally charged entities. Importantly, the results also complicate common assumptions about audience scale. Contrary to expectations, follower count did not systematically moderate polarisation through greater topical diversity. Instead, topical diversity itself was positively associated with polarisation, suggesting that broader issue repertoires may expose elites to contested moral arenas and heterogeneous audiences with conflicting expectations. This finding challenges the notion that diversification necessarily moderates polarisation and instead suggests that elite breadth of engagement may amplify exposure to divisive frames. In doing so, the study contributes to a growing body of work showing that elite communication strategies – not merely reach or ideology – play a central role in shaping polarised discourse online.
Taken together, these results underscore that polarisation should be understood as a communicative process rather than a fixed ideological condition. Elites influence polarisation not only by what they say, but by how they structure attention, manage emotional tone, and balance authorship versus amplification. By linking micro-level communicative behaviour to macro-level structural patterns in discourse, this study bridges elite communication research with broader theories of affective polarisation, agenda-setting, and networked public spheres. More broadly, the findings suggest that efforts to mitigate polarisation must move beyond content moderation alone and consider how elite engagement strategies and platform affordances jointly shape the structure and emotional dynamics of public debate.
Footnotes
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
The dataset (Bodaghi, 2025) and analytical code
1
used in this study are publicly available.
Declaration of Generative AI and AI-Assisted Technologies in the Writing Process
During the preparation of this work the author used GPT in order to check for grammatical errors and improve the flow of the text. After using this service, the author reviewed and edited the content as needed and takes full responsibility for the content of the publication.
