Abstract
The purpose of this study was to analyze the frequency and diversity of metacognitive and metalinguistic verbs (MCVs, MLVs) in the spoken language of Palestinian Arabic-speaking adolescents through narrative retelling and critical thinking tasks. An additional goal was to examine associations between the use of these verbs and syntactic complexity, macrostructure, and critical thinking. Forty-two seventh-grade students retold two fables and responded to critical thinking questions during individual interviews with an adult examiner. Transcripts were analyzed for the frequency and diversity of MCVs (e.g. think, know, believe) and MLVs (e.g. say, tell, explain), as well as for syntactic complexity, macrostructure, and critical thinking. Analyses included both frequentist and Bayesian approaches, allowing evaluation of evidence for both the presence and absence of associations. The critical thinking task elicited a higher frequency and greater diversity of MCVs than narrative retelling, whereas narrative retelling elicited a higher frequency and greater diversity of MLVs. Bayesian analyses provided decisive evidence for these task effects. Correlational findings were more selective: clausal density showed an uncorrected association with narrative MCV frequency proportion, whereas Bayesian evidence suggested moderate support for null associations with macrostructure and mean length of communication unit. Overall, adolescents’ use of MCVs and MLVs appears to be selectively recruited by task demands, with weaker and more context-dependent links to broader discourse and critical thinking measures.
Introduction
The use of metacognitive and metalinguistic verbs (MCVs and MLVs) in adolescent discourse provides valuable insight into cognitive and linguistic development (Astington & Olson, 1990; Nippold et al., 2017). MCVs such as think, know, and believe reflect one’s capacity for introspection, reasoning, and understanding others’ mental states—a key element of Theory of Mind (ToM; Astington & Olson, 1990; Feurer et al., 2015; Howard et al., 2008). In parallel, MLVs such as say, explain, and write indicate an awareness of language itself and an ability to discuss communication explicitly (Astington & Olson, 1990; Mertz & Yovel, 2009; Verschueren, 2004).
In adolescence, the production of such verbs increases in both frequency and diversity, especially in academic and narrative discourse (Nippold et al., 2017). Research with English-speaking adolescents shows that the frequency and diversity of MCVs and MLVs steadily increase across adolescence, paralleling gains in abstract reasoning and discourse sophistication. The use of these verbs is also associated with syntactic complexity, particularly the expansion of subordinate clauses, illustrating the “lexicon–syntax interface” (Nippold et al., 2020; Sun & Nippold, 2012). Adolescents who employ more MCVs and MLVs also tend to produce longer and more complex sentences and more coherent narratives (Nippold, in press; Nippold et al., 2020).
Nippold et al. (2020) demonstrated robust age-related gains in MCV production on a fable-based critical thinking task: 16-year-olds produced more MCV tokens and a larger set of distinct MCV types than 13-year-olds. Moreover, the use of MCVs among the older group correlated positively with the critical thinking scores, indicating a link between MCVs and the ability to demonstrate more elaborate, evidence-based reasoning. Findings also indicated that the semantic/pragmatic breadth of the older adolescents’ MCV inventory expanded beyond staple items (e.g. think, know, believe) to include stance- and evaluation-oriented verbs (e.g. convince, decide, disagree, trust, blame), suggesting a developmental shift from merely referencing mental states to expressing more nuanced meanings in writing. These patterns underscore that MCVs and MLVs are not mere lexical choices but reflect broader cognitive linguistic development. As Westerveld et al. (2023) further highlighted, these verbs often perform evaluative roles in discourse, helping narrators interpret events, attribute motives, and manage stance—capacities that bridge language, cognition, and social reasoning.
In Arabic, however, research on MCVs and MLVs remains limited despite their theoretical and practical importance. Arabic is a diglossic language, where children and adolescents grow up speaking a colloquial dialect (Spoken Arabic, SA) at home and in daily life, while acquiring Modern Standard Arabic (MSA) primarily through literacy and formal schooling. This dual linguistic environment shapes access to abstract vocabulary, including MCVs and MLVs, which may be more embedded in MSA discourse (Habib, 2022). Recent work with Palestinian Arabic-speaking preschoolers has shown that Internal State Terms (ISTs)—a broad category including perceptual, emotional, mental, and linguistic words are already used in narrative retellings, but that their frequency and diversity are influenced by diglossia and text complexity (Kawar et al., 2023). These findings highlight that younger Arabic-speaking children can produce concepts that reference cognitive and linguistic states, laying the foundation for the later emergence of more abstract MCVs and MLVs.
At the adolescent level, related phenomena have been documented through evaluative devices in narratives (frames of mind, hedges, direct speech)—categories that overlap with metacognitive and metalinguistic expressions. For example, Arabic-speaking adolescents (especially females) with normal hearing ability referred more frequently to thoughts and mental states than their deaf and hard of hearing peers (Kawar et al., 2024), underscoring the importance of these devices and the scarcity of research tracking their development in Arabic.
A broader cognitive framework relevant to these linguistic phenomena is ToM, which concerns the ability to attribute beliefs, intentions, and emotions to oneself and others. The present study draws on this literature to contextualize how mental-state reasoning may be reflected linguistically in discourse. Cross-cultural research on ToM shows that mental-state reasoning develops within cultural and linguistic environments that shape how internal states are discussed (Liu et al., 2008; Shahaeian et al., 2011; Westby, 2016).
In broad terms, individualist orientations—more characteristic of urban, Western, industrialized contexts—prioritize autonomy and unique perspectives and are often accompanied by elaborative parent–child talk about internal states. In contrast, collectivist orientations—more common in many non-Western contexts—emphasize interdependence, group harmony, and shared viewpoints, with socialization practices that may place relatively less weight on overtly individuated mental perspectives (Markus & Kitayama, 2014; Santos et al., 2017; Selcuk et al., 2023). These macro-level values shape children’s everyday micro-experiences (e.g. parental mental state talk, family discourse, schooling, socio-economic status [SES]), and together they influence both the timing and sequence of ToM development and the ways mental states are expressed in discourse (Selcuk et al., 2023; Slaughter & Perez-Zapata, 2014; Wellman & Liu, 2004). Importantly, contemporary societies do not necessarily fit a simple dichotomy: individualist and collectivist values often co-exist within the same country or even the same family, with children’s ToM trajectories varying by SES, residence (urban/rural), and caregiving practices (Ruffman et al., 2002; Selcuk et al., 2018, 2023; Shahaeian et al., 2011; Taumoepeau et al., 2019; Wellman & Liu, 2004). Accordingly, cross-cultural ToM research is invoked here not to evaluate adolescents’ ToM per se, but to provide a conceptual backdrop for understanding variability in MCV and MLV use across sociolinguistic contexts.
These orientations manifest in narrative socialization and discourse practices that plausibly condition adolescents’ use of MCVs and MLVs. In some Euro-American families, caregivers commonly scaffold children’s independent, coherent narratives, encouraging elaboration about thoughts, feelings, and reasons; by contrast, Latina mothers have been shown to favor collaborative storytelling that prioritizes shared participation over correction (Carmiol & Sparks, 2014). In some communities, storytelling is a collective activity in which audiences co-construct the tale, whereas other cultures discourage child storytelling in adult settings, pushing narrative practice into peer groups (Labov, 1972). Cultural values also shape themes and stance. For example, English-speaking families often emphasize autonomy and emotional expressiveness, yielding detailed, emotionally rich accounts. In contrast, Chinese families more often highlight communal obligations and moral lessons, fostering socially engaged, morally oriented narratives (Miller et al., 2012; Reese, 2013; Wang, 2013; Wang & Fivush, 2005). These contrasts map onto individualist–collectivist orientations rather than ethnicity or language per se. In Arab collectivist contexts, for example, family solidarity and group harmony are central, with extended-family participation common in social gatherings—practices that can tilt narrative purposes toward relational and moral evaluation (Bakalla, 2023; Bliss & McCabe, 2008; Vygotsky, 2012).
Cross-cultural ToM findings align with meta-analytic and comparative evidence indicating that, although most children reach explicit false-belief understanding around ages 4 to 5 years, sociocultural experiences—shaped by individualist vs. collectivist values—can shift the pace and order of ToM achievements (e.g. exposure to mental state talk, schooling, SES-linked resources). Notably, research has identified different sequences in the acquisition of ToM subskills where an “individualist” sequence (diverse desire → diverse belief → knowledge access → false belief → hidden emotion) was more often observed in the United States and Australia, and a “collectivist” sequence (diverse desire → knowledge access → diverse belief → false belief → hidden emotion) was often reported in China and Iran—patterns that likely reflect socialization priorities (e.g. valuing consensus and knowledge access vs. foregrounding divergent beliefs) (Shahaeian et al., 2011; Wellman & Liu, 2004). Evidence from Turkey illustrates the co-existence of these values within one society: macro-level norms interact with micro-level contexts (SES, caregiving, conversational style) to shape both timing and sequence of ToM, with middle-to-high-SES families exhibiting practices more typical of individualist contexts and lower-SES or institutional settings showing delays linked to fewer elaborative, mental-state-rich interactions (Selcuk et al., 2023).
Taken together, these cultural orientations bear directly on adolescents’ usage of MCVs and MLVs. Contexts that scaffold independent reasoning and elaborative mental-state talk are likely to elicit frequent and diverse MCVs (e.g. think, know, believe, infer, doubt) and MLVs (e.g. say, explain, argue, claim), especially when tasks invite stance-taking and justification. Conversely, contexts that prioritize group cohesion and moral propriety may favor didactic or normative framings, potentially channeling MCV and MLV use toward moral evaluation and knowledge access rather than divergent belief-ascription. Palestinian adolescents are therefore a compelling group: they develop in a diglossic Arabic environment (SA for everyday interaction; MSA through literacy/schooling) and are socialized within collectivist traditions alongside the formal influences of schooling in MSA—conditions that may differentially support access to abstract lexical items and shape when, how, and why MCVs and MLVs are used in narrative retelling and critical thinking discourse (Habib, 2022; Kawar et al., 2023).
Narrative retelling and critical thinking tasks provide ideal opportunities to elicit MCVs and MLVs. Narratives require reconstructing characters, settings, and events while attributing internal states, thereby encouraging the use of metacognitive language (Sun & Nippold, 2012). Narrative theory emphasizes two complementary “landscapes”: the landscape of action (events) and the landscape of consciousness (characters’ thoughts, feelings, intentions). Progress in narrative competence involves the increasing use of mental state language to explain why events unfold, not just what happens (Westby, 2016). Westby further notes that children’s narrative structure, cohesion, and understanding of deception correlate with the use of mental state language, indicating a close link between ToM and narrative proficiency.
These insights provided by previous researchers motivate further study of the use of MCVs and MLVs in adolescents’ narrative retelling and critical thinking. When speakers move beyond describing actions to articulating internal states and communicative intentions, they demonstrate advanced inferential reasoning and perspective taking.
Research using the Global Talking About Lived Experiences in Stories (TALES) protocol has shown that school-age children employ a wide range of evaluative devices—such as internal emotional states, mental states, causal explanations, hypotheses, and judgments—to convey what events mean to them (Westerveld et al., 2023). Global TALES is a standardized elicitation task for children’s personal narratives built around six open-ended prompts—excited/happy, worried/confused, annoyed/angry, proud, a problem situation, and something important—administered in a fixed order; when a child does not respond, examiners use a scripted follow-up and, if needed, neutral encouragers (e.g. “tell me more”), with prompts presented on laminated cards or a tablet. The protocol is translated by native speakers in each country to preserve instructions, order, and target emotions while ensuring cultural/linguistic appropriateness (Westerveld et al., 2022).
Factor analyses revealed clusters labeled causality, hypothesis, and judgment, underscoring that evaluative language coheres into functionally meaningful categories. Girls produced a wider range of evaluative devices than boys, though effects of prompt type were mixed (Westerveld et al., 2023). Although Global TALES focuses on spontaneous personal narratives, its emphasis on evaluative language aligns with the present study’s focus on MCVs and MLVs as resources for mentalizing, stance taking, and moral evaluation.
Critical thinking tasks, by contrast, prompt adolescents to evaluate morals, justify positions, and reflect on broader implications, often eliciting a richer range of MCVs (Nippold et al., 2017). Across several studies with English-speaking adolescents, critical thinking tasks using fables have elicited more frequent and diverse MCVs than narrative tasks and have highlighted age-related gains in critical thinking (Nippold et al., 2015, 2017, 2020). In particular, Nippold et al. (2020) used a written fable-based critical thinking task with 13- and 16-year-old adolescents and found that older adolescents showed (a) higher performance on critical thinking (more elaborated, reasoned responses), (b) more frequent and diverse use of MCVs, and (c) greater verbal productivity, measured in terms of the number of words and communication units (C-unit) produced. By definition, a C-unit consists of an independent clause and any modifiers attached to it. Importantly, syntactic complexity did not differ reliably by age and was not strongly associated with critical thinking, indicating that high critical thinking can be expressed without parallel increases in mean length of C-unit (MLCU) or clausal density (CD). These findings also showed that stronger critical thinking correlated with verbal productivity at both ages and with MCV frequency/diversity among older adolescents, supporting the idea that MCV use is a sensitive index of advanced reasoning in discourse.
Together, prior research suggests that task demands (narrative retelling vs. critical thinking) and developmental level jointly shape adolescents’ use of MCVs and MLVs, with critical thinking prompts—especially moral interpretation of fables—eliciting dense MCVs and MLVs usage and nuanced stance taking. In Arabic, the diglossic lexicon may further modulate access to abstract MCVs and MLVs, as such forms are often stabilized through literate, MSA-driven contexts yet used in SA-based, spoken tasks.
Building on this literature, the present study conceptualizes the use of MCVs and MLVs as discourse-level linguistic markers of advanced reasoning. Different discourse contexts are expected to place distinct cognitive and communicative demands on speakers, thereby differentially recruiting these verb classes. Narrative retelling primarily requires reconstructing events and reporting characters’ actions and dialogue, whereas critical thinking tasks require evaluating motives, interpreting themes, and justifying positions. These differing discourse demands may shape the extent to which adolescents draw on metacognitive and metalinguistic language resources. In turn, the use of these verbs is expected to relate systematically to broader discourse-level abilities, including syntactic complexity, narrative macrostructure, and performance on critical thinking measures. Within the diglossic context of Arabic, access to abstract lexical items—often associated with MSA—may further influence how these relationships are expressed in spoken discourse.
The Present Study
This study investigates the use of MCVs and MLVs as task-sensitive discourse resources in Palestinian Arabic-speaking adolescents across narrative retelling and critical thinking contexts. Specifically, we examine: (a) whether task demands differentially recruit MCVs and MLVs; (b) whether verb use is associated with syntactic complexity and narrative macrostructure; and (c) whether verb use is associated with critical thinking performance. By building on earlier research on ISTs in Arabic-speaking children and on evaluation devices in adolescents, this study extends prior work by focusing specifically on MCVs and MLVs as discourse indicators of cognitive linguistic development. The findings are expected to inform theoretical models of language development and applied approaches to supporting metacognitive and metalinguistic awareness among Arabic-speaking learners.
Guided by these aims, the following hypotheses were formulated to specify the expected patterns of task effects and linguistic–cognitive associations.
Method
Participants
The participants were 42 typically developing, monolingual Palestinian Arabic-speaking adolescents (21 boys and 21 girls), all enrolled in 7th grade in mainstream classrooms located in the central and northern regions of Israel. Participants ranged in age from 12 years and 1 month to 13 years and 1 month (M = 12;7 years, Standard deviation [SD] = 0.37).
All participants belonged to families of medium to high socioeconomic status, as measured by parental education levels. Inclusion criteria required that participants be typically developing, with no recorded or diagnosed reading, learning, social, or behavioral disabilities. This was confirmed through parental reports and school records, and none of the participants had any known reading difficulties. No formal standardized language assessments were administered as part of the present study. Participant status as typically developing was determined based on school records, parental report, and absence of diagnosed learning or language disorders. While this approach aligns with procedures used in prior discourse studies with adolescent samples, we acknowledge that the absence of standardized language profiling limits the precision with which “typical language development” can be operationalized.
Participants were recruited through social media platforms. Written informed consent was obtained from parents, and assent was obtained from the adolescents prior to participation. This study forms part of a larger research project on fable retelling and critical thinking among Palestinian Arabic-speaking adolescents.
Tasks
Narrative Retelling Task
Participants were presented with two of Aesop’s fables: The Fox and the Crow and The Dog in the Manger (see Appendix A for full texts). Throughout this article, they will be referred to as “Fox” and “Dog,” respectively. After listening to a story, the participant was asked to retell it in their own words.
Critical Thinking Task
Immediately following each retelling, participants completed a critical thinking task originally designed to assess their understanding of the fable’s deeper meanings and moral implications, as well as their ability to engage in reflective reasoning (Nippold et al., 2017, 2020; Shirley et al., 2024). In the present study, these questions required participants to attribute beliefs, intentions, and emotions to story characters—abilities conceptually related to ToM. Accordingly, the task was analyzed across three dimensions of comprehension—used as critical thinking subscales: (a) inferential understanding of story events, (b) mental states attribution to characters, and (c) moral reasoning about the fable’s lessons. The task consisted of eight structured questions for each fable, adapted from Nippold and Marr (2022), Nippold et al. (2015), and Shirley et al. (2024). These questions were grouped into three subscale categories, as outlined below.
1.
This category assessed participants’ ability to draw logical inferences from story events, targeting causal reasoning beyond surface-level recall:
• “Fox” (Q1): “Why did the crow take her stolen cheese to the branch of a tall tree?”
• “Dog” (Q2): “Why did the dog get so angry at the cattle?”
This category assessed participants’ ability to attribute beliefs, intentions, and emotions to characters—a key component of ToM (Astington & Barriault, 2001; Mitchell & Phillips, 2015; Saxe & Houlihan, 2017). These items probed both first-order ToM (attributing a character’s belief or emotion) and second-order ToM (inferring what one character thought about another’s thoughts or feelings):
• “Fox”: ○ (Q2) “Why did the fox tell the crow that she was beautiful?” ○ (Q3) “How did the crow feel when the fox paid her so many compliments?” ○ (Q4) “How did the crow feel when she dropped her piece of cheese?”
• “Dog”: ○ (Q1) “What was the dog thinking when he saw the cattle come into the barn?” ○ (Q3) “What did the cattle think when the dog snapped and snarled at them?” ○ (Q4) “What was the farmer thinking when he saw the dog’s reaction? How do you know?”
The final category assessed participants’ ability to interpret, evaluate, and apply the moral lessons of the stories. These items required participants not only to articulate their agreement or disagreement with the moral but also to provide reasons and real-life applications, reflecting higher-order reasoning skills (Astington & Olson, 1990; Nippold et al., 2015; Westby, 2016, 2021):
• “Fox”: ○ (Q5) “Is deception ever a good thing? Why or why not?” ○ (Q6) “Do you agree or disagree with the moral of this story, ‘Beware of flatterers’?” ○ (Q7) “Why do you agree (or disagree)?” ○ (Q8) “Can you think of a situation in real life where the moral might apply?”
• “Dog”: ○ (Q5) “Why do you think some people are selfish? Is selfishness ever a good thing, or is it always a bad thing?” ○ (Q6) “Do you agree with the moral of this story, ‘Do not begrudge others what you cannot enjoy yourself’?” ○ (Q7) “Why do you agree (or disagree)?” ○ (Q8) “Can you think of a situation in real life where the moral might apply?”
Analyses
The analyses were designed to address the three main research purposes: (a) examining differences in the frequency and diversity of MCVs and MLVs (as proportions of total number of words [TNW]) across narrative retelling and critical thinking tasks; (b) exploring associations between the use of MCVs and MLVs and syntactic complexity in both tasks, and macrostructure in narrative retelling; and (c) assessing whether the use of MCVs and MLVs is linked to accuracy on critical thinking.
All retellings and critical thinking responses were transcribed verbatim and segmented into communication units (C-units). Measures were first computed separately for each fable (i.e. “Fox” and “Dog”) and then averaged within each task to yield task-specific scores for narrative retelling and critical thinking. Importantly, indices were not averaged across tasks; rather, each task was analyzed independently to preserve task-specific variation.
Thus, for syntactic complexity, MLCU and Clausal Density (CD) were calculated separately for each fable, averaged within each task, and subsequently analyzed as distinct narrative retelling and critical thinking variables.
Comparable linguistic indices (syntactic complexity, MCV/MLV proportions, and diversity) were derived from both tasks. Narrative retelling provides a relatively structured discourse context requiring event sequencing and story grammar organization, whereas the critical thinking task elicits reflective, evaluative, and inferential reasoning. Computing parallel indices across tasks allowed us to test task sensitivity (within-participant shifts) in MCV/MLV use and syntactic organization. Macrostructure was assessed only in narrative retelling because story grammar is not applicable to the question-based critical thinking format. Analyses included paired-sample t tests (with false-discovery-rate correction) and Pearson correlations. To address concerns regarding multiple comparisons and to allow evaluation of evidence for both the alternative and null hypotheses, Bayesian analyses were also conducted. Bayes Factors (BF₁₀ and BF₀₁) were used to quantify the strength of evidence for the presence and absence of associations, respectively. For each task, transcripts were analyzed for the following:
Productivity (Control Variable)
Syntactic Complexity
Two indices were calculated for each task:
○
○
Metacognitive and Metalinguistic Verbs
Within each transcript, occurrences of MCVs—verbs denoting mental activities such as think, know, and believe—and MLVs—verbs denoting language-related activities such as say, explain, and write—were identified and counted. For each participant and task, four indices were computed:
Macrostructure
Macrostructure was assessed using a story grammar framework adapted from Gillam et al. (2017). Each narrative was coded for the presence or absence of seven canonical components, with each element scored as 0 = absent or 1 = present:
A total macrostructure score was calculated as the sum of elements present, with a maximum possible score of 7. Higher scores reflected more complete and coherent narrative organization.
Critical Thinking
Responses to the critical thinking questions were scored using a binary rubric adapted from Shirley et al. (2024). For Questions 1–8, with the exception of Question 6, correct responses were coded as 1, and incorrect, unclear, irrelevant, or missing responses were coded as 0. A response was considered correct if it provided a reasonable or relevant answer. Question 6 (“Do you agree with the moral?”) was excluded from the scoring system, as it required only a yes/no/unsure response. However, it was included descriptively to examine patterns of agreement and served as a precursor to Question 7, which required justification of the moral stance.
Thus, 7 of the 8 items per fable contributed to the critical thinking total score, yielding a maximum of 7 points per fable. With 2 fables administered, the maximum possible critical thinking total score was 14 points.
To capture different dimensions of comprehension, the critical thinking task was analyzed across three critical thinking subscales:
•
•
•
Composite critical thinking scores were calculated for each subscale as well as for the overall critical thinking total score.
Procedure
The instructional and assessment procedures were adapted from previous studies (Nippold et al., 2014, 2015, 2017, 2020). The interviews were conducted via Zoom. The interviewer (the first author, an experienced speech-language pathologist) introduced the task using the following standardized script, spoken slowly and clearly: This is a storytelling activity that involves fables. Fables are imaginary stories about animals, objects, and other creatures that act like people. I am going to read you a fable. Please listen carefully and be ready to tell the story back to me in your own words. Try to remember as much as you can so that you can tell the whole story. After you finish, I will ask you some questions about the story. There are no penalties for incorrect answers; I just want to know what you think about the stories. Are you ready?
The interviewer then paused to give the participant time to respond and address any questions before continuing. To control for potential order effects, the presentation order of the fables was counterbalanced across participants: approximately half of the participants were first presented with “Fox,” while the other half began with “Dog.”
Next, the interviewer presented the story text in MSA, accompanied by an illustration displayed on a screen. Prior to administration, participants were asked whether they were familiar with either of the fables. Students who reported prior familiarity were excluded from participation. The final sample of 42 participants therefore included only adolescents who indicated that they had not previously encountered the texts. The story was read aloud slowly and clearly in MSA, while the participant followed along using the text on the screen. The interviewer directed the participant’s attention to the illustration to aid comprehension. Although the stories were presented in MSA, participants were not instructed to respond in a particular language variety. They were invited to retell the story “in their own words,” and no guidance was provided regarding register choice. All instructions and critical thinking questions were delivered in SA to ensure comprehension and minimize performance pressure.
After reading the story, the interviewer confirmed the participant’s readiness to proceed and explained that the session would be audio-recorded solely for analysis purposes. Verbal consent to record was obtained before activating the audio recorder. The participant was then asked to retell the story in their own words, with only the illustration shown and the text removed from view.
Once the retelling was complete, the interviewer asked a set of predetermined critical thinking questions related to the story. Participants were given sufficient time to consider their responses and were encouraged to answer thoughtfully, without pressure.
Reliability
For reliability purposes, an independent judge (a trained third-year undergraduate student) scored all 84 narratives. Interrater reliability was high across all measures:
•
•
•
•
•
•
•
All correlations were significant at p < .001, indicating strong consistency between raters.
Results
Narrative level indices—TNW, MLCU, clausal density (CD), MCV frequency proportion (MCV tokens divided by TNW), MCV diversity proportion (distinct MCV types divided by TNW), MLV frequency proportion (MLV tokens divided by TNW), MLV diversity proportion (distinct MLV divided by TNW), and macrostructure—were computed separately for each fable (“Fox” and “Dog”) and then averaged to yield one narrative retelling score per participant. For the critical thinking task, the equivalent indices were computed for each fable and then averaged across the two fables to yield one critical thinking score per participant; macrostructure was not computed for the critical thinking task.
In addition, three critical thinking subscales were created: Inferential Comprehension (“Fox” Q1 + “Dog” Q2), Mental states attribution (“Fox” Q2–Q4 + “Dog” Q1, Q3, Q4), and Moral Reasoning (“Fox” Q5, Q7, Q8 + “Dog” Q5, Q7, Q8). A critical thinking total score per narrative (0–7), as well as a combined total across narratives (0–14) and an averaged total scaled to 0–7, were also computed. Analyses included paired-sample t tests (with false-discovery-rate correction) and Pearson correlations. Effect sizes are reported as Cohen’s d(z) for within-subjects (paired) comparisons, computed as t/∙n. Magnitude labels follow conventional guidelines (≈0.20 small, ≈0.50 medium, ≈0.80 large). Descriptive data are presented in Table 1. Macrostructure scores were relatively high and clustered near the upper end of the scale (see Table 1), suggesting limited variability in this measure.
Extended Descriptive Statistics for Narrative Retelling, Critical Thinking Task, and Critical Thinking Responses.
Note. Values represent means with standard deviations in parentheses. Critical thinking subscales represent summed accuracy scores across the respective items. CT = Critical thinking; TNW = Total number of words; MLCU = Mean length of C-unit; CD = Clausal density; MCV = Metacognitive verbs; MLV = Metalinguistic verbs.
MCVs and MLVs were relatively infrequent in both discourse contexts. Across tasks, MCV proportions ranged approximately from 0.07 to 0.11 of total words, whereas MLV proportions were lower, ranging approximately from 0.01 to 0.03 of total words. Thus, even statistically reliable task differences should be interpreted as differences in relatively low frequency but theoretically meaningful lexical-discourse markers.
Hypothesis 1. Narrative Retelling versus Critical Thinking
Compared with narrative retelling, the critical thinking task elicited significantly higher MCV frequency proportion (M = 0.109, SD = 0.030) than narrative retelling (M = 0.071, SD = 0.025), t(41) = –7.04, p < .001, d(z) = 1.09 (large), 95% confidence interval [CI] [0.027, 0.049]. The Bayesian paired comparison provided decisive evidence for this task effect (BF10 = 2.57 × 106).
Likewise, MCV diversity proportion was higher in the critical thinking task (M = 0.065, SD = 0.019) than in narrative retelling (M = 0.054, SD = 0.021), t(41) = 2.96, p = .005, d(z) = 0.46 (medium), 95% CI [0.003, 0.018], with moderate Bayesian evidence for the effect (BF10 = 9.12).
By contrast, narrative retelling yielded significantly higher MLV frequency proportion (M = 0.028, SD = 0.013) than critical thinking (M = 0.012, SD = 0.007), t(41) = 7.41, p < .001, d(z) = 1.14 (large), 95% CI [0.012, 0.020], with decisive Bayesian evidence (BF10 = 8.79 × 106).
Similarly, MLV diversity proportion was higher in narrative retelling (M = 0.019, SD = 0.009) than in critical thinking (M = 0.007, SD = 0.004), t(41) = −7.68, p < .001, d(z) = 1.18 (large), 95% CI [0.009, 0.015], again with decisive Bayesian evidence (BF10 = 2.06 × 107).
These findings supported H1 and indicate a clear task-sensitive dissociation: critical thinking recruited MCVs more strongly, whereas narrative retelling recruited MLVs more strongly.
Correlational Analyses
Table 2 reports Pearson correlations among narrative retelling (superscript N) and critical thinking (superscript CT) measures—including TNW, MLCU, clausal density (CD), MCV and MLV indices MCV and MLV, as well as critical thinking scores (CT Total, CT Mental States, CT Inferential, and CT Moral).
Correlation Matrix of Narrative Retelling and Critical Thinking Measures.
Note. Pearson correlations are reported (upper triangle only). N = 42. Significance markers reflect uncorrected p-values (*p < .05. **p < .01. ***p < .001). Bonferroni-adjusted p-values and Bayesian analyses are reported in the Results section. CT = Critical thinking; superscript N = Narrative retelling; TNW = Total number of words; MLCU = Mean length of C-unit; CD = Clausal density; MCV = Metacognitive verbs; MLV = Metalinguistic verbs.
Given the number of correlations examined, results should be interpreted with caution. To address the risk of Type I error, Bonferroni-adjusted p values were calculated for the family of correlations reported in Table 2. In addition, Bayesian analyses were conducted to evaluate evidence for both the presence and absence of associations. Bayes Factors (BF₁₀ and BF₀₁) were used to quantify evidence for the alternative and null hypotheses, respectively, allowing a distinction between lack of evidence and evidence supporting the null hypothesis.
For correlational analyses, Bayesian Pearson correlations were computed, whereas for paired task comparisons, Bayes Factors (BF₁₀) were used to quantify evidence for task differences. Following conventional interpretive guidelines, Bayes Factors around 1 were interpreted as inconclusive, values of approximately 3 to 10 as moderate evidence, and values above 10 as strong evidence.
Hypothesis 2 (MCVs and MLVs and Syntactic Complexity)
Consistent with the hypothesis, which focused on frequency proportions, clausal density in narrative retelling (CDN) showed a positive uncorrected association with MCV frequency proportion in narrative retelling (MCVN) (r = 0.35, p = .024); however, this effect did not survive Bonferroni correction (p = .401), and Bayesian analysis indicated only anecdotal evidence for the association (BF₁₀ = 1.53).
Additionally, CDN was not significantly associated with MLV frequency proportion in narrative retelling (MLVN) (r = 0.17, p = .289; Bonferroni-adjusted p = 1.000), and Bayesian analysis provided moderate evidence in favor of the null hypothesis (BF₀₁ = 4.75), suggesting that an association between MLV frequency and clausal density is unlikely.
Exploratory analyses further examined diversity measures. CDN showed an uncorrected association with MLV diversity proportion in narrative retelling (MLVdivN) (r = 0.31, p = .044); however, this effect did not remain significant after correction (p = .746), and Bayesian results were inconclusive (BF₁₀ = 0.90; BF₀₁ = 1.11), providing no clear evidence for either the presence or absence of an association.
Associations involving mean length of C-unit in narrative retelling (MLCUN) were not statistically significant, and Bayesian analyses provided moderate evidence for the null hypothesis for both MCV frequency proportion (BF₀₁ = 7.81) and MLV frequency proportion (BF₀₁ = 7.04).
Taken together, these findings provide only limited support for Hypothesis 2, with a weak and non-robust tendency linking clausal density and MCV frequency proportion, and converging evidence indicating no reliable association between MLV frequency proportion and syntactic complexity.
Hypothesis 3 (MCVs and MLVs and Macrostructure)
Macrostructure scores clustered near ceiling (M = 5.86 out of 7, SD = 0.97), indicating restricted variability. Macrostructure did not correlate significantly with narrative MCV or MLV frequency proportions (MCV: r = −0.10, p = .513, BF01 = 6.72; MLV: r = 0.06, p = .696, BF01 = 7.70). Bayesian results therefore provided moderate evidence for the null hypothesis for the frequency-proportion associations. Associations with diversity proportions were also not significant and yielded weak-to-moderate evidence for the null (MCV diversity: r = −0.23, p = .137, BF01 = 2.77; MLV diversity: r = −0.22, p = .156, BF01 = 3.06). Accordingly, the predicted positive association between MCV/MLV proportions and macrostructure was not observed; however, this result should be interpreted in light of the restricted range of macrostructure scores.
Hypothesis 4 (MCVs and MLVs and Critical Thinking)
During the critical thinking task, inferential comprehension showed an uncorrected positive association with MCV frequency proportion (r = 0.34, p = .026), Bonferroni-adjusted p = .445; however, Bayesian evidence for the alternative hypothesis was weak/anecdotal (BF10 = 1.40). Moral reasoning showed an uncorrected negative association with MCV diversity proportion during critical thinking (r = –0.31, p = .043), Bonferroni-adjusted p = .724, but the Bayesian analysis was inconclusive (BF10 = 0.93; BF01 = 1.08). Moral reasoning also showed an uncorrected positive association with MLV frequency proportion during critical thinking (r = 0.32, p = .037), Bonferroni-adjusted p = .631, again with inconclusive Bayesian evidence (BF10 = 1.04; BF01 = 0.96). Mental-states attribution was not associated with MCV or MLV frequency proportions during critical thinking, with Bayesian evidence supporting the null (MCV: BF01 = 8.31; MLV: BF01 = 8.12). These results indicate that H4 received only tentative and component-specific support in the frequentist analyses, whereas Bayesian analyses suggest that several non-associations—particularly for mental-states attribution—are supported by the data.
Register Distribution (Qualitative Observation)
To illustrate how diglossia manifested in the dataset, we conducted a qualitative inspection of register use in participants’ responses. Although the fables were presented in MSA, nearly all retellings were produced in SA. The two participants who initially began retelling in MSA shifted to SA within the first few utterances.
MSA forms appeared primarily at the lexical level rather than as full syntactic alternations. Examples included phonologically MSA realizations such as الثعلب (with interdental /θ/), lexical items such as , , , الغنائم, and abstract evaluative expressions such as بالغرور والسعادة and بالخيبة. These insertions clustered around descriptive and evaluative meanings.
Summary of the Main Results
Task Differences (Narrative Retelling vs. Critical Thinking)
The critical thinking task showed a higher MCV frequency proportion than the narrative retelling task (large effect), and higher MCV diversity (medium effect).
Narrative retelling task yielded higher MLV frequency and diversity than the critical thinking task (both large effects).
Thus, the prediction that critical thinking would elicit denser MCV use than narrative retelling was supported, whereas MLVs patterned in the opposite direction, appearing more often in narrative retelling.
Associations with Syntactic Complexity and Macrostructure
Within narrative retelling, clausal density (CDN) showed a weak, non-robust association with MCV frequency proportionN, which did not survive correction for multiple comparisons and was supported only by anecdotal Bayesian evidence.
No reliable association was observed between CDN and MLV frequency proportionN, with Bayesian analyses providing moderate evidence in favor of the null hypothesis.
No significant association was observed between CDN and MCV diversity proportionN.
The association between CDN and MLV diversity proportionN reached significance at the uncorrected level but did not remain significant after correction and was supported by inconclusive Bayesian evidence; therefore, it is interpreted as exploratory and non-robust rather than as evidence supporting the hypothesized relationship.
No significant associations were observed between narrative macrostructure and MCV or MLV proportions; however, this null finding should be interpreted cautiously given the restricted variability and near-ceiling distribution of macrostructure scores.
Links to Critical Thinking
• Inferential comprehension was positively associated with MCV frequency proportionᶜᵀ (medium): greater MCV use during critical thinking aligned with better inferential accuracy.
• Moral reasoning was negatively associated with MCV diversity proportionᶜᵀ (medium negative) and positively associated with MLV frequency proportionᶜᵀ (medium).
Overall, these findings indicate that MCVs and MLVs function as task-sensitive discourse resources in adolescent language, with distinct patterns emerging across narrative retelling and critical thinking contexts.
Discussion
The current research examined how Palestinian Arabic-speaking adolescents use MCVs and MLVs across narrative retelling and critical thinking tasks, and how these measures relate to syntactic complexity, narrative macrostructure, and critical thinking. The strongest finding concerned task sensitivity. Consistent with Hypothesis 1, critical thinking prompts elicited higher MCV frequency proportion and diversity proportion, whereas narrative retelling favored MLV frequency proportion and diversity proportion; Bayesian analyses provided decisive evidence for these task effects. In contrast, the correlational findings were more limited. After correction for multiple comparisons, none of the Pearson correlations remained statistically significant, and Bayesian analyses indicated that several non-associations—particularly those involving MLCU, macrostructure, and mental-states attribution—were better interpreted as evidence leaning toward the null rather than merely as failures to detect effects.
Task Effects on the Use of MCV and MLV: Critical Thinking Task Draws MCVs; Narrative Retelling Draws MLVs
A central finding was the task-specific recruitment of MCVs and MLVs: the critical thinking task elicited higher MCV frequency proportion and diversity proportion, whereas narrative retellings elicited higher MLV frequency and diversity. These results support the view that task demands recruit specific verb classes depending on the type of reasoning required. This pattern accords with prior work showing that critical thinking prompts—particularly those that elicit stance, require justification, and invite moral evaluation—mobilize a metacognitive lexicon as adolescents articulate beliefs, appraise motives, weigh alternatives, and defend claims (e.g. think, know, believe, doubt, infer, decide, agree/disagree). In the current study, these functions were explicitly targeted by items such as: Stance: “Do you agree with the moral . . . ?” (Q6, both fables). Example (Dog, Q6): “For example—say, a bottle of water. I have a bottle of water, but I don’t want it [MCV—desire]; he wants it [MCV—desire], yet I’m not willing to give it to him [MCV—intention/volition].”
These prompts encourage epistemic stance and evaluative reasoning, which naturally increase MCV tokens and types (e.g. I think / I believe / he decided / they trusted), often realized in complement clause constructions (e.g. I think that . . ., He believed that . . . ). In this sense, MCV use functions as a linguistic index of evaluative and inferential processing rather than as a direct measure of cognitive ability per se.
By contrast, narrative retelling scaffolds MLVs: speakers manage reported speech and discourse representation to track who said what to whom and to maintain episode structure (e.g. say, tell, ask, explain, answer). This helps account for the higher MLV frequency proportion and diversity proportion in narrative retelling relative to critical thinking. Together, these findings support Hypothesis 1 and reinforce the interpretation of MCVs and MLVs as task-sensitive discourse resources.
At the same time, it is important to contextualize these findings by noting that the overall proportions of MCVs and MLVs were relatively low across both tasks. MCVs accounted for only a modest proportion of total words (up to approximately 11%), while MLVs were even less frequent (generally below 3%). This pattern is consistent with prior research indicating that mental-state and MLVs typically constitute a small proportion of total word use in discourse, even among older children and adolescents (Nippold et al., 2017; Westby, 2016).
Importantly, these relatively low proportions should not be interpreted as reflecting limited use. Rather, they likely reflect the functional specificity of such verbs: MCVs and MLVs are recruited selectively at points in discourse that require evaluation, perspective-taking, explanation, or stance marking, rather than being distributed uniformly across utterances. Consequently, even low-frequency occurrences may carry disproportionate weight in signaling higher-level cognitive and discourse processes.
MCVs and MLVs, Syntactic Complexity, Macrostructure and Critical Thinking
Syntactic Complexity
Only partial support emerged for the second hypothesis concerning syntactic complexity. In narrative retelling, clausal density showed a positive association with MCV frequency proportion; however, this relationship was not robust, as it did not survive correction for multiple comparisons and was supported only by anecdotal Bayesian evidence. This suggests a weak tendency for MCV use to co-occur with greater clausal embedding, consistent with their role in complement clause constructions (e.g. “he thought that . . . “).
In contrast, MLV frequency proportion was not associated with clausal density, and Bayesian analyses provided moderate evidence in favor of the null hypothesis, indicating that such an association is unlikely.
Exploratory analyses revealed that clausal density showed an uncorrected association with MLV diversity proportion; however, this effect did not remain significant after correction and was supported by inconclusive Bayesian evidence. Accordingly, this association is interpreted cautiously as non-robust and exploratory rather than as evidence supporting the hypothesis.
Importantly, this pattern differs from earlier findings reported in English-speaking samples, where MCV use has been linked more consistently to syntactic complexity in written language (Nippold et al., 2017; Sun & Nippold, 2012). Several factors may account for this divergence. First, the present study examined Palestinian Arabic-speaking adolescents operating within a diglossic linguistic system, whereas prior studies focused on English. Cross-linguistic differences in clause structure and discourse organization may influence how mental-state verbs are integrated into syntactic constructions.
Second, whereas prior work examined written narratives, the present study analyzed oral narrative and critical thinking discourse. Oral production is more constrained by real-time processing demands and may rely more on discourse-pragmatic strategies than on syntactic elaboration. This distinction is consistent with broader evidence showing that different discourse genres (e.g. conversational vs. narrative) place different demands on linguistic complexity.
Third, narrative tasks themselves vary in the extent to which they elicit complex syntax. For example, fable-based tasks can not only support complex language production but also constrain responses through shared story structure and prompts. In the present study, such task constraints may have limited the variability in syntactic constructions, thereby weakening the association between verb use and syntactic complexity.
Taken together, these cross-linguistic, modality-related, and task-related considerations highlight that the relationship between mental-state language and syntactic complexity may not generalize uniformly across contexts.
However, associations involving MLCU were not statistically significant, and Bayesian analyses provided moderate evidence for the absence of association, suggesting that verb use may relate more specifically to clause packaging (density) than to overall utterance length.
Taken as a whole, these findings indicate that the relationship between verb use and syntactic complexity is selective and limited, applying primarily to MCV frequency and clausal density, rather than extending to MLV use or to syntactic complexity more broadly. Importantly, the Bayesian analyses allow us to conclude that some expected associations—particularly those involving MLV frequency and MLCU—are not merely undetected but are likely absent or weak in this dataset.
Macrostructure
Contrary to Hypothesis 3, macrostructure did not correlate significantly with narrative MCV/MLV proportions. However, this finding should be interpreted cautiously given the restricted variability and near-ceiling performance observed in macrostructure scores. The clustering of scores near the upper end of the scale suggests that many participants produced relatively complete story structures, which may have limited the sensitivity of the measure to detect associations with verb use.
Accordingly, the absence of significant correlations may reflect a measurement constraint rather than a true absence of relationship between macrostructure and MCV/MLV use. This pattern suggests that adolescents may achieve relatively complete story grammar organization without necessarily increasing their use of MCV or MLV. Macrostructure, as coded here, reflects inclusion of canonical story grammar elements (characters, setting, problem, attempts, consequence) and may be driven primarily by event sequencing and narrative planning rather than by explicit mental-state or speech-reporting language.
Critical Thinking
Hypothesis 4 received only partial and component-specific support. Inferential comprehension showed a positive association with MCV frequency proportion during the critical thinking task, consistent with the idea that MCVs support epistemic stance and causal explanation. However, this association did not survive correction for multiple comparisons, and Bayesian evidence was only anecdotal, indicating that the effect should be interpreted cautiously.
Similarly, the associations involving moral reasoning—specifically, the positive association with MLV frequency proportion and the negative association with MCV diversity proportion—were significant at the uncorrected level but yielded inconclusive Bayesian evidence, suggesting that these effects are tentative rather than robust.
By contrast, Bayesian analyses provided clearer evidence in favor of the null hypothesis for associations involving mental states attribution and MCV/MLV frequency proportions, indicating that these relationships are likely absent rather than merely undetected.
Taken together, these findings indicate that associations between verb use and critical thinking are selective rather than global. MCV use appears to contribute specifically to inferential reasoning, whereas moral reasoning reflects a different linguistic profile.
In particular, moral reasoning was positively associated with MLV frequency proportion and negatively associated with MCV diversity proportion during the critical thinking task. This suggests that moral reasoning responses may rely more on extended explanation and discourse framing (including speech/reporting verbs such as “say,” “tell,” and “explain”) than on lexical diversity in MCVs. One interpretation is that once adolescents articulate a moral stance, they elaborate using discourse moves that reference what characters “said” or what one “should say or do,” rather than expanding the range of MCV types.
Overall, these findings refine Hypothesis 4 by showing that MCV and MLV use relates differently to distinct components of critical thinking, and that these relationships are modest and context-dependent rather than broad or consistently robust.
Sociolinguistic Context: Diglossia and Lexical Layering
The Palestinian Arabic context is diglossic (SA in daily life; MSA in schooling), which likely shapes access to abstract MCVs and MLVs. Schooling in MSA may enrich the metacognitive lexicon, while performance in SA during oral tasks introduces register-bridging demands (Habib, 2022).
As reported in the Results, although the fables were presented in MSA, nearly all participants retold the narratives in SA, and MSA forms appeared primarily at the lexical level. This pattern suggests that SA remains the dominant frame for spontaneous discourse, while MSA lexical items may be selectively activated.
Cross-cultural literature suggests that socialization emphases may modulate mental-state talk (Shahaeian et al., 2011; Wellman & Liu, 2004; Westby, 2016); however, the present findings are best interpreted as reflecting task affordances and sociolinguistic layering rather than direct evidence of cross-cultural differences in ToM development. In this study, ToM is treated as a theoretical backdrop that informs mental-state language use, not as an outcome variable.
Earlier Arabic findings on ISTs and evaluative devices (Kawar et al., 2023, 2024; Westerveld et al., 2023) foreshadow the adolescent patterns observed here, reinforcing continuity from early ISTs to later, more abstract MCVs. For example, research on Arabic-speaking preschool children demonstrated that despite exposure to MSA narratives, children overwhelmingly retold stories in their spoken vernacular (PA), while selectively incorporating MSA lexical items when encoding internal states (Kawar et al., 2023; Ravid et al., 2014). The adolescent pattern observed here extends this trajectory into later developmental stages.
Importantly, the MSA lexical insertions observed in the present study clustered around abstract, evaluative, and mental-state meanings (e.g. اعتقد “believed,” يتفاخر “boasts,” أجنحته لامعة “his wings shining”), suggesting that MSA may function as a lexical reservoir for epistemic and evaluative vocabulary within otherwise SA discourse.
Unlike monolingual English contexts, Palestinian Arabic-speaking adolescents operate within a diglossic system in which abstract evaluative vocabulary is primarily stabilized through MSA exposure in schooling yet used within spoken discourse. Thus, MCV use in Arabic reflects not only cognitive reasoning but also active navigation between linguistic registers. The present findings indicate that task demands may differentially activate this layered lexicon, highlighting how MCV production in Arabic is shaped by sociolinguistic structure in addition to developmental maturation.
Overall, these findings position MCV use in Arabic as a discourse-level phenomenon emerging at the intersection of cognitive development, narrative organization, and diglossic language structure, while also reflecting task-sensitive and context-dependent patterns observed in the present study.
Register Distribution of Verb Forms
Although a full quantitative register analysis was beyond the scope of the present study, qualitative inspection of transcripts revealed systematic patterns in the distribution of SA and MSA forms. Nearly all retellings were produced in SA, even though the fables were presented in MSA. MSA forms appeared primarily at the lexical level rather than as full syntactic shifts, suggesting lexical insertion rather than structural code-switching.
Examples included phonological MSA realizations such as الثعلب (with interdental /θ/), lexical items such as , , , , أجنحته لامعة, and abstract evaluative expressions such as بالغرور والسعادة and بالخيبة. Notably, many of these MSA insertions clustered around evaluative, descriptive, or abstract meanings. This pattern suggests that exposure to the MSA narrative input may have temporarily activated MSA lexical representations, particularly for abstract or literary vocabulary, while the overall discourse frame remained in SA.
These observations indicate that diglossia in the present dataset manifested primarily as lexical-level register layering influenced by immediate input, rather than as the use of distinct SA and MSA grammatical systems.
Educational and Clinical Implications
The findings suggest that instructional tasks can be deliberately shaped to elicit the kinds of language associated with stronger critical thinking. Because the critical thinking task in this study drew relatively more MCVs, classroom and clinical activities that ask students to take a stance, justify a position, and evaluate a moral are likely to encourage explicit references to beliefs, intentions, and evaluations (e.g. I think/ believe/ decide), thereby supporting clearer ToM articulation. Because coherent story grammar supports causal and inferential reasoning, teaching the components of narrative macrostructure (characters, setting, problem, internal response, attempts, and consequences) may indirectly support students’ ability to explain events and generate inferences during discussion tasks. In practical terms, modeling a brief narrative plan before discussion, or prompting students to identify the problem–response–consequence chain in a fable, may facilitate both the organization of retellings and the quality of subsequent critical thinking.
Given the Fox–Dog pattern in the critical thinking task, educators should note that encouraging longer responses (as seen with Dog) does not guarantee richer mental state language or better accuracy in critical thinking; targeted prompts that invite stance and belief-ascription (as in Fox) may yield more efficient gains in the use of MCVs and MLVs and in critical thinking.
These implications are particularly relevant in a diglossic context. Although students produced their responses in SA, they sometimes recruited MSA verbs when expressing metacognitive and communicative meanings. Short, explicit bridges between MSA and SA—such as presenting an MSA stem for “I think that . . . “ and inviting an SA completion—can reduce retrieval effort for abstract verbs and make it easier for students to use MCVs and MLVs during oral tasks. Clinicians and teachers can also model sentence frames that align with the tasks used here (e.g. “I think that X believed . . . because . . . “), then fade supports as students begin to produce these forms independently.
Finally, for assessment and progress monitoring, the results indicate that indices tied to the use of MCVs and MLVs and macrostructure are informative alongside traditional syntactic metrics. Importantly, consistent with the results, effective critical thinking responses were observed without robust or consistent associations with syntactic complexity measures (e.g. clausal density), highlighting the selective and context-dependent nature of these relationships. Educators and clinicians may wish to track MCVs and MLVs frequency proportion and diversity proportion, and to document growth in macrostructure, in addition to reporting MLCU and clausal density. Aligning instruction with these measures—by designing prompts that elicit MCVs and MLVs, by teaching macrostructure explicitly, and by bridging MSA–SA usage—offers a coherent pathway to support adolescents’ critical thinking, ToM-related reasoning, and inference drawing within the same instructional sequence.
Limitations and Future Research
The sample in the present study skewed toward medium-to-high socioeconomic status and interviews were conducted online via Zoom; replication with broader SES bands and in-person administration would help assess ecological validity and potential modality effects.
In addition, given the modest sample size, the correlational findings should be interpreted cautiously. Although Bayesian analyses allowed evaluation of evidence for both the presence and absence of associations, the modest sample size may limit the precision and stability of Bayes Factor estimates. The present analyses were exploratory in nature, and replication with larger samples is needed to clarify the stability and strength of the observed associations between verb use and critical thinking performance. Replication with larger samples is therefore essential to confirm the robustness of these findings.
An additional limitation concerns the absence of standardized linguistic screening measures (e.g. vocabulary or syntax assessments). Because MCV and MLV use is closely tied to academic language proficiency, reliance on school and parental reports may not fully capture variability in language ability within the sample. Future research should incorporate standardized assessments to better characterize linguistic profiles and examine how verb use patterns vary across ability levels.
Although diglossia provided an important sociolinguistic context for interpreting the findings, the present study did not systematically quantify register distribution (SA vs. MSA) across verb tokens or examine register choice as an independent analytic variable. The qualitative observations reported here suggest lexical-level layering influenced by task demands and input language; however, future research should employ explicit coding schemes to analyze register form, code-switching patterns, and task-based variation in SA and MSA verb use. Such analyses would clarify how diglossia interacts with discourse-level reasoning and lexical access in Arabic-speaking adolescents.
Future research should broaden the range of discourse contexts by incorporating additional genres (e.g. argumentative and expository tasks), a wider set of topics, and culturally varied narratives. Finally, intervention studies are warranted to test whether instruction that explicitly targets macrostructure, combined with MCV and MLV modeling (e.g. sentence frames and think-alouds), improves critical thinking and inferential comprehension.
Conclusion
Overall, this study advances understanding of how MCV and MLV signal the integration of language, cognition, and discourse during adolescence. By contrasting narrative and critical thinking contexts, the findings reveal that Arabic-speaking adolescents modulate their mental state language according to communicative purpose: using MLVs to construct story dialogue and MCVs to articulate reasoning.
Consistent with the results, these relationships are selective and context-dependent rather than uniformly robust across all components of critical thinking. Rather than positioning the use of MCVs and MLVs as a standalone predictor of cognition, the results suggest that verb use, discourse organization, and task demands interact to shape critical thinking performance within a diglossic linguistic environment.
Footnotes
Appendix A
Ethical Considerations
We hereby attest that the submitted manuscript contains original, unpublished work and is not under consideration for publication elsewhere. The study was approved by the Beit Berl College Ethics Committee (Approval No. 150-2025), and all procedures adhered to established ethical guidelines for research with human participants.
Author Contributions
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
Data supporting the findings of this study are not available due to privacy and ethical restrictions.
