Abstract
Contemporary research on language and communication has expanded beyond its traditional focus on spoken and written forms to encompass signing, gestures, facial expressions, and other bodily actions. This shift has been accompanied by methodological advancements that extend beyond classical tools, such as tape recorders or video cameras, and include motion-tracking systems, depth cameras, and multimodal data-fusion techniques. Although these tools enable richer empirical insights, they also introduce significant conceptual and practical challenges, particularly for researchers new to multimodal data collection. In this article, we present a structured, decision-oriented workflow for multimodal data collection in language and communication research. We introduce a flexible framework that guides researchers through key methodological choices, including the alignment of research questions with data streams; study-design and acquisition strategies; synchronization and technical requirements; ethical governance; and data management, dissemination, and reuse. The framework is illustrated with case studies spanning controlled laboratory experiments, large-scale annotated sign-language corpora, and field-based research, including nonhuman primates. Rather than advocating a one-size-fits-all approach, in our discussion, we emphasize key decision points, trade-offs, and real-world examples to help researchers navigate the complexities of multimodal data collection. By integrating perspectives from different disciplines, our flexible decision-making framework is intended as a practical tool for researchers seeking to design, implement, and address common conceptual and methodological challenges in the rapidly developing area of multimodal data collection.
Keywords
Over the past decades, researchers across linguistics, psychology, neuroscience, and related fields have increasingly embraced the notion that language and communication are fundamentally multimodal and not simply accompanied by optional nonvocal cues (Abner et al., 2015; Bavelas, 1990; Hagoort & Özyürek, 2025; Holler & Levinson, 2019; Keevallik, 2018; Kendon, 2004; Kendrick et al., 2023; Mondada, 2016; Özyürek, 2021; Perniss, 2018; Rasenberg et al., 2022; Sandler, 2024; Trettenbrein et al., 2025; Vigliocco et al., 2014). Rather than relying solely on speech, manual signs, or text, real-world language use and communicative interactions involve gestures, facial expressions, body posture, gaze, or tactile and haptic cues (Checchetto et al., 2018; Edwards & Brentari, 2020). Although research in the language sciences has traditionally focused on humans, in the last decade, several studies have started exploring parallel phenomena in nonhuman primates, which appear to also frequently integrate communicative signals across modalities to enhance meaning, improve signal effectiveness, and adapt to diverse social and environmental contexts (Aychet et al., 2021; Fröhlich et al., 2019, 2021; Hobaiter et al., 2017; Liebal et al., 2013; Mine et al., 2024; Wilke et al., 2017). This increasing cross-disciplinary interest in multimodality is directly reflected in the increasing number of publications on the topic over the past decades (see Fig. 1), presumably driven by both theoretical advancements and technological progress in capturing and analyzing multimodal phenomena. However, the concept of multimodality remains diffuse, encompassing a range of interpretations and definitions (Fröhlich et al., 2021; Sandler, 2022). Here, we distinguish between (a) multimodality in language and communication as a property of the phenomenon and (b) multimodality in data collection as a property of the measurement strategy and review these definitions in more detail below.

Publications matching the search term “multimodal language or communication” in PubMed. The increasing interest in multimodality in the language sciences is evidenced in the number of articles published referring to the concept. The histogram shows the total number of articles matching a targeted search on PubMed using the search term “((language) OR (communication)) AND ((multimodal) OR (kinaesthetic))”. Because the search yielded no relevant findings before the 1970s, the figure presents data from that period through the end of 2025, excluding partial data for the current year (2026). This bibliometric query should be read as a rough indicator of rising interest in multimodality rather than as an exhaustive operationalization of the full range of modality constellations relevant to this article.
Despite growing attention to the multimodal nature of communication, current guidelines and standards for multimodal data collection remain discipline-specific and fragmented. Previous works offer valuable insights and practical recommendations for specific fields, such as gesture or sign-language research (e.g., Fenlon & Hochgesang, 2022; Mittelberg, 2007; Müller, 2024; Orfanidou et al., 2015), including specific guidance for video-data quality and annotation (Hanke & Fenlon, 2022; Perniss, 2015). However, researchers investigating multimodal language and communication would benefit from a broader approach that incorporates more diverse populations and explicitly addresses the use of multiple data streams, such as motion capture, physiological sensors, eye tracking, and audio-visual recordings. Such multimodal data not only enable new research questions but also introduce conceptual and methodological challenges when integrating multiple sources in a single study. In addition, ongoing technological developments—such as semiautomatic annotation tools, computer-vision-based markerless motion tracking, masking tools, and methods for synchronizing and integrating multimodal data—expand the possibilities for data collection and analysis, raising further conceptual and methodological considerations. These developments create new methodological opportunities and challenges, highlighting the need for a coherent decision-making framework for researchers across disciplines that supports comprehensible, reasonable, and replicable practices for multimodal data collection.
An extensive review of the state of the art of existing research tools is available in Gregori et al. (2023); thus, we will be brief in this respect here. The goal of this article is to outline a decision-making framework for multimodal data collection based on practical experiences. The framework is intended to apply to multiple scenarios across disciplines rather than prescribing a single one-size-fits-all protocol. We introduce this approach to help researchers navigate the methodological complexities inherent in multimodal data collection involving both human and nonhuman communicative behavior. To bridge interdisciplinary gaps and refine best practices, we illustrate updated recommendations through three case studies to highlight key aspects. First, we offer structured guidelines for integrating cross-disciplinary methods, combining approaches from linguistics, psychology, and computer science. The case studies demonstrate how these different fields can collaborate effectively. Second, we address the challenges posed by new technologies, ensuring the ethical, transparent, and reproducible use of existing tools, as illustrated through the case examples. Finally, although we think that there is not a single solution to all research domains and practices because there is not one framework that could address all questions in cognitive science (Dale et al., 2009), we still want to emphasize the importance of standardizing data collection, improving metadata documentation, and refining annotation practices, all of which are demonstrated in the case studies to enhance research interoperability.
What Is “Multimodality” in the Cognitive Sciences?
The concept of multimodality can be understood from at least two different perspectives. On the one hand, it can be viewed as a fundamental property of communication or—more narrowly—as a fundamental property of language itself. From a broad perspective, multimodality is approached as a basic characteristic of communicative practice, constituted by combinations of different semiotic resources, such as speech, gesture, pictures, text, diagrams, and others (Bateman et al., 2017; Cohn & Schilperrord, 2024). From a narrower perspective, multimodality can be considered a fundamental property of language itself. Language-based communication is inherently multimodal, spanning multiple physical channels rather than being confined to a single perceptual or production system (Fricke, 2013; Hagoort & Özyürek, 2025; Holler & Levinson, 2019; Ladewig, 2020; Perniss, 2018). In this sense, the notion of modality refers to the physical channel through which language is produced or perceived. Spoken languages, for example, include the well-known vocal-auditory signals but also somatosensory signals (e.g., the touch of the two lips for producing a closure in stops, the touch of the tongue at the palate producing a fricative; Ito et al., 2009; Kent, 2024). In addition, it is very common in face-to-face interactions that language systematically involves nonvocal signals, including manual and nonmanual gestures (i.e., movements of the hands, head, torso, and face). Such movements engage both the bodily visual and somatosensory modalities (i.e., information about body position and movement; Momsen & Coulson, 2025). Signed languages primarily use the visual-kinesthetic modality and may recruit multiple articulators simultaneously such that linguistic signals are conveyed through coordinated movements of the hands, head, and facial features. In addition, signed languages may incorporate contact-induced structural elements, such as mouthing, that is, mouth movements that resemble spoken or written words of the surrounding spoken language. Such contact elements are considered part of natural signing (Bauer, 2019; Bauer & Kyuseva, 2022; Boyes-Braem & Sutton-Spence, 2001; Mohr, 2012), but they are clearly shaped and influenced by another modality (e.g., the vocal-auditory one) or have cognitive representations in another modality (Cohn & Schilperoord, 2024). Like speech, sign language also includes somatosensation or even touch as a sensory stimulus, as in tactile sign languages, which likewise span multiple perceptual systems, such as visual and tactile (Checchetto et al., 2018; Edwards & Brentari, 2020; Mesch & Raanes, 2023).
Complementing this conceptual perspective, multimodality can also be considered from a methodological standpoint, as a property of the data itself. From this second perspective, “multimodality” refers to the integration of multiple parallel data streams, instruments, measurement devices, or acquisition techniques used to capture the same phenomenon. This perspective encompasses modalities such as audio and video recordings, physiological signals (e.g., breathing patterns, heart rate, and skin conductance), electroencephalography (EEG), magnetoencephalography (MEG), eye tracking, motion capture, and other. In this context, each acquisition method constitutes a modality and yields a distinct data set. For example, investigations of the relationship between co-speech gesture and speech may combine high-quality audio and multiangle video recordings with eye-tracking data (Brône & Oben, 2015) because understanding how gaze shapes predictive coding during interactive language processing requires the linking of eye gaze and neural activity (Brilmayer et al., 2025), whereas studies exploring the forms and functions of head movements in signed languages may combine multiangle video recordings with motion-capture (Puupponen et al., 2015) or computer-vision measurements (Bauer et al., 2024). A key characteristic of multimodal data is complementarity: Each modality provides information that is not readily inferable from the others (Lahat et al., 2015). In this second perspective on multimodality, researchers speak of “multimodal corpora”—repositories that contain multiple forms of media and textual representations, such as transcriptions and annotations (Bauer et al., 2026; Knight & Adolphs, 2021).
Whereas the former view of multimodality as an intrinsic property of language and communication has its roots primarily in sign-language research (Perniss, 2018; Pfau et al., 2012; Sandler & Lillo-Martin, 2006) and gesture studies (Kendon, 1980, 2004; McNeill, 1992; Müller et al., 2013, 2014), the latter view (multimodality as a property of the data) is becoming central to empirical research practices in the language sciences and beyond. In practice, these two perspectives intersect: Acknowledging the inherently multimodal nature of language and communication, in which various signals (vocal, manual and nonmanual, somatosensory, and others) intertwine with each other, necessitates the collection of richer, multifaceted data sets (e.g., audio, video, and additional physiological or behavioral measures) to obtain a fine-grained view.
Thus, the term “multimodality” is still used ambiguously in the literature to refer both to the intrinsic organization of language and communication and to methodological choices in data collection—two conceptually distinct uses. To avoid this ambiguity and to improve conceptual clarity, we explicitly distinguish between (a) “multimodal language and communication,” referring to the inherently multimodal nature of linguistic systems and language use, and (b) “multimodal data collection,” referring to the methodological practice of integrating and aligning multiple data streams. We adopt this distinction throughout the remainder of the article.
The Workflow for Multimodal Data Collection: A Flexible Decision Framework
Throughout this article, we refer to a flexible decision framework for multimodal data collection, presented here as a diagram (Fig. 2), highlighting the important aspects involved in the process of collecting multimodal data. Although the list of factors is not exhaustive and other similarly significant decisions may arise, we believe these are the most crucial and warrant attention. Considering recent technical advances and the interconnectedness of these factors, as depicted in Figure 2, in this framework, we aim to guide researchers in making informed decisions for the collection of multimodal data in language sciences. Although the framework is organized into three sequential steps, the process is typically iterative: Later decisions can require revisiting and revising earlier choices. In other words, if a requirement emerges at a later stage that cannot be met by the current design, then the relevant choice must be revised until the workflow is consistent. For a more explicit examination of these interdependencies, see Figure S1 in the Supplemental Material available online, which provides an expanded map of cross-links between decision areas for a specific example study and illustrates how constraints in later steps can prompt revisions of earlier design choices.

A flexible decision framework for multimodal data collection. The workflow defines three steps. Step 1 defines what is studied (research question, phenomenon), who is studied (target population), and how this can be captured (study design). Step 2 (implementation) translates these choices into an executable plan and covers governance (ethics, legal regulations, constraints) and technical implementation, such as the definition of data streams to record, an analysis plan, acquisition protocols, synchronization, and data management. Step 3 (dissemination) determines data sharing and reuse. For an expanded depiction of these dependencies and example if-then decision points, see Figure S1 in the Supplemental Material available online.
How to use this framework
Figure 2 summarizes the workflow, which is organized into three steps. Step 1 (defining the research space) helps to specify what is investigated (phenomenon, units of analysis, and level of granularity), who needs to be studied for that, and a first draft of how this phenomenon can be captured (candidate study designs and recording strategies). Step 2 (study implementation) is concerned with the translation into an executable protocol and covers governance decisions (ethics, consent, sensitivity, and access), selection of data streams and an analysis plan, acquisition setup, synchronization, quality control, and data management. Step 3 focuses on dissemination and reuse, including repository and licensing choices, documentation, and access procedures. Iterative refinements become necessary when constraints emerge: For example, pilot recordings may reveal that the chosen analysis requires higher image resolution than initially planned, ethics review boards may impose limitations on which data can be recorded, and dissemination goals may require additional metadata collection (prompting updates to the acquisition protocol). Thus, from each step, prior decisions may be revisited until all aspects are jointly consistent.
Three illustrative case studies
Throughout this article, we illustrate the complexities of multimodal data collection by drawing on examples from three distinct studies of our own: eliciting multimodal data in controlled laboratory settings (Case Study 1), constructing annotated corpora for large-scale analysis (Case Study 2), and gathering naturalistic data in field environments (Case Study 3). Although these examples do not exhaust the range of possible approaches, focusing on three case studies allows us to systematically discuss key decision points and their potential consequences in multimodal data collection. The example studies contain multimodal data collection for signed, gestural, vocal, and other bodily interaction and are intended as illustrations of the workflow rather than an exhaustive map of multimodal research. In the following, we provide brief descriptions of the case studies. Because of space constraints, readers are referred to the relevant literature for further details.
Case study 1
As an exemplification of eliciting multimodal data in the lab, we examine a gesture-elicitation task conducted to inventory the gestural repertoire of nonsigning speakers of German (Spruijt et al., 2025). This experiment is part of a larger project investigating the influence of the gestural repertoire on the acquisition of individual parameters of a sign—hand shape, location, and movement—by second-language learners of a language in a different modality than their native language. For the elicitation study, silent gestures were elicited from 18 native nonsigning speakers of German to investigate the systematicity of gestural representations for particular concepts. None of the participants had prior exposure to a signed language. Participants were presented with a concept in written form and were asked to express this concept gesturally and without pointing to objects in their surroundings.
Case study 2
To exemplify a construction of an annotated multimodal corpus for large-scale analysis, we refer to the German Sign Language (DGS) Corpus, which is a collection of signed conversations in DGS that have been enriched with additional information (e.g., annotations of manual and nonmanual features or OpenPose measurements) to facilitate linguistic analysis. It is part of the DGS Corpus project, a long-term initiative aimed at developing a data set suitable as a linguistic reference corpus, cultural heritage collection, and the basis for a corpus-based dictionary of DGS (Prillwitz et al., 2008). In its first data-collection campaign (2010–2012), the project collected 560 hr of discourse from 330 participants. A second campaign is currently ongoing (Konrad et al., 2024). Recordings cover participants from all regions of Germany, and all participants, moderators, and technicians present during recording sessions use DGS as their primary language of daily life. Participants were grouped into pairs and engaged in various conversational tasks, including discussions on assigned topics or historical events, free dialogue, and story retelling (Nishio et al., 2010). For further details on corpus curation and demographic composition, see Schulder et al. (2021) and Schulder and Hanke (2022). Fifty hours of the DGS were made available publicly as the Public DGS Corpus, which includes the research data set “MY DGS—Annotated,” which provides full sign annotations and translations in German and English (Hanke et al., 2020; Konrad et al., 2020).
Case study 3
Although lab studies have a variety of advantages over the control of the environment, not all studies are possible or ideal in the lab. As an example of gathering naturalistic data outside of the lab, we refer to the collection of multimodal data from nonhuman primates. Here, instead of focusing on a specific case study, we refer to a group of existing empirical studies on one of humans’ closest living relatives, the chimpanzee (Hobaiter et al., 2017; Mine et al., 2024; Wilke et al., 2017). All of these studies involve video observations of naturalistic interactions in wild settings (although semiwild and captive settings are also possible) and annotation of communicative signal production in different modalities.
Discussion of Case Studies
Across the three case studies, trade-offs between experimental control, ecological validity, cost, and measurement precision become apparent (see Table 1). Case Study 1, the laboratory-based gesture elicitation, affords strong control over stimuli and experimental conditions, enabling precise measurement of specific behaviors and causal inference. However, this comes at the expense of ecological validity, generalizability, and sample diversity. Case Study 2, the DGS Corpus of signed conversations, captures richer multimodal patterns and broader generalizability while requiring substantial resources for annotation and yielding largely descriptive and correlational findings. Case Study 3, the naturalistic observation of chimpanzee communication, ensures ecological validity and evolutionary relevance but provides limited control over the study conditions and reduced statistical power in the analysis of small and hard-to-standardize samples. As summarized in Table 1, these examples illustrate how decisions regarding study design, participant populations, and data-collection methods directly shape the interpretability, scope, and robustness of research outcomes. Building on these illustrative examples and the trade-offs they reveal, we now move from case-based insights to a structured decision framework that systematizes the key considerations for multimodal data collection.
Overview of Methodological Trade-Offs Across Three Case Studies
Note: The table summarizes the populations studied, behaviors targeted, data-collection methods, and the consequences of study-design choices for control, ecological validity, generalizability, and measurement precision. WEIRD = Western, educated, industrialized, rich, and democratic. (Henrich et al., 2010).
Step 1: Defining the Research Space
Step 1 is to define the broad conceptual building blocks of a specific study. Defining these is the first important decision that cascades into requirements for technical equipment, recording and analysis plans, and associated legal and ethical regulations. The key goal in this step is to make these choices explicit, that is, (a) the overall research question, the phenomenon that will be studied, its most relevant units of analysis and level of granularity (what); (b) the investigated population and the associated setting (who); and (c) early drafting of a study design and capture strategies in terms of technical and legal feasibility (how).
What: research questions and phenomena
The first step in any empirical study is to clarify the research question(s) and phenomena, what kind of data are required to address a specific research question (operationalization), and what type of conclusions one aims to draw from that data. The nature of the phenomenon under investigation plays a central role in these decisions.
A useful early distinction is between examining highly specific phenomena (e.g., individual signs or phonemes, isolated gestures, or specific facial expressions) with a relatively short time frame and studying broader, continuous interactions (e.g., dialogues, group interactions, or social behavior). Importantly, this distinction does not imply that a given data set is restricted to a single level of analysis: Multimodal data sets often support analyses at multiple temporal and structural scales even if a study foregrounds one primary level. Specific phenomena often require designs with a high level of experimental control in which specific conditions are elicited in a defined way and later compared (as in Case Study 1). Research on continuous interactions often cannot be as tightly controlled and is therefore typically implemented using observational designs, although strict experiments are also possible (e.g., a specific aspect of a social interaction is manipulated, and observed changes in behavior are compared). Observational designs are better suited for capturing behavior occurring in naturalistic settings but allow for drawing only looser, correlational conclusions (as in Case Study 2 from the DGS Corpus).
Decisions about dependent variables and their measurement not only reflect the research question and hypotheses but also shape the choice of the general setting and recording devices. Early decisions about data-collection protocols—such as whether to conduct focal sampling or ad libitum video recordings—can substantially influence the observed range of behaviors (i.e., the behavioral repertoire). Overall, decisions concerning the type, scope, granularity, and modality of data should align closely with the target phenomenon to support robust and meaningful conclusions. Note that many data sets from nonhuman primate research still show a preference for one modality over others rather than being truly multimodal (Liebal et al., 2022).
The choice of unit of analysis and granularity of measurement is a fundamental decision for subsequent analysis strategies: Short sequences (e.g., isolated gestures, signs, words, or phrases; e.g., in lexicon studies or imitation experiments) allow for fine-grained examination of specific motoric signals, including the kinematic (speed, amplitude, duration) and dynamic aspects (forces, acceleration, mass) and acoustic properties (temporal, spectral) of utterances or even individual segments. At the same time, such data can often be reaggregated to examine higher-level patterns (e.g., sequences or interactional structure), just as continuous interaction data can be segmented for fine-grained analysis. In continuous interaction data, such as conversations, interviews, or debates, the focus of analysis can be shifted to overarching aspects, such as pragmatics and discourse dynamics.
The phenomena captured in data collection, the research questions posed, and the levels of analysis are closely interrelated, and these relationships are reflected in the trade-offs summarized in Table 1. When the goal is to characterize brief, well-defined events and their timing (e.g., coupling of gesture and speech or discrete signal features and their interrelation), controlled elicitation designs, such as in Case Study 1, enable precise measurement and allow for causal inference. When the focus is on recurring multimodal themes that structure interactions, a corpus-based investigation (as illustrated in Case Study 2) allows for detecting broader generalizations about language patterns but at the cost of labor-intensive annotation and primarily descriptive or correlational inference. In naturalistic field studies, ecological validity is prioritized, and observations can focus on extended behavioral sequences and convergent multimodal patterns that form larger thematic units (Ladewig & Horst, 2024; Müller & Kappelhoff, 2018; Müller & Ladewig, 2013). In nonhuman primate research, signal analysis is likewise determined by the research questions, and there is increasing emphasis on systematic and transparent protocols for annotating and classifying signal types to increase measurement control as much as possible even when the context may be multifactorial (e.g., ChimpFACS, Vick et al., 2007; ChimpLASG, Zulberti et al., 2024).
Taken together, these aspects highlight the importance of aligning data granularity, population, and data-collection methods with research objectives: Fine-grained control of the experimental setting facilitates causal inference, whereas broader, naturalistic sampling enhances ecological validity and generalizability. Examples for related practical decision rules are the following: If the primary target of measurement is a brief, fine-grained event, then acquisition and annotation must prioritize temporal and spatial precision; if the target is extended interaction, then broader contextual capture and tolerance for greater variability is a priority.
Who: populations and research settings
Multimodal data can be collected from specific demographic or social groups (e.g., children, adults, elderly, bilingual speakers, individuals with communication disorders), language communities (e.g., users of a particular sign language), or even other species (e.g., nonhuman primates). The choice of population is tightly linked to the research question, that is, either a specific phenomenon is investigated in a particular population (e.g., signing in the Deaf community) or compared across different groups (e.g., language development from childhood to adulthood). In addition to selecting the population of interest, researchers need to define (a) inclusion/exclusion criteria and recruitment channels and (b) the research setting (e.g., lab vs. field, single site vs. multisite), (c) take into account constraints that are unique to the target population, and (d) anticipate respective consequences for governance and acquisition protocol. For example, in Case Study 1, selecting a task-compliant adult sample in a standardized lab-based setting is a straightforward choice to achieve tight experimental control. In Case Study 2 (DGS Corpus), community needs were addressed through multisite recordings, which also accounted for regional linguistic variation of DGS in Germany, and through the involvement of community-aligned staff to reduce barriers to participation. In Case Study 3, the target population (captive vs. wild) is inseparable from the research setting and has direct theoretical implications because communicative behaviors may be shaped by rearing history and environment settings.
A useful decision rule is to align the population and research setting closely with the targeted primary research question: If the hypothesis is about causal inference and requires timing control and stimulus manipulations, strict experimental control is possible only in lab-based environments, preferably at a single site (as in Case Study 1). If the goal is to capture typical language patterns and population variation (as in Case Study 2), multisite recordings in a standardized recording setting can be appropriate, and the focus is observational and seminaturalistic (i.e., less explicit stimulus control and no strict experimental manipulation). If the target is naturally unfolding behavior, a field observation setting is often the best choice (Case Study 3).
Studies that target specific populations, such as infants, children, the elderly, clinical populations, or rare/marginalized communities, need to take into account from early on their respective requirements and constraints. Important aspects to consider are, for example, tolerance for specific measurements and experimental procedures, participant burden, expected fatigue or attrition, increased privacy, reidentification risks, and potential recruitment bottlenecks (e.g., reliance on clinics, schools, community organizations). These constraints should be treated as limiting factors for certain decisions (for more details, see Step 2: Study Implementation), and the goal is often to determine a “minimum viable” set of data streams and procedures that operates within regulatory and practical limitations. For a compact overview of typical constraints in special populations and recommended workflow adaptations, see Table S1 in the Supplemental Material.
How: study design and capture strategy
Once what and who are specified, the next goal is to draft a feasible study design and data-collection strategy that can capture the target multimodal behavior with sufficient precision in the appropriate setting (see Step 2: Study Implementation).
Conceptually, data-collection methods in language research range from elicited (high experimenter control) to naturalistic (minimal intervention, focus on observation) and have respective trade-offs between control, ecological validity, logistical effort, and data quality (see Table 1).
In designs with elicited data, participants are prompted to produce specific responses (e.g., in a stationary or mobile laboratory). This approach is particularly useful for studying specific linguistic features or behaviors under controlled conditions. Both Case Studies 1 and 2 employed such an approach, although there are obvious differences: Whereas in Case Study 1, participants were recorded in a single university laboratory, the DGS Corpus (Case Study 2) was collected at various sites throughout Germany to avoid participants adjusting their dialectal language use to that of the recording region; the data-collection team would travel instead and bring along the necessary equipment.
In designs with naturalistic data, participants are observed in their natural environments without the intervention of experimenters, enabling, as in studies from Case Study 3, the investigation of multimodal signal use across a variety of contexts and therefore providing a comprehensive understanding of the communication system under investigation. This is particularly relevant in the field of nonhuman communication because methods for the inference of signals’ meanings and functions can be derived only by contextual information given by the naturalistic settings of interactions.
In practice, elicited and naturalistic data collection can be implemented in different research settings, for example, in laboratory studies, remote/online studies, and field studies, each with its characteristic constraints. Laboratory studies provide high control through standardized stimuli and close supervision, enabling reliable measurement and strong causal inference. This method requires participants to be physically present at a specific location, which can involve complex scheduling—especially when multiple participants are needed—and often results in smaller sample sizes, as exemplified by Case Study 1 (Offrede et al., 2021).
Large-scale multimodal data collection can also occur in lab settings with less control over participant behavior, as in DGS Corpus exemplified by Case Study 2. Here, seminaturalistic interactions are recorded, capturing multiple modalities—including manual and nonmanual signals—across larger and more diverse participant groups than in Case Study 1. This approach balances ecological realism with logistical feasibility, enabling systematic annotation and quantitative analysis of multimodal communication.
Web-based data collection, by contrast, offers broader accessibility and faster turnaround, typically completing within hours or days. Yet it poses challenges for quality control because data are gathered through participants’ personal devices under all kinds of circumstances and activities. Prespecified quality-control steps and extensive capture of metadata are helpful to mitigate these challenges (for more details, see Protocol and Acquisition Setup).
Field or on-site data collection prioritizes ecological validity by recording participants in their natural environments. In naturalistic settings, data collection typically involves recording participants in real-world interactions, such as everyday conversations, interviews, or group discussions, under minimal researcher intervention. Ethnographic methods, such as participant observation, can also be employed, in which researchers spend extended periods immersed in the group or community of interest. Environmental factors, such as noise, visually occluded environment, and multiple individuals that overlap their vocalization, speech, gestures, and signs in time and space, reduce experimental control of what was produced from which individual. On the other hand, they may allow authentic observation of communicative behavior and access to large participant groups. Naturalistic multimodal data collection, as in Case Study 3, captures gestures, facial expressions, and vocalizations in real-world interactions, illustrating the trade-off between ecological realism and experimental control.
In sum, laboratory, seminaturalistic lab/studio, and field approaches each provide unique opportunities for multimodal data collection, requiring researchers to balance control, ecological validity, and practical constraints. Across all approaches, governance decisions (e.g., informed consent, privacy, implementation of the right to withdraw, community expectations, and other ethical issues) constrain what can be collected and shared.
Step 2: Study Implementation
Once the conceptual design has been specified, the next challenge is to implement it in a way that is technically robust, ethically sound, and feasible. The goal of this step is to come up with an actionable plan that contains all implementation details with respect to technical considerations, analyses, and dissemination of results and all associated ethical and legal requirements. Many decisions need to be made during this step that influence each other and sometimes also refine or modify the initial definitions of the research question and population to be studied.
Governance
Early consideration of ethical and regulatory requirements is essential because approvals can take substantial time and directly shape what is permissible during implementation. Systematic governance ensures that research is conducted responsibly and that data management is professionally handled from the outset. Five interrelated decision points are central to this process: data sensitivity and sharing goals, planning informed-consent procedures, anonymization and access policies, storage and backup strategies, and repository selection. Trade-offs between objectives—such as maximizing openness versus protecting participant privacy—are often unavoidable. However, careful planning can optimize data availability, facilitate reproducibility, and preserve ethical standards.
A key principle is that governance decisions depend on the target population and the associated research setting: For example, when working with clinical populations or very small communities, both the likelihood of reidentification and the potential harm from unintended disclosure are increased, typically resulting in the need for stricter access control and constraints for data sharing. Likewise, children, older adults, and many clinical participants often face higher participant burdens from complex or lengthy procedures and may require additional procedures (e.g., legal-guardian consent, additional regulatory requirements from clinics) that can further restrict which data can be collected and shared. In practice, these constraints can become a primary driving factor in study implementation. These population-dependent constraints and practical mitigation strategies are summarized in Table S1 in the Supplemental Material, which links common issues (e.g., participant burden, reidentification risk) to concrete adjustments and their consequences.
Data sensitivity and sharing goals
A key initial decision is how much of the data set can or should be shared publicly. This depends on factors such as whether audio-video recordings contain identifying information, who funded the project, and whether participants expect or require anonymity. Some researchers choose to avoid video recordings for privacy reasons. Case Study 1 illustrates a scenario in which only short clips were shared openly and the complete data set remained with restricted access. In contrast, large public corpora, such as the DGS Corpus, emphasize comprehensive data sharing: Participants gave informed consent for open use of their videos in linguistic research (Schulder et al., 2021; Schulder & Hanke, 2022). Even then, demographic metadata (region, age) were coarsened before publication to balance researcher needs with privacy considerations, and all dialogues mentioning personal identifiable information of participants or third parties were anonymized (Bleicken et al., 2016). The fully curated set, accompanied by annotations, was released under a specific license, and more sensitive or raw material is restricted to trusted researchers under additional agreements.
Specific considerations with respect to identifiable audio-video recordings arise when these are linked to clinical diagnoses or vulnerable populations. When collecting multimodal data from special or vulnerable populations, such as children, Deaf individuals, or even students training to become sign-language interpreters, ethical and practical considerations are closely tied to the research setting. Vulnerable populations often belong to small, tightly knit communities where many members know one another. In Deaf communities, both nationally and internationally, sharing video recordings across contexts—for example, with bilingual signers in another country—may inadvertently reveal personal information, such as an individual’s residence or community ties. Likewise, students training as sign-language interpreters may be hesitant to have early production errors exposed, which could affect their professional development. Researchers must therefore carefully balance the benefits of sharing multimodal data sets with the need for strict de-identification protocols and robust consent procedures, ensuring that participants and their guardians fully understand the implications of data sharing.
Plan informed consent
Multimodal data sets are often inherently identifiable (face, voice, body shape, movement style, interactional content). For this reason, informed consent needs to cover sufficient detail: (a) what exactly is recorded, (b) how it will be processed (including attempted anonymization procedures), (c) where the data will be stored, (d) data-access policies (who and for what purpose), and (e) regulations for withdrawing consent and removal or anonymization of data. All these aspects are essential for ethics approval and constrain the concrete study implementation and dissemination of collected data. Clear communication about these issues is vital (Meyer, 2018), especially when multiple institutions and/or countries are involved. Ensure that procedures satisfy institutional review boards and legal frameworks (e.g., General Data Protection Regulation, Health Insurance Portability and Accountability Act) and discipline-specific established guidelines. Often, layered consent forms are required when participants are allowed to freely choose which aspects of data sharing they approve (internal use only, controlled access for qualified public release of selected derived data), as is detailed information on the respective consequences (e.g., revoking access for data under controlled access is possible but for openly published data often not possible, including potential reidentification risks for certain types of derived data).
Anonymization, access policies, storage strategies, and repository selection
Closely related methodological decisions concern the extent of anonymization and the openness of data access, which must balance privacy protection with research transparency. Multimodal data frequently consist of audiovisual recordings of identifiable individuals. When participants belong to vulnerable populations, such data are particularly difficult to store in open repositories.
In sign-language corpora or laboratory experiments involving elicited gestures, complete anonymization (e.g., facial blurring or voice distortion) may remove information that is crucial for analysis. An important consideration, therefore, is whether partial anonymization is feasible without compromising the study’s objectives or whether access-restricted data sharing is the more appropriate strategy. One compromise is to publish derived features alongside anonymized video data. For example, this may include time-aligned motion-capture data of body joints or facial-expression contours over blurred bodies and faces (see e.g., the “masked-piper” approach; Owoyele et al., 2022). Note, however, that even extensive kinematic information may still carry reidentification risks (Battisti et al., 2024). In less controlled language-production settings, participants may also mention private information, such as personal identifiable information about themselves or others, that may require anonymization in both primary recordings and secondary materials, such as transcripts, annotations, or motion-capture data (Isard, 2020). Different types of settings and tasks may be more or less likely to necessitate such steps (e.g., personal narratives and retelling picture stories).
When full anonymization cannot be achieved, controlled-access repositories can help protect participant identities while still enabling research collaboration, review, and reuse of multimodal data. For particularly vulnerable populations, such as children or patients, openly shared data are often limited to carefully screened derivatives (e.g., transcriptions, categorical facial-expression time courses, or kinematic vectors), and raw audiovisual recordings remain stored at the institution where the data were collected.
Gregori et al. (2023) reported a growing number of tools that automatically enable partial identity masking in video and audio while extracting nonidentifiable information for multimodal analysis (see e.g., Khasbage et al., 2023; Rachow et al., 2023). Large-scale data sets may require automated or semiautomated approaches for face blurring, body masking (Owoyele et al., 2022), or voice distortion. Nevertheless, reidentification risks persist as analytical methods evolve (Gadotti et al., 2024), making periodic reassessment of anonymization strategies essential.
In nonhuman primate fieldwork, anonymizing incidental appearances of humans is relatively straightforward. In contrast, field studies centered on human participants typically require substantially greater effort to obtain informed consent in oral or written form, depending on literacy, and ensure adequate anonymization, which is often impractical in fully open settings. Clinical data linked to personal health information pose additional challenges and must be handled on a case-by-case basis (Kaur & Cheah, 2024). In such cases, fine-grained, layered access strategies offer a viable solution (Marcotte et al., 2023).
Data streams (what to record)
In multimodal data collection, each data stream (i.e., each measurement/acquisition modality) constitutes a continuous flow of data generated through a distinct acquisition method. In the language sciences, video and audio streams are foundational and may be supplemented by specialized measures, such as eye tracking, motion capture, physiological monitoring, or neurophysiological recordings (e.g., EEG or MEG). Depending on the phenomenon, primary streams may also include screen recordings, digital-pen trajectories, interaction logs, or button presses. Because the choice of data streams determines which phenomena can be observed, deciding which streams to record is a core implementation step.
For example, studying the relationship between co-speech gesture and speech may require dedicated high-quality audio capture in addition to high-quality video, rather than relying on the camera’s integrated audio signal, or additional specialized motion-capture equipment. Although additional data streams can substantially enrich a data set, each stream also increases acquisition complexity, synchronization effort (see Gregori et al., 2023), and the risk of technical failure. Researchers should therefore include only those streams whose unique and essential contributions to the research question are clearly justified and nonredundant. Furthermore, it should be kept in mind that more data streams are always associated with increased risk for unintentional reidentifying an individual and should thus be kept to a minimum. When working with children, clinical populations, and small communities, tolerability and burden need to be considered carefully. For example, setting up multiple additional sensors (e.g., physiology, EEG, eye tracking) can substantially increase time requirements; a lab setting with a large array of cameras and microphones may intimidate children and patients.
The inclusion of additional data streams is justified when they contribute complementary information that cannot be derived from other streams and is essential for addressing the research objective (Lahat et al., 2015). For some research questions, multiple streams are required because they are codependent, and this codependency is itself central to the research objective. For instance, events in one stream may serve as analytical anchors for another, such as fixation-related EEG effects defined on the basis of eye-tracking data (Brilmayer et al., 2025; Dimigen & Ehinger, 2021). Likelihood of equipment failure and associated cost of additional recording sessions may also factor into decisions regarding equipment that provides redundant or overlapping data. In other words, if an additional stream does not contribute unique information that is necessary for the research objective, then it should not be recorded; if it is indispensable, then researchers must plan for the added synchronization, governance, and quality-control demands.
Analysis plan
At least a coarse analysis plan should be drafted before finalizing the acquisition plan because the chosen analyses determine which features must be measurable at which temporal and spatial resolution and thus, the required annotation-effort, synchronization-precision, and recording specifications. An exhaustive overview of available up-to-date analysis tools is beyond the scope of this article (but see Gregori et al., 2023); here, we briefly mention some core choices and their consequences.
Annotation ecosystems
Most multimodal projects rely on manual coding of at least some behavioral streams (e.g., conversational structure, gesture categories, sign boundaries, turn transitions). This typically requires a dedicated annotation environment to link multiple audio/video tracks with structured annotations exportable for further analyses. Tool choice should be guided by the complexity of the annotation scheme (simple event coding vs. multitier linguistic annotation) and existing community standards in the target scientific field. Using the most widely adopted tool in a research community is often the best choice (e.g., ELAN-style tier-based workflows in sign language and gesture research; event-logging tools, such as Behavioral Observation Research Interactive Software (BORIS) in behavioral/animal research). In cases in which established annotation schemes or conventions exist, reusing these is often preferable to developing idiosyncratic schemes from scratch, unless the research question clearly demands otherwise.
Semiautomatic pipelines extract features from video using computer-vision methods (e.g., OpenPose; Cao et al., 2021). Using such pipelines requires substantial efforts for quality control and postprocessing (e.g., identity tracking, handling missing frames, smoothing, and targeted manual correction; Schulder & Hanke, 2020; Trettenbrein & Zaccarella, 2021) to be a useful complement of manual annotation. The decision to use these tools directly influences other steps, such as choice of data streams and acquisition setup: For example, specific camera views (e.g., multiview to minimize occlusion), higher frame rates, and optimized shutter speeds to reduce motion blur and lighting and background to enhance contrast (see Protocol and Acquisition Setup) are needed. If pose estimation outputs will be disseminated, governance decisions must still address privacy concerns because pose data may also contain unique and identifiable features depending on the type of analysis. It should also be considered that computer-vision tools are often trained with adult input stimuli and may be less accurate for children or infants.
Multimodal integration
Multimodal analyses require explicit decisions regarding the integration of different data streams, for example, whether modalities are analyzed separately and linked post hoc via time-aligned events or fused into joint representations (Lahat et al., 2015). Further choices that govern the alignment of data streams are the unit of analysis (discrete events vs. continuous trajectories), the harmonization of different sampling rates, and integration of spatial-coordinate systems (e.g., mapping of gaze position from a mobile eye tracker and world video). In turn, these aspects directly influence the required precision of synchronization, tolerance for drifts, and acceptability of missing data. For example, reconstruction of precise three-dimensional pose and the analysis of movement kinematics require calibrated cameras and frame-accurate synchronization (e.g., Hanke et al., 2010). Other analyses of dyadic interaction (e.g., coordinated dynamics of facial expression, gesture, and speech) may tolerate small timing deviations as long as event order and interval boundaries remain stable.
Protocol and acquisition setup
In this section, we outline key practical considerations for turning the design into a concrete recording protocol. Choices should be made with the analysis plan and governance constraints in mind. We focus on recording setups, including video-equipment selection, camera configurations, background setups including lighting conditions, and microphone placement. Drawing on the case studies introduced earlier—most notably, the DGS Corpus (Case Study 2) and gesture experimental study (Case Study 1)—we illustrate how acquisition choices directly affect data quality, annotation, feasibility of automated analysis pipelines, and the suitability of recordings for reuse in diverse research contexts. Throughout, we highlight trade-offs and provide concrete recommendations aligned with common research objectives.
Video-recording requirements
Selecting appropriate recording equipment is crucial for ensuring data quality and long-term usability. As a general baseline, high-definition (HD) cameras with a minimum resolution of 1,920 × 1,080 pixels are recommended. However, higher resolutions, such as 4K or even 6K, can be advantageous when data sets are intended for long-term reuse or analyses that require fine-grained visual detail, such as facial expressions or subtle articulatory movements. This motivated the DGS Corpus to upgrade from HD to 6K video in its second data-collection campaign.
If the research focuses on precise motion analysis—particularly rapid hand movements, such as those observed in signing or co-speech gestures—researchers should prioritize fast shutter speeds (at the very least, 1/100 of a second) to minimize motion blur. Depending on the camera type, this can be set either explicitly or by choosing a sufficiently high frame rate (here: 100 frames per second or more). Standard indoor recordings at 50 frames per second are typically insufficient for capturing such movements reliably. High frame rates, however, require adequate and evenly distributed lighting; insufficient illumination will otherwise reintroduce blur or noise. Image blur substantially impedes both human annotation and automated computer-vision approaches, making lighting conditions a critical part of the acquisition protocol (Crasborn & Morgan, 2022).
Camera configurations
Camera placement and the number of camera views should be determined by the interactional structure of the study and the intended analyses. In studies involving dyadic interaction, a common and effective setup consists of at least three cameras: one frontal view for each participant and one additional camera capturing both interactants simultaneously (see Fig. 3). This configuration supports both individual-level analyses and interactional perspectives. For experimental noninteractive studies (e.g., Case Study 1), a second camera at an angle may also be helpful to capture partially occluded expressions or a distinct angle to access the precise form.

Technical setup of a DGS Corpus recording session (screenshot from dgskorpus_koe_01_free conversation; source: https://doi.org/m59x). The informants are seated facing each other at a distance of approximately 3 m. The setup includes five cameras: Two high-definition (HD) cameras provide frontal views of the informants, two HD cameras (not visible in the picture) capture the signing from above for reconstruction of depth information, and one HD camera records the overall scene, including the moderator (seated between the informants, in the chair that is empty in the picture). Elicitation materials are presented on screens placed low between the signers to avoid obstructing their views. The studio arrangement ensures separate filming of each informant while also offering a comprehensive view of interactions, assisting with subsequent transcription and analysis (Hanke et al., 2010).
In sign-language recordings, occlusion and the lack of depth information in standard two-dimensional video present major challenges. If the study involves detailed manual articulation, spatial relations, or translation and gloss annotation, researchers should consider adding a top-down (bird’s-eye) camera view, as implemented in the DGS Corpus (Hanke et al., 2010). This perspective can substantially improve the visibility of hand configurations and spatial relations.
If the research goal includes computational analyses, such as full three-dimensional reconstruction or fine-grained kinematic modeling, additional camera angles—such as 45° side views—are often necessary. In the DGS Corpus, each recording is supplemented with pose estimation data generated using OpenPose (Cao et al., 2021) and extensively postprocessed to correct common errors (Schulder & Hanke, 2020). These corrections address issues such as false detection of nonexistent persons, erroneous splitting of a single individual into multiple entities, misattribution of limbs between participants, and missing frames. The resulting refined pose data enable detailed investigations of multimodal kinematic phenomena, including head nods (Bauer et al., 2024), and stimulus control when videos are reused as experimental materials (Trettenbrein & Zaccarella, 2021).
Background setup
Background selection is another critical design decision. Uniform backgrounds that provide strong contrast with participants’ clothing are generally preferable. Although blue and green screens are commonly used, any evenly colored background that maximizes contrast can be effective. In practice, participants wearing unicolored clothing—preferably colors that contrast with their skin color—facilitate both visual inspection and automated processing (see Fig. 3).
If researchers anticipate using chroma key techniques or automated segmentation, blue screens often offer a good compromise: They function as a neutral background in unaltered video while also supporting background replacement. Green screens are widely used for chroma keying but may be perceived as visually uncomfortable by some viewers. Importantly, highly artificial backgrounds or overly simplified recording environments can negatively affect participant comfort and the naturalness of social interaction. Researchers should therefore weigh technical advantages against potential impacts on ecological validity, particularly in interaction-focused studies.
Microphone selection
Audio-recording quality is equally important in multimodal acquisition (Gregori et al., 2023). High-fidelity microphones ensure that speech and other vocalizations are captured accurately and are suitable for both manual analysis and automated processing. Microphone placement should maximize proximity to the respective speaker while minimizing visual interference and cross talk, especially in interactive settings with multiple participants.
If the study requires automatic speech recognition or speaker diarization, individual microphones for each participant are strongly recommended. Researchers should also attend carefully to ambient noise, room reverberation, and other sources of acoustic interference because these factors can substantially degrade audio quality and downstream analyses.
Remote data collection
Remote data collection (e.g., browser-based online studies) has specific challenges because it introduces variability in device characteristics. Therefore, it is particularly important to record detailed metadata (e.g., operating system, browser version, audio and video modes) and a brief assessment of context conditions (e.g., room noise, lighting). Based on requirements of the analysis plan, define minimum requirements and verify these via quality-control procedures (e.g., calibration routines, test recordings, and attention or catch trials). Moderate control can still be achieved in online perception experiments, whereas language-production tasks tend to suffer more from device variability and participant positioning.
Protocol adaptations for specific populations
Protocols must be adapted to accommodate specific capabilities and constraints of the target population. For instance, sensors typically designed for use with healthy adults may require different calibration procedures, hardware modifications, or adaptation protocols when used with children or clinical populations. In addition, individual differences in motor and cognitive development over time can affect the reliability of certain data readings. This is particularly important when setups integrate multiple data channels, requiring careful control to ensure data accuracy and consistency.
Sensor placement (e.g., for motion capture) and camera framing should be adjusted to account for frequent postural shifts or smaller body sizes. Stimuli and instructions are critical and must also be tailored to the developmental stage of participants to elicit behaviors across multiple modalities (e.g., speech and gesture). Shorter tasks, game-like interactions (e.g., Ćwiek & Fuchs, 2025), and visually engaging stimuli can improve participant engagement, particularly in multisensor setups or when motion-tracking equipment is used. In addition, real-time monitoring of children’s fatigue levels is crucial because tired or distracted participants may produce incomplete or lower-quality multimodal data.
Individuals with motor or cognitive impairments may face difficulties following specific instructions required for certain recording setups. To minimize participant burden, when possible, researchers should adapt the experimental setup, including, for example, task instructions and sensor configurations, to the needs of the target population (e.g., using remote eye-tracking instead of head-mounted devices) and consider collecting fewer but higher-quality data streams. Reducing complexity in the experimental setup can improve data reliability while ensuring a more inclusive research approach.
In specific cases, for example, when working with marginalized or mobile language communities, such as Deaf Ukrainians who have migrated to Germany, data collection should be conducted by researchers with established relationships in the community. Care must be taken not to overburden participants with repeated data-collection requests. At the same time, studies should aim to contribute meaningfully to the community, for example, by supporting integration initiatives or other activities that directly benefit participants. The key takeaway is that careful attention to population characteristics, research setting, and community engagement is essential not only for ethical compliance but also for ensuring the quality, relevance, and sustainability of multimodal-language data.
Synchronization
Synchronization is essential in multimodal data collection, particularly when integrating heterogeneous streams such as audio, video, and sensor data, an issue that was partially addressed in the literature by Gregori et al. (2023). When possible, time-code coordination should be used during recording to enable automatic alignment of different video and audio streams (i.e., synchronizing the internal timer generators of all recording devices and using data-recording formats with integrated time stamps; e.g., Brilmayer et al., 2025). If not feasible, external synchronization tools should be employed. A common approach involves a clapperboard or an LED flash at the session’s start, ensuring a clear visual and auditory marker on all recording devices. Some equipment requires synchronization via external trigger signals or remote protocols, which either ensure synchronized timing or embed signals for post hoc alignment. For nonvisual/auditory devices (e.g., physiological sensors, motion trackers, thermal cameras), this is often the only option. Device documentation should be consulted to assess synchronization precision (e.g., millisecond vs. 100 ms range); synchronization errors should be significantly smaller than the expected experimental effects (e.g., ≈200–400 ms).
An example of a controlled multimodal-data-collection setup is the gesture-elicitation study (Case Study 1). Participants sat against a solid blue background to enhance contrast, and consistent lighting was provided by four studio lights. A laptop monitor was placed at a 45° angle on the dominant side of the participant, out of reach to avoid interference with natural gestures. Each trial began with a fixation cross and an auditory cue, followed by a written stimulus displayed for 4 s. Participants gestured during this period, recorded simultaneously from two camera angles: one front-facing and one at a 45° angle from the nondominant side. To synchronize the data streams—including the two video recordings and the randomized stimulus presentation—button presses were logged in each modality. PsychoPy, the experiment software, automatically recorded these events. In addition, cameras captured the laptop keyboard, enabling later alignment in ELAN (Wittenburg et al., 2006).
Data management
In this section, we outline practical recommendations for managing multimodal data. We first address how to choose storage formats and document metadata to ensure data quality and reproducibility. We then discuss strategies for data safety, backup, and long-term archiving, including concrete guidance for both lab-based and field-based studies.
Data management is a core requirement of multimodal data collection because video, audio, sensor streams, and metadata vary widely in format, size, and temporal resolution. If researchers want their data to remain usable, reproducible, and shareable, data-handling strategies must be planned from the outset rather than addressed after data collection.
Data storage
Given the large file sizes associated with multimodal data collection, efficient data-storage and data-management strategies are essential. If the goal is to preserve maximal data quality, original recordings should be stored using lossless or minimally compressed formats (e.g., FFV1 for video; FLAC or WAV for audio). If the goal is long-term storage, accessibility, or data sharing, researchers should additionally create derivatives in widely supported formats, such as MP4 using H.264 or H.265 compression.
Large data sets should be stored in secure repositories with proper metadata documentation to facilitate reproducibility. Technical metadata should include detailed information about the recording devices used, specifying exact models, firmware/software versions, and hardware configurations (e.g., camera-sensor specifications, frame rates, shutter speeds, microphone sensitivity, sampling rates). In addition, precise settings for each device, such as resolution (e.g., HD, 4K, or 6K), compression format (e.g., lossless FFV1 or H.264/H.265 compression), audio-encoding standards, and synchronization methods (e.g., embedded time stamps, external trigger signals), should be documented. Information about the physical environment, such as lighting setup, background color and material (particularly for chroma keying), camera and microphone placement, and distances and angles relative to participants, should also be included. Furthermore, metadata should clearly describe data-management procedures, including file-naming conventions, storage formats, version histories, and postprocessing steps.
Data safety, backup, and long-term archiving
Given the typically large file sizes of multimodal data (e.g., up to 4K video recordings, sensor data with high sampling rates), reliable storage solutions and backup protocols are essential. Researchers should always make sure to take care of redundancy (at least two physically separate storage locations), keeping logs of structured metadata (e.g., sensor calibration, environment conditions, participant demographics, and other recording-session details). For example, in a lab-based study as the Case Study 1, each camera feed was backed up daily to external hard drives, and session logs were carefully time stamped to facilitate subsequent annotation in software tools, such as ELAN. For fieldwork, data redundancy and safety is at the same time more difficult to achieve and more critical: In the lab, data backup and redundant storage can often be achieved using automated scripts and simple procedures; more planning is involved in fieldwork (e.g., changing, backup of SD cards). Equipment is typically safer in the lab and protected against environmental influences and data loss than “in the wild.”
If long-term reuse or data sharing is intended, adopting the FAIR (making data findable, accessible, interoperable, and reusable; Schulder & Hanke, 2022; Wilkinson et al., 2016) principles provides a useful framework. These data principles promote robust documentation and standardize and open metadata. Creating a formal data-management plan can further support consistent practices for backup, versioning, and archiving (Michener, 2015).
Recent advances enable the collection of high-quality multimodal data across video-, audio-, and sensor-based streams (e.g., Gregori et al., 2023), although effective implementation still depends on careful choices regarding equipment, synchronization, and data management (Offrede et al., 2021).
In sum, Step 2 translates the conceptual design of a study into an executable protocol. This requires aligning governance constraints, technical acquisition setup, exact experimental protocol, choice of multimodal data streams to record and their synchronization, a solid concept for planned analysis, quality control, and the data-management strategy. These aspects need to be jointly satisfied in a technically feasible and ethically acceptable way and often cause further adjustments of the study goals and the research question (Step 1).
Step 3: Dissemination
Once the study goals and the implementation are sufficiently specified, data collection can begin; however, dissemination planning is equally important to satisfy open-science principles. Step 3 therefore addresses documentation, selection of those aspects of the study that can and should be shared for reuse, and appropriate licensing. These considerations should be made from the outset (i.e., before data collection has started) because they also affect governance and implementation decisions (Step 2).
Repository choice for publication and sharing
For vulnerable participant groups and data with only partial anonymization, only secure institutional hosting and storage of the data may be feasible. Such repositories may offer only limited data- and access-managing features that differ substantially across institutions. On the other hand, numerous publicly accessible repositories are available that offer flexible features, such as layered data-access restrictions, custom licenses, and advanced data management and version history (for available repositories and their features, see re3data.org, https://www.re3data.org/; Pampel et al., 2013). Usage licensing and repository choice is highly dependent on anonymization constraints and needs for access restrictions and ultimately needs to be in accordance with legal regulations and the explicit consent obtained from the participants. Thus, the ultimate check before publishing a multimodal data set always needs to be whether the storage and sharing strategies and policies align with the written informed consent. Whenever possible, persistent identifiers (e.g., DOIs) should be attached to archived data sets to allow for citation and tracking. Tools and web services, such as OSF, Zenodo, or other specialized repositories, often support versioning and fine-grained access controls (see e.g., Soderberg, 2018). It is also advisable to clarify the procedure for “take-down requests,” that is, how participants or their guardians may request partial or complete removal of their data at a later point.
Licensing for dissemination
Selecting an appropriate license is distinct from defining an access-control policy but an important step in data sharing (Ball, 2014). Access control defines who can initially obtain materials, and a license defines the conditions under which the material can be used (attribution, redistribution, derivative works, commercial use). Permissive licenses (e.g., Creative Commons [CC], https://creativecommons.org/chooser/) are most suitable for fully anonymous or anonymizable data and materials (metadata, annotations, derived data, analysis scripts). In data sets from academic research, CC BY provides a clear attribution requirement (e.g., a citation or acknowledgment in academic publications), and CC0 prioritizes facilitation of broad reuse without any restriction. Other CC variants (e.g., noncommercial [NC], no derivatives [ND], share alike [SA]) may be attractive in principle but can hinder scientific reuse in practice: For example, NC may be problematic in contexts with industry cofunding, ND can block many legitimate forms of reuse and transformation, and SA can complicate aggregation across sources. Because CC licenses are typically nonrevocable, this choice is not recommended for identifiable data, data with a potential risk of (re)identification, or otherwise sensitive materials. In these cases, either controlled access with customized, research-only data-usage agreements are needed to restrict redistribution and usage (e.g., as in the Public DGS Corpus). Depending on the informed consent in the respective legal context and studied population, such data-usage agreements must be designed in a way that allows for revoking access and usage rights of specific materials. Software and analytical pipelines should use established open-source licenses (MIT, Apache 2.0, GPL) rather than CC licenses. Stimuli and third-party materials often require separate handling because of copyright concerns. Practical resources for selecting licenses include the CC license chooser and institutional research-data-management guidance.
Summary
In this article, we addressed multimodal data collection as a methodological challenge that cuts across subfields of the language and communication sciences. Building on a clear distinction between multimodal language and communication (i.e., the inherently multimodal organization of linguistic systems and communicative behavior) and multimodal data collection as a research practice (i.e., integrating various data streams), we have identified key methodological, ethical, and practical considerations that arise when studying multimodal phenomena. Using three contrasting case studies on (a) controlled gesture elicitation (Case Study 1), (b) large-scale sign-language-corpus construction (Case Study 2), and (3) field-based primate research (Case Study 3), we demonstrated how these considerations manifest across laboratory, corpus-based, and naturalistic research settings.
A central insight emerging from these case studies is that multimodal data collection necessarily involves trade-offs among experimental control, ecological validity, and scalability. Highly controlled laboratory designs allow researchers to target specific phenomena with precision but may limit generalizability to everyday interaction. Corpus-based approaches offer breadth, demographic coverage, and ecological richness but require substantial investments in synchronization, metadata standards, and annotation infrastructure. Field-based studies maximize ecological validity yet pose unique logistical and technical challenges that constrain control and standardization. Rather than treating these trade-offs as methodological weaknesses, we argue that they should be understood as predictable consequences of studying inherently multimodal behavior.
The primary contribution of this article is a flexible, decision-oriented workflow for multimodal data collection that makes these trade-offs explicit and actionable. Organized around three steps—defining the research space, implementing the study, and enabling dissemination and reuse—the framework supports transparent alignment between research questions, data streams, ethical governance, and data-management strategies. By foregrounding decision points (e.g., level of granularity, synchronization requirements, robustness of acquisition, documentation standards), the workflow helps researchers anticipate constraints and select methods that are fit for purpose rather than nominally optimal.
This framework is intended to improve methodological rigor, reproducibility, and cumulative knowledge building. Multimodal data sets are complex, expensive to collect, and difficult to replicate. Making methodological decisions explicit—and documenting them systematically—enhances interpretability, facilitates reuse, and supports meaningful comparison across studies and laboratories. Importantly, the framework is designed to be adaptable: It does not prescribe a single “best” method but provides structured guidance that can be tailored to different populations, settings, and resource constraints.
More broadly, this framework highlights why multimodal data collection is essential for psychological and language sciences. If cognitive and communicative processes unfold across multiple physical channels, then methods that reduce these processes to a single modality risk omitting theoretically relevant structure. Multimodal data collection is therefore not simply a technical extension of existing practices but a methodological necessity for studying complex multimodal behavior in ecologically valid and theoretically informed ways.
In conclusion, advancing multimodal research requires not only better tools but also clearer methodological frameworks that connect theory, data, and practice. By articulating key decision points and offering a reusable workflow, we aim to support more transparent, ethical, and robust multimodal research practices.
Supplemental Material
sj-docx-1-amp-10.1177_25152459261442338 – Supplemental material for Data Collection in Multimodal Language and Communication Research: A Flexible Decision Framework
Supplemental material, sj-docx-1-amp-10.1177_25152459261442338 for Data Collection in Multimodal Language and Communication Research: A Flexible Decision Framework by Anastasia Bauer, Patrick C. Trettenbrein, Federica Amici, Aleksandra Ćwiek, Lisa-Marie Krause, Anna Kuder, Silva Ladewig, Marc Schulder, Petra B. Schumacher, Door Spruijt, Chiara Zulberti, Susanne Fuchs and Martin Schulte-Rüther in Advances in Methods and Practices in Psychological Science
Footnotes
Transparency
Action Editor: Rogier Kievit
Editor: David A. Sbarra
Authors Contributions
A. Bauer and P. C. Trettenbrein contributed equally to this work and should be considered joint first authors. All authors are listed in alphabetical order, except for the two first authors and two senior authors.
ORCID iDs
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
