Abstract
New federal research policies support data-driven science by requiring data management planning and data preservation for open access to scientific data and research products. This article shows a “platform effect” in how scientists plan for data preservation. By examining 976 data management plans for projects funded between 2011 and 2021 by the United States’ National Science Foundation (NSF) in five research areas, we document a shift in planning for data preservation on corporate-owned platforms away from open-source solutions. These findings raise questions about the potential impact of cloud storage services and corporate platforms on scientific knowledge production. We argue that the dynamics of platform capitalism and open access mandates create new challenges for scientists, policymakers, and proponents of open science because of how value and control are distributed within digital preservation infrastructures. As research data increasingly becomes managed and made accessible by platforms and cloud storage services, scientific knowledge risks becoming enclosed, with implications for the reconfiguration of knowledge infrastructure, control over long-term preservation, open access to data, and the circulation of scientific research.
Keywords
Introduction
Modern scientific knowledge production is dependent on access to data. Such access relies on the digital preservation and circulation of data, research products, and peer-reviewed science publications. Platforms are used in nearly every area of society to coordinate interactions between users, institutions, and third parties to manage and integrate data resources (Srnicek, 2016; Van Dijck et al., 2018). Scientists and research data repositories increasingly depend on platforms and cloud architecture to store, share, and analyze their research outputs (Fast and Rinner, 2017; Ruediger, 2021). As such, platforms and cloud storage services are becoming central to how scientific data is made open and how it is accessed and preserved. This reality is further bolstered by the open science movement, which aims to provide long-term public access to research data and scientific knowledge by making datasets, research products, and publications available through digital infrastructures (David, 2003; Fecher and Friesike, 2014; Kim, 2019). This study brings empirical analysis of data management plans written by scientists in the United States (US) to this open science debate to examine the role of platforms and cloud storage services in the long-term preservation of scientific data.
Preservation planning that enables access to data in institutional repositories is foundational to open science policy and federal funding priorities in the US (National Academies of Sciences, Engineering and Medicine, 2017; National Academy of Sciences (US), National Academy of Engineering (US) and Institute of Medicine (US) Committee on Ensuring the Utility and Integrity of Research Data in a Digital Age, 2009; National Research Council et al., 1999). Here we present a survey of platform products and cloud services used in scientific data preservation by tracking their appearance in planning documents from the first year that data management plans were required by the National Science Foundation (NSF) in 2011, through 2021. We read data management plans as empirical sources of policy compliance to understand the complex web of practices and institutions that scientists say will be part of their preservation planning for future access. In examining the archiving and preservation practices from a corpus of research data management plans over the first decade of the policy requirement, we observe a shift in scientists reporting that they plan to use platforms and commercial cloud storage products to provide access to research data. At the same time, many of these commercial services challenge the open science values that public access mandates and federal research data management policies seek to address.
This shift in planning with platforms and cloud storage services points to what Birch (2020) and Rikap (2022) characterize as “technoscientific rentierism,” where access to scientific knowledge becomes commodified and shaped by private firms (Birch, 2020; Rikap, 2022). The tension between open science goals and commercial platform interests creates a conflict where efforts to make science more accessible potentially strengthen software service providers and big tech firms providing digital preservation and access. As Mirowski (2018, 2023) argues, the increasing dependence on commercial platforms for scientific data practices like collaboration and preservation turns open science efforts into an extension of platform capitalism, where scientific knowledge and data become assets for extraction and enclosure rather than public goods.
As scientific documents, data management plans (DMPs) tell us what scientists think they will do or reasonably expect their institutions to carry out, including activities for data preservation and archiving (components that the NSF policy requires). But as readers know, plans can change. As such, we argue that DMPs should be considered an anticipatory discursive infrastructure for scientific data management (Ananny and Finn, 2020). While acknowledging the provisional limits of early planning in project timelines and research proposals for funding, we take plans as evidence of scientists’ intended goals for their research output and reception, not as an objective reality of what will or has occurred. As aspirational, anticipatory documents, DMPs may make promises that overestimate future resources or have unrealistic characterizations of constraints and possibilities. Even so, we find these declarations as evidence of compliance with and sustained commitments to policy requirements, because they are reviewed by program officers and peers, and they are integral to grant reporting and project evaluation. Thus, DMPs are rich object-lessons in the rapidly changing landscape of research data activities necessary for preserving open science and simultaneously show how platforms and cloud storage services are reshaping the political economy of scientific knowledge production.
We begin with background literature on research data policy in the context of US federally funded science, showing how these policies support open government, open science, and big data research goals. Following the background, the design, and methods, findings from a thematic analysis of a DMP corpus are presented. Then we discuss how the adoption of platform services in institutions can transform academic knowledge production, and, specifically in this article, the platformization of research data management when platforms and cloud storage services are planned to support scientific data access.
There are limits to these documents as empirical evidence, in that they promise future research activities that have yet to occur. While acknowledging that DMPs are anticipatory documents rather than contractual guarantees, we approach them as legitimate expressions of scientists’ preservation intentions that provide perspective into their expectations as researchers and the capabilities of their institutions. Our analysis distinguishes between broader data management practices during projects and the archival activities for post-award data access by focusing on specific “preservation statements” within DMPs that capture scientists’ specific plans for archiving and making project data accessible. Although preservation statements may be the most prospective components of plans (because they extend beyond project funding periods after research has ended), they provide direct evidence of anticipated preservation infrastructure from NSF-funded scientists (Bennett et al., 2021).
Over the decade of examined plans, we show that scientific research data and project outputs are steadily hosted on corporate-owned cloud architectures such as Amazon Web Services (AWS) and that platforms like GitHub appear integral to scientists’ preservation and long-term access points for post-award project data. The significance of our approach lies in the examination of these occluded documents, offering the first large-scale corpus study of how scientists characterize their data archiving and preservation responsibilities as part of research data management in response to the NSF policy.
Background
The DMP policy (now called data management and sharing plan policy) requires that researchers submit data management plans as part of project proposals for funding to the NSF. Plans are two-page documents that describe the processes for managing the types and formats of data produced during projects, and activities for providing access to data produced from projects. The NSF publishes guidelines for developing DMPs across the foundation's numerous directorates and funding programs, and the guidance covers various strategies for documentation, instrumentation, and data deposit depending on program area (Tian et al., 2021).
Plans are intended to capture many parts of the scientific research cycle from the beginning to the end, and for some knowledge products and data beyond the proposed project itself. Plans include precise descriptions of knowledge production activities from immediate research during the project, as well as data production activities that aim to prepare accessible data for preservation and reuse (Baker and Mayernik, 2020). The NSF advises researchers that providing access should be, “at no more than incremental cost and within a reasonable time, the primary data, samples, physical collections and other supporting materials created or gathered in the course of work” during funded projects (National Science Foundation, 2023). For preservation, agency guidance suggests that proposals discuss, “[p]lans for archiving data, samples and other research products, and for preserving access to them” (National Science Foundation, 2023).
The NSF began requiring DMPs for all funded research in 2011, following President Obama's 2009 open government initiative, which aimed to increase transparency and promote public access to government data (NSF, 2010). The move to require DMPs not only fulfilled legal requirements but supported open science practices and coincided with the “big data” boom and the data empiricism that followed (Borgman, 2015). The turn towards data-intensive science was bolstered by the increasing volume or complexity of digital data, but its broad availability, marking the beginning of big data empiricism driven by open access (Kitchin, 2014). Effective preservation planning in research data management is essential to realize the value of federally funded research, and as many have argued, for government efforts to democratize data access with broader availability for the public (Lane and Potok, 2024). Preservation for access is also critical as scientific data becomes viewed as an asset with potential for reuse beyond its original context (Leonelli, 2019, 2022).
While the NSF policy requires DMPs as part of every funded project proposal, it does not directly address post-award management, compliance, or further evaluation as is common in European research schemes, such as the European Commission's Horizon program (European Commission, 2022; Horizon Europe, 2021). As part of proposals, DMPs are evaluated in closed-door panels by peer reviewers and program officers, and they are not publicly accessible. Guidelines for writing these plans vary by directorate and scientific domain, with some programs offering specific templates or preferred repository lists (Tian et al., 2021). Some scientists may elect to deposit their research products in particular repositories that have their own data management policies or deposit guidelines that can then be incorporated into DMPs. Many research universities in the US also offer data management planning services and data deposit to support researchers, and open source templating tools exist, like DMPTool, to help build comprehensive plans. Even so, we hypothesize that DMPs in the US have received less attention since they are not public records and lack formal standardization across research domains. 1
Making one's research data available is a hallmark of open science (Borghi and Gulick, 2022; Willinsky, 2005). Positive scientific, economic, and political outcomes have all been ascribed to open access to research data (Aspesi and Brand, 2020; Borgman and Bourne, 2022). For scientists, it promotes transparency, enhances careers through data citations, increases interdisciplinarity, reduces duplicate collection efforts, and prevents data misuse (Arzberger et al., 2004; Mosconi et al., 2019). However, meta-evaluations show that researchers’ attitudes toward data sharing and making it publicly accessible are more complex. Sharing data is viewed as “largely positively among researchers,” but most report hitting obstacles and constraints that prevent sharing (Thoegersen and Borlund, 2021). These challenges include resource constraints (both human and technical) and data governance concerns (Birkbeck et al., 2022), reflecting the distinction between knowledge production processes and the additional labor needed for data production to support data reuse (Baker and Mayernik, 2020; Borgman and Groth, 2025).
Public funding for scientific data archives goes back to the 1960s in the US (Bisco, 1966; Shankar et al., 2016). The NSF was one of the first federal agencies to coordinate and support federated scientific data archives in the 1970s, leading to large-scale national research data repositories such as Biological and Chemical Oceanography Data Management Office (BCO-DMO) and the National Center for Atmospheric Research (NCAR). Universities have also built institutional repositories such as the Inter-university Consortium for Political and Social Research (ICPSR) at University of Michigan that provide public and research access to data archives (Downey et al., 2019: 201; Geraci, 1988; Hemphill et al., 2018). Much of today's research data management has been driven by a hub-and-spoke model of infrastructure and computing resources support, where scientists generate project data at their home institution and deposit it in repositories like NCAR or ICPSR, which may then be further funded by federal agencies. As Borgman and Bourne have argued, successful data sharing requires coordination among numerous stakeholders with different roles and responsibilities (Borgman and Bourne, 2022). Information studies researchers have found that institutional arrangements have historically emphasized knowledge production while gradually developing infrastructure for data production as digital technologies like linked-open data and APIs have evolved to support access (Acker and Kriesberg, 2019; Pasquetto et al., 2015).
Academic libraries and research computing units in universities have developed institutional repositories to support long-term preservation and sharing digital data (Acker, 2021; Cragin et al., 2010). These repositories often rely on open source software stacks like DSpace or Fedora and operate within a landscape shaped by digital preservation standards, audit requirements, and emerging principles such as FAIR and CARE (Carroll et al., 2020, 2021; Wilkinson et al., 2016). The NSF and other federal funding agencies have historically supported domain-specific repositories, but institutional investments in data repositories have become more common, particularly as universities are called on to fulfill open access mandates for their researchers (MacDougall and Ruediger, 2024). However, the gap between institutional capabilities and scientists’ expectations is a recurring concern across research data management literature (Birkbeck et al., 2022; Martone and Nakamura, 2022).
Research on DMPs and scientists’ data practices primarily comes from efforts at training for compliance with funding agencies’ submission policies (Cox et al., 2017). Research data managers and data librarians responsible for digital preservation planning and access efforts have evaluated institution-wide samples from university repositories (Bishoff and Johnston, 2015; Briney et al., 2023; Parham et al., 2016; Rolando et al., 2015). Early data curation research found that while scientists desired open, accessible data, they expected more support from their institutions in hand-over, preferring librarians to offboard their data and generate metadata to make datasets accessible (Wallis et al., 2013). But interviews with research data curators and data librarians reveal a friction in offboarding datasets from scientists (Cragin et al., 2010; Donaldson, 2019; Donaldson and Koepke, 2022). Such tensions are exemplary of the “data relations” between institutional actors and community members that may value data—and have various, even conflicting expectations about the work needed to maintain access (Lee and Dourish, 2024). Librarians and research data managers continue to call for more data management training of scientists, lab managers, and graduate student researchers, particularly regarding documentation and metadata in preparation for deposit (Hudson-Vitale and Moulaison Sandy, 2019; Pasek, 2017).
Beyond training scientists in data management, some research data librarians and archivists have called for DMPs to be more accessible as dynamic documents (Praetzellis et al., 2023). Advocates of openly available plans see them as living documents that could be at the heart of a digital research ecosystem, bringing together persistent identifiers and linking relevant information and outputs related to a research project that would all be linked together allowing DMPs to be both machine-readable and machine-actionable public documents (Bakos et al., 2018; Miksa et al., 2019; Simms et al., 2016).
We bring a critical perspective to plans by examining the preservation practices that scientists articulate within their DMPs: that is, the specific statements used to describe preservation tactics to support future access. Our focus on preservation activities described in plans allows us to analyze the data production aspects of DMPs that extend beyond project-based research lifecycles and award funding to examine how platforms and cloud storage services are shaping preservation and access to publicly funded research data. By focusing on the statements about data preservation for access, we can shift the emphasis from training and compliance to understanding how planning shapes access to the long tail of scientific data (Borgman et al., 2016; Wallis et al., 2013).
Research design
This study comes from a multi-phase project investigating DMPs from the first decade of the NSF policy requirement (see Figure 1). Prior analyses of the corpus in the first three phases have identified and evaluated key themes and assessment protocols (Sharma et al., 2023a, 2023b). This article specifically reports on findings from the fourth phase of analysis, which focuses on the appearance of platforms and cloud storage products within plans featuring “Preservation Statements” (hereafter, preservation statements) found in plans (Phase 4). Preservation statements were determined during coding analysis of plan corpus by identifying activities related to data reuse, archiving, or accessing project related data after the project ends (Phase 2). To generate thematic insights and identify trends in digital preservation and data archiving practices, we applied a defined list of platforms and cloud storage services to the preservation statements (Phase 3). As discussed above, a limitation of plans is that they are not confirmed outcomes, however as intended activities for data management, they capture scientists’ expectations of data activities during projects and after their research. Here below we provide more details about the four phases of data collection, preparation, coding, and analysis.

Phases of research design, Phases 3 and 4 are reported on in this article.
Phase 1. Data collection via email campaign
We designed an email campaign to solicit DMPs directly from scientists with funded projects. Using publicly available award information from the NSF awards database (https://www.nsf.gov/awardsearch/) we contacted 5499 scientists through email. The campaign used the Google workspace mail merge feature to customize email requests (Hawksey, 2024). Recipients were asked to reply to the email received with a DMP document attached from a specific awarded project. 1032 attached documents were received, anonymized, and then collected into a document corpus, and verified for suitability for analysis. 2 Table 1 details the overview of the email campaign by response rates, documents received, and DMPs analyzed across five research areas.
Results of email campaign by research area, resulting in a corpus of 976 data management plans (DMPs).
Of the 1032 documents received in response to the email campaign, 976 met the requirements of our protocol criteria for the study. These DMPs were then organized by research program area and year and analyzed using MAXQDA Team Cloud qualitative coding software. The analysis was conducted with a codebook comprising 10 themes; project instruments available online at (Acker et al., 2023).
Phase 2. Theoretical framework for qualitative Institutional Analysis Development (IAD) coding analysis
In Phase 2, we coded each sentence of each DMP into segments using a codebook grounded in institutional analysis of commons resources, known as the Institutional Analysis Development (IAD) framework, inspired by Elinor Ostrom and Charlotte Hess’ extensive scholarship on governing knowledge commons (Hess and Ostrom, 2005, 2006a, 2006b, 2007; Ostrom and Hess, 2005). The NSF's commons approach to open science appears in several policies shaping publicly funded resources. Policies such as the DMP requirement then exert normative pressure to adopt a commons framework for open science. Ostrom and Hess argue that the commons of scholarly knowledge from information and data can be enclosed through privatization, monitoring, encryption, and licensing regimes (2005).
When applied to our corpus, we use the IAD framework to position DMPs as an action arena to understand decisions contributing to a data commons and open science more broadly. The plans describe the common pool resource being shared, delineate how data lifecycles progress, and reveal decisions for governing data sharing within knowledge infrastructures. This approach allows us to follow various scientific communities and their rules, gather specifics about the materiality and formatting of the resource unit (datasets or data-related products), and examine the resource system of the infrastructure used by scientists for storing, sharing, and reusing data, as well as the impacts of adopting such systems. Iterative initial coding with the research team resulted in a final set of 10 codes. After confirming 70% intercoder reliability between two team members with a subset of DMPs, the corpus was subsequently coded over 2 years.
Phase 3. Defined list of actors: Platforms and cloud storage services
In Phase 3, we generated a defined list of actors as part of the qualitative coding analysis, which included repositories, institutions, professional organizations, and platforms. Using a sample of 150 DMPs from the corpus, 30 plans from each of the five programs were randomly selected to build a selected list of defined corporations, platform products, and cloud storage services. Further coding analysis and literature review resulted in a comprehensive list of known repositories, institutional data services, and platforms used by scientists. We then classified these services as “for-profit” and “not-for-profit.” The complete list of platforms and cloud services classified is shown in Table 2.
Defined list of platforms and cloud storage services.
Asterisk denotes services currently owned by Microsoft.
It is noted that this classification does not encompass the complete differentiation between these platforms and cloud services. For example, not-for-profit services like Dspace can be self-hosted using open software or can be used through a registered service provider, while services like Dryad can be integrated with other repositories through their API, while others, like Zenodo, are solely hosted at CERN. For-profit services like GitHub have social features that track reputation and impact (e.g. Commits, Badges, and Stars), while Google Apps offer project management features, while others like Box and DropBox act more like traditional cloud storage services for file management.
Phase 4. Platforms and cloud storage services in preservation practice statements
In Phase 2, the 976 plans were coded at the sentence or “segment” level using the IAD framework and codebook. In this first cycle of descriptive coding, we found that 743 plans (76.13% of the total corpus) contained 2535 segments coded as “Preservation Statements.” After reviewing the first cycle of coding segments, we determined that platforms and cloud storage appeared in both internal data practices, as well as external data management efforts for preservation and archiving. For the second cycle of analysis presented in this article, we used pattern coding to identify emergent patterns (Flick, 2014). Using the defined list generated in Phase 3 (Table 2), we examined appearances of these 21 platform products and cloud storage services within the 2535 Preservation Statement segments.
Findings
Here we discuss findings that emerged from 743 of the DMPs that contained statements about preservation practices. The appearance of specific statements about preservation practices demonstrates that digital preservation and data archiving are central to many researchers’ plans over time, as well across research programs. However, the percentage of DMPs with preservation statements varied from year to year, from a low of 64.00% in 2021 to a high of 82.20% in 2018 (Table 3). We begin by sharing some descriptive statistics from pattern coding the DMPs with the preservation statements, then we share themes drawn from qualitative coding of the preservation statements themselves.
Preservation practices statements found in the data management plan (DMP) corpus 2011–2021.
Pattern coding with the defined list of platforms and cloud services resulted in some services appearing more than others in preservation statements, as well as some services appearing at higher rates in certain research programs. We found a marked increase in mentions of cloud services in DMPs over time. Table 4 demonstrates this upward trend from 2015 onwards, with plans that mention cloud services rising to 43.75% in 2021.
Data management plan (DMPs) with preservation statements by year and the proportion that reference platforms and cloud service providers for preservation.
A steady rise in platforms and cloud storage service providers for preservation appears over the decade examined. We note that 2011 has a high mention of cloud service providers because of the over-representation of DMPs from the Division of Biological Infrastructure document group for that year (58.54% of DMPs in 2011 were from DBI). Biological Infrastructure is a research program that supports computationally-intensive research and supports research infrastructure projects, and on average plans used platforms and cloud service providers for preservation more than any other research group we examined (16.31% of Division of Biological Infrastructure DMPs contained a cloud service providers overall while only 10.25% of all the other program areas contained a cloud service providers in their preservation statements).
We further analyzed whether DMPs discussed for-profit or not-for-profit services in their preservation statements (Figure 2). For-profit cloud services, such as Amazon Web Services or GitHub, rose over time, particularly after 2016. GitHub and Dryad make up the most frequently seen services across preservation statements overall, followed by AWS. But Dryad, the second most common service in preservation statements, is exclusively seen in two programs: the Division of Ocean Sciences and the Division of Biological Infrastructure. Dryad is often used as a public repository for data or code that is associated with a publication. Scientists who mention Dryad often talk about how it is publicly accessible, widely used in their community, and its long-term prospects.

Distribution of cloud service providers mentioned in data management plan (DMP) preservation statements by year (percentages of annual totals).
These trends suggest that for-profit cloud services are becoming integral to data access and digital preservation strategies (Table 5), while not-for-profit services continue to play a significant role, especially for institutional repository offerings. We can expect for-profit services to increase in research data management and continue to appear in scientists’ plans.
Appearance of for-profit or not-for-profit platforms and cloud storage services in preservation statements by year.
DMPs: data management plans.
Persistent platforms and assumed resilience
GitHub emerged as the most frequently cited platform across the corpus, reflecting its critical role in scientific research for versioning, code access, and software development. Established in 2008 and acquired by Microsoft in 2018, GitHub now serves over 100 million users, with support for research, government, not-for-profit, and education sectors. A closer examination of the DMPs that discuss for-profit cloud services reveals that GitHub is a key driver of this trend compared to the other for-profit services we identified. Over the decade analyzed, GitHub is discussed in 38 DMPs being used for preservation, while the other 12 for-profit services make up 32 other instances of for-profit platforms being used for preservation. This means that GitHub makes up 54% of cases of scientists planning to use for-profit platforms for data preservation in their DMPs.
From the preservation statements that mention GitHub, we find that it is described as a comprehensive digital preservation solution as well as an access point for research-related outputs, including code, datasets, documentation, and publications. For example, in this plan from a team of biologists, their GitHub project repository will act as a clearing house for all outputs of the project's API: GitHub comprises our long-term preservation and distribution mechanism. The source code, change log, binary installers, installation instructions, API documentation, release notes, links to papers describing the library and API, and other information and files are available. (Division of Biological Infrastructure, 2014)
Project plans such as this one, that rely on GitHub, speak to the faith that researchers have in the platform as a long-term preservation solution. It is trusted to preserve project outputs from development to release, but also a reliable mechanism for data distribution too.
Throughout preservation statements, we see the language of long-term availability that assumes platforms like GitHub and Google Drive will be sustained indefinitely, much like institutional repositories. For example, we found across the corpus that scientists applied the same phrase “preserved indefinitely” to both institutional collections and GitHub alike. We take this to mean that there are blurred lines between scientists’ expectations of institutional and platform-driven preservation services. While some plans describe government, university repositories, and archives, permanent preservation implies the responsibility of institutions to maintain collections for governance, posterity, and knowledge: “Data in the USGS [United States Geological Survey] database will be preserved indefinitely.” (Division of Ocean Sciences, 2014) “Data deposited at the [Redacted – Museum] will be preserved indefinitely in the museum's archive facility.” (Science and Technology Studies, 2019)
We note that when the same phrase is used in statements featuring GitHub, but it is primarily leveraged as a repository because of its version control system: “The codes used here will be preserved indefinitely through version control on GitHub” (Division of Ocean Sciences, 2019). Preservation here is dependent on the platform's technical infrastructure, and scientists, like other platform users, rely on GitHub's corporate services for publishing and accessing code. Despite assumptions about GitHub's preservation resilience, backup plans that account for contingency appear in DMPs too. A nearly identical quotation from another team of oceanographers 2 years later includes the caveat: “The codes will be preserved indefinitely through GitHub's version control, assuming the platform's sustained operation” (Division of Ocean Sciences, 2021).
Occasionally, plans characterize data preservation and availability as contingent on the survivability of the platform. For example, a Science and Technology Studies project notes, “A substantial portion of the project will be replicated and versioned on GitHub and continue to be available there as long as the platform survives” (Science and Technology Studies, 2014). Such statements speak to scientists’ pragmatism in the face of platform options, wherein they assume GitHub's resilience but acknowledge its potential vulnerabilities. This caution regarding GitHub's assumed resilience reflects an understanding of the power (and possible pitfalls) in planning for data preservation with platforms.
Uncertain custodianship
All data management plans discuss future activities and identify responsible parties to support access to data and project outcomes. However, as the timelines of plans stretch further beyond the immediate project lifecycle, the specifics around who will sustain such preservation efforts become less detailed. While research data management workflows and custodial roles are mentioned, details regarding long-term commitments from institutions and platforms alike are often incomplete or ambiguous. A case in point: one plan says that the “[p]rimary responsibility for curating and preparing the data for archiving rests on the Data Librarians at the University” (Secure and Trustworthy Cyberspace, 2013). Although this suggests a clear designation of responsibility to the university libraries, there is a lack of detail about the sustainability of these efforts or whether data librarians are aware of being designated as such.
Some PIs express uncertainty about future custodianship and continued institutional support in plans. In a Science and Technology Studies project, the PI highlights the uncertainty around institutions’ support: “The University repository guarantees the data's survival, but live access depends on continued institutional backing” (Science and Technology Studies, 2014). Other plans shift custodianship to external repositories without clarifying long-term control: “The products of this research may also be archived at the [REDACTED] repository following their policies” (Secure and Trustworthy Cyberspace, 2018). While this indicates potential institutional partnerships, custodianship arrangements are ambiguous.
We observed another trend where custodianship is wholly left to external repositories’ control. In one case, a PI states, “Once these data sets have been submitted for archiving, the repository will have primary responsibility for long-term data curation and access control” (Secure and Trustworthy Cyberspace, 2015). Such reliance on repositories speaks to researchers’ institutional expectations but also raises questions about the long-term preservation capacities of external repositories or whether such designated responsibility should be confirmed with data librarians or research data managers. This is especially so when PIs assume that simply submitting data for deposit ensures its indefinite preservation and custody.
We also find that multiple temporalities are embedded in plans, and that scientists tend to envision ideal use cases, and sometimes anticipate contingencies (Tian, 2025). When DMPs describe multiple timelines and potential preservation pathways, they often acknowledge the uncertainty of institutional or platform support, placing future responsibility on unspecified stakeholders. For instance, a PI discusses long-term plans for storage, noting, “If data collection is to be terminated or transitioned … the data archive will be brought up to date, and then backed up on two separate hard drives” (Division of Civil, Mechanical, and Manufacturing Innovation, 2015). Although this plan outlines a specific procedure for data archive transition, it assumes that future actions with two hard drives (such as securing longer-term storage or migrating project data) will occur without providing concrete steps for delegating those preservation activities. Again, it is less clear who will be responsible for undertaking such stewardship.
Many plans reflect an implicit assumption that institutional backing or external platforms will probably evolve to meet the future needs of the project data. For instance, “Data submitted to the [REDACTED] repository will become the responsibility of [REDACTED] College Library and will be retained for the life of [REDACTED] College unless removed to a satisfactory discipline-specific repository” (Division of Biological Infrastructure, 2020). Here, preservation is directly tied to the institution's longevity, yet the underlying assumption is that the institutional actors will ensure appropriate data management and transfer.
Our own assessment of the DMPs is that institutional custodianship is assumed to be stable, despite potential risks and constrained resources. For instance, one STS project suggests that as the digital preservation community grows, “[REDACTED] Library will modify its succession plans so as to continue in its role as a trusted digital repository” (Science and Technology Studies, 2019). Such preservation statements reflect archival optimism but signal the uncertainties surrounding institutional commitments to digital preservation, which often depend on external factors beyond the control of the PIs. Both succession planning and uncertain custodianship uncover scientists’ concerns about the sustainability of data preservation efforts while designating responsibility to both platforms and cloud services or institutions that license and pay for their services.
DMPs with multiple timelines tend to describe ideal use cases for data access and reuse. But plans carry further uncertainties, where future costs are assumed to be incurred by the researcher's university or home institution. Scientists may mention the lack of support from their own university repositories, but leave hopeful possibilities of future preservation actions: While the University repository can guarantee the survival and access of the data generated by the project in perpetuity through [redacted], live access to the digital edition will be maintained as long as the University and its Libraries can provide institutional support to match outside funding. Should the University no longer be capable of supporting access to the production site, we will take snapshots of the website and the database and store these in dark storage until future interest can revive it. (Science and Technology Studies, 2014)
As timelines extend farther into the future, plans become less precise, and it becomes evident that PIs assume future institutional curatorial efforts (and resources) that are far from guaranteed. Such statements, as above, reflect a trend where scientists express hope that future institutions or stakeholders, whether institutions or corporate firms, will bear the responsibility for preservation and carrying data forward, even as they acknowledge the challenges of sustaining access and threats to funding.
Platform services in research data management
DMPs show that key platform actors are increasingly relied upon for the management of data, with 121 (16.28%) plans in our overall corpus mentioning platforms and cloud storage services as a part of their preservation and access strategies over the decade we examined. For-profit services have increased in prominence from 2015 to 2021, with a sharp rise in the last few years, growing from 8.97% in 2017 to 25% in 2021, nearly tripling their presence in preservation statements. Scientists frequently reference platforms like GitHub and Google in DMPs, and on occasion, anticipate the possibility of migration due to the loss of services. For example, one plan from a computational biology team reflects their previous experiences adapting to platform changes: “Should GitHub discontinue services, we will migrate our project to another hosting site, as was previously done when moving from Google Code” (Division of Biological Infrastructure, 2014). Migration planning is an example of realistic preparedness in the face of platform effects on research data management. It highlights dependencies on platform services in research workflows that can get disrupted, as well as the realities of firms’ discontinuing services, and the need for scientists to have adaptable strategies for platform-driven data management.
Where plans may mention platform hurdles and strategies to mitigate them, some plans may contain unrealistic claims of unlimited, never-ending cloud storage. While contingency planning and assumptions about unlimited resources appear to be in contradiction, they both speak to a platform effect on research data management and scientists’ expectations of these services for preservation. For instance, one DMP from the Division of Biological Infrastructure states, “The University [Redacted] offers free access to Office 365 OneDrive (5 TB limit) and Google Drive (unlimited) for long-term cloud storage” (Division of Biological Infrastructure, 2018). These assumptions for free access and unlimited storage reflect scientists’ expectations of platform services at their institutions and universities. But these assumptions have proven to be inaccurate for most universities, as seen in Google's rollback of unlimited storage for educational institutions in 2022 (Hickey, 2023). Shifts in platform service offerings expose the challenges researchers face when relying on external services for long-term data management and demonstrate how platforms shape current practices and future preservation strategies for research data management. While universities may outsource their digital tooling and data storage to external platforms and cloud services, researchers cannot trust that these storage services will remain free or unchanged forever.
Discussion: Platform services and implications for open science
As platforms become the dominant mode of data access across many sectors of society, we are witnessing data management workflows conforming to platform services, interfaces, and cloud storage architectures in scientists’ plans. This is the process of platformization, where platforms like GitHub and AWS extend their reach to reshape access infrastructures, transforming both scientific data practices of sharing and research data repositories alike (Plantin et al., 2018; Plantin and Thomer, 2023). Science and technology studies scholars assert that open science efforts can lead to the platformization of science, where platforms re-engineer approaches to accessing and managing data, while making academic research scientists more and more dependent on academic publishing conglomerates, research sharing platforms, cloud storage repositories, and collaboration platforms to coordinate labor and circulate research data as assets (Birch, 2020; Plantin et al., 2018; van Dijck et al., 2018).
Science historian Philip Mirowski argues that open science is the continuation of neoliberalism in academic research (Mirowski, 2023). Open science, he writes, “seeks to maximize data revelation as a means to eventual monetization” (Mirowski, 2018: 193). Similarly, Levin and Leonelli argue that access, dissemination, and reuse of scientific data can reinforce commercial capture: “[w]hen openness is codified in policy, it not only enacts particular things as open and closed but also performs certain values, such as defining some research outputs and practices as more valuable than others” (Levin and Leonelli, 2017: 296). The path of research data becoming assets follows broader patterns of technoscientific capitalism, where knowledge itself becomes subject to rent-seeking behaviors and market logics. But the conceptualization of research data as assets depends on storage and preservation infrastructures to maintain their value over time. Thus, cloud storage becomes the “hidden abode” of value in data economies, where preservation services transform from institutional responsibilities into commercial software offerings (Banoub and Martin, 2020: 1101).
Our findings show that scientists are increasingly planning to use platforms and cloud storage services for data preservation activities, with mentions of for-profit services nearly doubling during the second half of our data collection (26 mentions of for-profit services for preservation from 2011 to 2016, to 50 mentions of for-profit services for preservation from 2017 to 2021). GitHub appears most frequently, accounting for 54% of all for-profit platform mentions. This trend reflects what we describe as “the platform effect,” where successful platforms consolidate their position by offering initially free or low-cost services, managing network effects, and deterring disintermediation. Cloud storage services transform preservation into a subscription-based service rather than a long-term institutional commitment with designated custodians like archivists and librarians. These strategies reinforce the data-as-asset model, creating ongoing financial relationships around research data produced with public funding but then made accessible by platforms and cloud storage (Acker, 2025).
The arrangements articulated in DMPs, from file formats to cloud storage services, reflect interactions between scientists’ data practices, platforms, and commercial cloud services used by institutions. This has been observed in other empirical accounts of universities as knowledge infrastructures, both descriptive and critical (Katz, 2010; Sørensen and Traweek, 2021). Professional digital preservation organizations and institutions themselves have begun to report the costs and benefits of incorporating platforms and commercial cloud storage for digital preservation (Rieger et al., 2022; Ruediger, 2021). For over a decade, the National Digital Stewardship Alliance (NDSA) has periodically surveyed its members (including universities, government agencies, and non-profit organizations) about their digital preservation storage practices. The goal of these comparative surveys is “to develop a snapshot of storage practices within the organizations of the NDSA and insight into how those practices change over time” (Gallinger et al., 2017). The majority of NDSA members report using Amazon Web Services, Microsoft's Azure, Google, and DuraCloud, among others (2019 Storage Infrastructure Survey, 2020: 21). Cumulatively, the NDSA storage surveys from 2011 to 2023 have found a steady increase in cloud storage service providers, with the use of commercial cloud storage services increasing to 55% in 2023 (2019 Storage Infrastructure Survey, 2020; Allen et al., 2024; Altman et al., 2013; LeFurgy, 2012). Our findings follow on from the evidence gathered from the NDSA storage surveys, reflecting larger patterns from research institutions leveraging cloud storage and platforms for both data preservation and access.
The embedding of platform and cloud storage services within university and researchers’ data archiving practices carries profound implications for preserving science data in the near term, but for accessing it in the far future too. Cloud storage and platform services transform how research data is valued and circulated. They shift scientific practice from community-oriented sharing that builds knowledge commons toward market-driven approaches where data becomes an asset (Leonelli, 2019).
Corporate cloud storage and platform services offer research universities scalable, cost-effective solutions for digital preservation, ostensibly democratizing access and enhancing compliance with funding policies. For example, Miller described the challenges smaller and mid-sized institutions face when budgeting for digital preservation and the advantages for adopting platform services when confronting the technical complexities of multiple workflows (Miller, 2015). For organizations with limited resources and technical staffing, adopting cloud-based digital preservation and storage products that scale easily and can streamline layers of ingest, processing, and maintenance workflows into one platform solution can be a boon. But the benefits of vendor-hosted and corporate third-party solutions come with well-known tradeoffs to libraries and archives. By streamlining services, platforms’ centralized control over collections management introduces dependencies that may further undermine institutional autonomy or expose organizations to costly lock-in effects (Rosenthal, 2019).
Another dimension of platform services is asset externalization, where data assets are made accessible online or rendered machine-actionable through cloud infrastructures. Platforms that externalize assets via services like virtualization depend on persistent data storage and active data management (Narayan, 2022: 923). Narayan and others have argued that platform capitalism is driven by hyper-scaled computing infrastructures, concentrating cloud storage among a few dominant firms (Narayan, 2022; Thylstrup et al., 2024). Following this view, the externalization of scientific data alters preservation from maintaining scholarly integrity to enabling ongoing value extraction through controlled access points via scaled storage infrastructure and cloud services. Platforms can then act as gatekeepers by reducing technical barriers, but may compromise institutions’ ability to manage reuse or track impact if they outsource storage, too (Mattern et al., 2024).
The platformization of research data management has been further theorized in Plantin and Thomer's work (2023), interviewing curators and research data managers about their experiences with using Figshare for institutions. The adoption of platforms like Figshare to support data management workflows moves repositories, libraries, and archives away from historically community-developed and open source infrastructures typically used towards subscription-based models, where they have less control over discoverability standards. They find that the adoption of platform services for repository data management may increase workers’ satisfaction by offloading technical services to the platform, but the tradeoff for increased usability is less institutional control over description and discovery. This shift to platform-driven data management threatens the autonomy of institutions, making them more dependent on external vendors for maintenance, updates, and technical support.
Conclusion: The platform effect
At issue here is the platformization of scientific knowledge, but more directly, access to data from publicly funded scientists in the public interest. Our findings show how reliance on for-profit platforms illustrates what we have termed the platform effect. Presently, policies around open access sharing, data management, and digital preservation are tied to requirements that result in DMPs are generated at the beginnings of a project lifecycle. The platform effect in data management planning impacts scientists’ preservation strategies, embedding risks into the knowledge infrastructures that researchers and institutions depend upon for data access.
Suchman famously theorized that plans were both structured and structuring artifacts (Suchman, 1987). Here we find that DMPs are structuring documents that serve a dual purpose for accessing scientific data: plans describe types of data produced by research projects, but they also describe the strategies and technologies that will be used for managing future access through preservation planning. Often, plans enroll and leverage existing institutional support and data management resources afforded to scientists as members of research institutions and universities (Bishoff and Johnston, 2015; Carlson, 2017). Our analysis of DMPs from the first decade of the NSF policy indicates a steady increase in commercial platforms and cloud storage services in planning to archive scientific research data for access alongside institutional services. In time, this shift is likely to involve reliance and enclosure that challenge principles of open science and ensuring open access to research data because of the incurring costs and subscription pricing models that most of these commercial platforms employ.
How might we productively disrupt situations that are structured in DMPs now and going forward? From a long-term digital preservation perspective, growing dependence on corporate-owned platforms for research data management is problematic. There are no guarantees that these firms will prioritize the preservation of open scientific research over commercial interests or that these services won’t sunset or shutter. The appearance of platforms and cloud storage services in DMPs thus reveals a broader transformation in the political economy of science and institutional data archiving, where the infrastructures supporting data management are driven by the imperatives of platform capitalism and signal the likelihood that data access will soon be subject to rents, among other market forces. The implications of the platform effect for research data management represent a fundamental transformation in the political economy of science, as well as the means of opening access to publicly funded research data. This reconfiguration positions researchers not as independent scientists, but as users within a data economy dominated by corporate interests.
Footnotes
Acknowledgements
The authors wish to thank the scientists who shared their plans with our research team and the National Science Foundation, Science and Technology Studies program, for supporting our research. We would also like to acknowledge the contributions of Dr Sarika Sharma for the corpus preparation; the initial design of qualitative coding analysis; and data analysis.
Ethical approval and informed consent statements
This research has been IRB approved under an exempt determination for Protocol Number 2020-05-0017 for The University of Texas at Austin.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research has been supported by the US National Science Foundation, Award numbers 2020604 and 2429325.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
