Abstract

Keywords
We read with interest the recent methods paper by O’Keefe et al. discussing the evolving role of information specialists in the era of large language models (LLMs) (O’Keefe et al., 2026). The article provides a thoughtful conceptual overview of how traditional information retrieval skills may intersect with prompt engineering, context engineering, and related AI-assisted approaches within evidence synthesis workflows. The discussion is timely and important given the rapidly expanding use of generative artificial intelligence in systematic reviews, underscored by the RAISE recommendations and the joint position statement published by Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence (Flemyng et al., 2025).
O’Keefe and colleagues conclude that the skills required to develop and run complex searches are transferable to LLM engineering with appropriate continuing professional development (CPD). Such CPD-based structured learning opportunities, particularly in the context of LLMs and Artificial Intelligence (AI), would help information specialists update their professional competencies as evidence synthesis methods evolve. Building on the conceptual foundation laid by this article, which successfully highlights parallels between information retrieval and prompt engineering at a theoretical level, the next step is to consider how these transferable skills can be operationalized in practical, transparent, and evaluable workflows, particularly for researchers with limited access to methodological expertise.
We extend this discussion by proposing one practical application of such CPD. Information specialists trained in prompt engineering can help design, refine and evaluate reusable, expert-informed prompt-based workflows for evidence synthesis. Such workflows that encode information-retrieval rules into a structured guidance note could instruct an LLM to perform a narrow, clearly specified, and human-verifiable task of translating user-supplied search terms (controlled vocabulary terms and free-text keywords) into database-specific search strategies with appropriate Boolean logic, field tags, and syntax.
This is different from asking a Generative AI (GenAI) system to independently retrieve search terms, design a search strategy or retrieve literature through an opaque search-and-retrieval interface, which would increase the risk of bias and hallucinations in the absence of transparency on the reasoning behind such tasks. Recent evidence supports the need for this boundary. In a systematic review of GenAI use in evidence synthesis, Clark et al. (2025) included 19 comparative studies, of which three evaluated GenAI for literature search tasks. In these search-related studies, recall ranged from 4% to 32%, indicating that GenAI tools missed 68%-96% of relevant studies identified by manual human searching. Therefore, the authors concluded that researchers should not use GenAI tools for searching in their current form (Clark et al., 2025). Similarly, Lieberum et al. (2025) in a scoping review of LLM applications in systematic reviews, found that literature search was the most frequently addressed SR step, reported in 15 of 37 included LLM articles. However, they also noted that the literature search was the step most frequently judged as non-promising by study authors, and concluded that many LLM approaches remain exploratory and are not ready for direct transfer to research practice without adequate human supervision (Lieberum et al., 2025).
In line with these findings, our proposed use of LLMs is not autonomous searching, independent term generation, or opaque record retrieval. Instead, while using the reusable, prompt-based workflow, the LLM, guided by the expert-designed “chain-of-thought” prompt (which forms the backbone of this workflow), only helps translate user-provided terms into syntactically accurate search blocks. At the same time, the reviewer remains responsible for defining the review question, compiling appropriate terms, checking the output, running the search in each database and documenting the process. Such a structured workflow for constructing search strategies helps ensure that the task remains human-verified and transparent, in line with the RAISE recommendations (Flemyng et al., 2025). This structuring is also consistent with the distinction documented by O'Keefe et al. between structured prompt engineering and informal “vibing” because, for systematic review searching, the goal is not casual, conversational interaction with an LLM, but a documented, constrained instruction process that can be checked, repeated, and reported (O’Keefe et al., 2026).
Reusable prompt-based workflows could, thus, provide a pragmatic bridge between expert-led searching and unsupported searching by non-specialist reviewers. Such an application is not proposed as a replacement for information specialists, but as a way to equitably extend information-specialist expertise to reviewers working in settings where such expertise is not readily available. This is especially relevant for low- and middle-income country (LMIC) institutions. Many researchers in these settings conduct systematic reviews on locally important health, environmental, and policy questions, but lack routine access to trained medical librarians or expert information specialists. The recommendation to involve an information specialist remains methodologically sound, but it is not always feasible. As a result, non-specialist reviewers often need to construct search strategies on their own. Cochrane guidance emphasizes that systematic review searches should be sensitive, reproducible, and tailored to the review question, covering core sources such as CENTRAL, MEDLINE/PubMed, and Embase, while also considering trial registries, specialized registries, national or regional databases, subject-specific databases, citation indexes, and other supplementary sources (Lefebvre et al., 2025). This requires reviewers to handle controlled vocabulary, free-text terms, Boolean logic, truncation, field tags, and database-specific syntax correctly. Lack of familiarity with the multitude of databases, syntactic knowledge, and clarity of concepts inherent in creating a high-quality search strategy can result in errors at this stage, reducing search sensitivity, compromising reproducibility, and weakening the credibility of the evidence synthesis.
We have previously attempted to explore this particular syntax-generation application of AI using a reusable two-part, prompt-based workflow in which fixed instructions encode database-specific search rules and an editable section contains user-supplied concepts, controlled vocabulary terms, and free-text keywords (Yadav et al., 2025). While this initial reusable prompt-based workflow was developed as a response to the prevalent access gap in LMIC settings, its further refinement and evaluation would benefit greatly from the expertise of information specialists trained in LLM engineering through CPD, as envisaged by O’Keefe and colleagues, before the broader evidence synthesis community can practically utilize such a strategy.
RAISE guidelines state that evaluation and validation studies are critical to determining whether an AI system or tool performs adequately for evidence synthesis in general and for the specific topic under review (Flemyng et al., 2025). In this context, information specialists trained in LLM engineering could evaluate whether a reusable prompt-based workflow preserves all user-supplied terms, avoids hallucinated terms, applies database syntax correctly, produces reproducible outputs across LLM versions and sessions, and remains usable for reviewers with limited search experience. Where feasible, workflow-generated strategies that combine AI and non-expert human reviewer roles could be compared with strategies developed by expert information specialists using known relevant study sets, sensitivity, precision, syntax-error tracking, and user-time requirements. This could be done using emerging evaluation frameworks such as Cochrane Evaluation of (Semi-) Automated Review Methods (CESAR), which proposes an adaptive platform study-within-a-review design to assess AI tools for selected Cochrane review tasks using accuracy, efficiency, response stability, error impact, and usability metrics (Gartlehner et al., 2026). Experts could help improve the handling of proximity operators, database-specific wildcards, and controlled vocabulary mapping through more in-depth context engineering and even adaptive prompting with validation checks.
While applauding the evolving role of information specialists envisioned by O’Keefe and colleagues, we suggest that CPD should include not only conceptual training in LLM engineering but also practical training in designing, refining, reporting, and evaluating reusable prompt-based workflows. Such training would help information specialists use LLMs in their work and enable them to extend their high-quality search expertise to broader communities of reviewers. For LMIC researchers, these workflows could improve search quality where expert collaboration is unavailable, while still preserving the principles of human expertise, transparency, reproducibility, and methodological accountability in evidence synthesis.
Footnotes
Author Contributions
Conceptualization: TT, VY; Methodology: TT, VY, UKM, YDS; Writing - original draft: TT, VY; Writing - review & editing: TT, UKM, YDS; Supervision: TT; Final approval of manuscript: TT, VY, UKM, YDS.
Funding
No specific funding was received for this work.
Declaration of Conflicting Interests
The authors declare no conflict of interest.
Data Availability Statement
No datasets were generated or analyzed for this commentary.
AI Use Disclosure
AI tools were used only for language editing and formatting support. All intellectual content and final approval were provided by the authors.
