Mapping AI performance metrics reported on the websites of digital evidence synthesis tools: A cross-sectional study
Evidence synthesis plays an important role in complementary and integrative medicine. Digital evidence synthesis tools increasingly use artificial intelligence (AI) to support researchers in handling large volumes of scientific literature. This cross-sectional study examines how the performance of AI is described on the public websites of these tools. By mapping reported performance metrics to different stages of the evidence synthesis workflow, the study offers an overview of what information is currently communicated to users and where important gaps remain.
Background and rationale
Evidence synthesis, such as meta-analyses, is essential for informed decision-making in healthcare and policy. For complementary and integrative medicine, evidence synthesis is particularly important due to the heterogeneity of interventions, outcomes, and study designs. In recent years, digital evidence synthesis tools including AI have promised faster and more efficient workflows. However, information of the performance of the tools is not always available.
Understanding if and how AI performance is presented in the public websites of digital evidence synthesis tools is important for users who rely on these tools to judge their suitability and reliability This study focuses on what is reported publicly, not on whether the claims are accurate or comparable across tools.
Study design
This research uses a cross-sectional study design. Websites of AI Tools will be analyzed to systematically extract AI performance information from publicly available sources. Reported metrics will be collected exactly as described by the tool developers, without reinterpretation or recalculation.
All extracted metrics will be then mapped to specific stages of the evidence synthesis workflow, such as literature search, screening, data extraction, and synthesis. This mapping helps clarify where AI performance is most often emphasized and where reporting is limited or absent.
Implications for users and developers
By providing an organized overview of how AI performance is communicated, this study helps users better understand the limitations of website-based claims. It also highlights opportunities for developers to improve clarity, consistency, and transparency when reporting AI performance. Ultimately, more standardized and validated reporting could support more informed adoption of digital evidence synthesis tools.
Project team
- Claudia Witt (Principal Investigator)
- Jesús López-Alcalde (Contact)