The problem
Science moves forward when results can be checked and built on. In the life sciences, that check often fails before it starts. When researchers set out to repeat 193 experiments from 53 high-impact cancer papers, none was described in enough detail to repeat without asking the original authors, and the raw data was public for only 4 [1]. Of the 50 experiments they could repeat, the median effect was 85% smaller than first reported [2]. Companies see the same gap. At Bayer, published data fully matched in-house results in only about 20–25% of 67 target-validation projects [3], and Amgen scientists confirmed the findings of only 6 of 53 landmark preclinical cancer papers [4].
Much of the gap sits in the computational record itself. In published research data packages, 74% of R files failed to run without error [5]. In a large sample of Python notebooks on GitHub, 24% ran without errors and 4% gave the same results [6]. Even a spreadsheet can change the record: 30.9% of articles with supplementary Excel gene lists contained gene names that Excel had silently changed [7]. The cost is large. An estimated US$28 billion a year is spent in the US on preclinical research that is not reproducible [8].
In numbers
Cancer experiments with public raw data
4of 193
None could be repeated from the paper alone.
Research R files that failed to run
74%
Code shared with published data packages, run again.
Drug-target projects where published data matched
20–25%
67 target-validation projects at Bayer, checked against in-house results.
AI agents now make this work much faster, and they add errors that are harder to see. In 2023 tests, 55% of the citations from GPT-3.5 and 18% of those from GPT-4 were invented [9]. In a 2025 study, AI summaries of research were nearly five times more likely than human summaries to overstate a study's conclusions [10]. On real bioinformatics analyses, agents built on GPT-4o and Claude 3.5 Sonnet reached 17% accuracy on open answers [11]. Researchers notice: 64% now worry about inaccurate AI output, up from 51% a year earlier [12].
Newer models and better tools help, but they do not close the gap. With the same model, Claude Opus 4.5, the share of papers whose results an agent could reproduce rose from 42% to 78% when only the agent harness changed [13]. On full computational-biology studies, the best of 13 frontier models scored 0.48 out of 1 in August 2026 [14]. Asked for references without web search, GPT-5 invented 51% of them [15]. And an auditor that read only an AI-written paper caught 55% of planted flaws; with the run logs and code, it caught 82% [16].
The common thread is a missing record. A result that cannot be traced to the data, code and steps behind it cannot be checked, and a result that cannot be checked cannot safely be built on. As AI raises the number of results, the gap grows.
What we are doing
BioVantageLab builds the missing record into the work itself. Each finding should keep its source: the session, the file and the tool call that produced it. Findings should connect to what a project already knows, so new work builds on checked results instead of starting from zero.
When results disagree, the disagreement should stay visible. We do not claim to decide what is true; we make it possible to check. And you stay in control: your work stays on your computer, and data leaves it only when you pick a cloud model.
Sereh, the research workspace we are building, is our first tool to do this. See how Sereh works →