BioVantage Lab Request access

The problem

Much of life-science research cannot be traced or checked.

Published results often cannot be re-run or traced to their data, and AI is widening the gap. Here is the problem, and what we are doing about it.

The problem

Science moves forward when results can be checked and built on. In the life sciences, that check often fails before it starts. When researchers set out to repeat 193 experiments from 53 high-impact cancer papers, none was described in enough detail to repeat without asking the original authors, and the raw data was public for only 4 [1]. Of the 50 experiments they could repeat, the median effect was 85% smaller than first reported [2]. Companies see the same gap. At Bayer, published data fully matched in-house results in only about 20–25% of 67 target-validation projects [3], and Amgen scientists confirmed the findings of only 6 of 53 landmark preclinical cancer papers [4].

Much of the gap sits in the computational record itself. In published research data packages, 74% of R files failed to run without error [5]. In a large sample of Python notebooks on GitHub, 24% ran without errors and 4% gave the same results [6]. Even a spreadsheet can change the record: 30.9% of articles with supplementary Excel gene lists contained gene names that Excel had silently changed [7]. The cost is large. An estimated US$28 billion a year is spent in the US on preclinical research that is not reproducible [8].

In numbers

Cancer experiments with public raw data

4of 193

None could be repeated from the paper alone.

Errington et al., eLife, 2021

Research R files that failed to run

74%

Code shared with published data packages, run again.

Trisovic et al., Scientific Data, 2022

Drug-target projects where published data matched

20–25%

67 target-validation projects at Bayer, checked against in-house results.

Prinz et al., Nat Rev Drug Discov, 2011

AI agents now make this work much faster, and they add errors that are harder to see. In 2023 tests, 55% of the citations from GPT-3.5 and 18% of those from GPT-4 were invented [9]. In a 2025 study, AI summaries of research were nearly five times more likely than human summaries to overstate a study's conclusions [10]. On real bioinformatics analyses, agents built on GPT-4o and Claude 3.5 Sonnet reached 17% accuracy on open answers [11]. Researchers notice: 64% now worry about inaccurate AI output, up from 51% a year earlier [12].

Newer models and better tools help, but they do not close the gap. With the same model, Claude Opus 4.5, the share of papers whose results an agent could reproduce rose from 42% to 78% when only the agent harness changed [13]. On full computational-biology studies, the best of 13 frontier models scored 0.48 out of 1 in August 2026 [14]. Asked for references without web search, GPT-5 invented 51% of them [15]. And an auditor that read only an AI-written paper caught 55% of planted flaws; with the run logs and code, it caught 82% [16].

The common thread is a missing record. A result that cannot be traced to the data, code and steps behind it cannot be checked, and a result that cannot be checked cannot safely be built on. As AI raises the number of results, the gap grows.

What we are doing

BioVantageLab builds the missing record into the work itself. Each finding should keep its source: the session, the file and the tool call that produced it. Findings should connect to what a project already knows, so new work builds on checked results instead of starting from zero.

When results disagree, the disagreement should stay visible. We do not claim to decide what is true; we make it possible to check. And you stay in control: your work stays on your computer, and data leaves it only when you pick a cloud model.

Sereh, the research workspace we are building, is our first tool to do this. See how Sereh works →

Request access

References

  1. Errington et al., "Challenges for assessing replicability in preclinical cancer biology", eLife 10:e67995 (2021). elifesciences.org/articles/67995 Preclinical cancer biology only: 193 experiments from 53 high-impact papers.
  2. Errington et al., "Investigating the replicability of preclinical cancer biology", eLife 10:e71601 (2021). elifesciences.org/articles/71601 The 85% is a median effect size. It is a different measure from the share of effects that replicated (46%).
  3. Prinz et al., Nat Rev Drug Discov 10:712 (2011), doi:10.1038/nrd3439-c1 Based on a questionnaire to 23 lab heads; 47 of the 67 projects were in oncology.
  4. Begley & Ellis, Nature 483:531 (2012), doi:10.1038/483531a The papers were hand-picked and Amgen did not publish its data. It does not show that most cancer research is false.
  5. Trisovic et al., Sci Data 9:60 (2022), doi:10.1038/s41597-022-01143-6 The unit is R files, not studies.
  6. Pimentel et al., MSR 2019, doi:10.1109/MSR.2019.00077 These are general GitHub notebooks, not only research notebooks.
  7. Abeysooriya et al., PLoS Comput Biol 17(7):e1008984 (2021), doi:10.1371/journal.pcbi.1008984 Only articles with Excel gene lists (3,436 of 11,117).
  8. Freedman et al., PLoS Biol 13(6):e1002165 (2015), doi:10.1371/journal.pbio.1002165 A model-based estimate of money spent on research that is not reproducible, not a measured amount wasted.
  9. Walters & Wilder, Sci Rep 13:14045 (2023). pmc.ncbi.nlm.nih.gov/articles/PMC10484980/ Older models (GPT-3.5 and GPT-4, 2023).
  10. Peters & Chin-Yee, R. Soc. Open Sci. 12(4):241776 (2025). pmc.ncbi.nlm.nih.gov/articles/PMC12042776/ 26–73% is a range across models, not one rate.
  11. Mitchener et al., BixBench, arXiv:2503.00096. arxiv.org/abs/2503.00096 Older models. 17% is open-answer accuracy, not a failure rate across all tasks.
  12. Wiley ExplanAItions press release, 7 Oct 2025 (2,430 researchers). newsroom.wiley.com Press release, not a paper.
  13. Kapoor et al., Holistic Agent Leaderboard, CORE-Bench Hard results, Dec 2025. hal.cs.princeton.edu/corebench_hard; paper: arXiv:2510.11977 (ICLR 2026). 45 tasks; 78% is the automatic score (95.5% after manual regrading). Which harness wins depends on the model.
  14. Koch et al., "BixBench3", arXiv:2608.25286 (2026). arxiv.org/abs/2608.25286 20 analyses, one agent harness; authors work for the benchmark's publisher. The live leaderboard has since moved slightly higher.
  15. "GhostCite", arXiv:2602.06718v2 (2026). arxiv.org/abs/2602.06718 Preprint; models asked for references without web retrieval; computer-science topics.
  16. Luo, Kasirzadeh, Shah, arXiv:2509.08713, NeurIPS 2025 AI4Science workshop. arxiv.org/abs/2509.08713 The auditor was itself an LLM; 100 projects from two automated research systems.

Each number was checked against its original source on 8 October 2026. The short note after each reference says what the number does not show. If you find an error, tell us and we will correct it.