BioVantageLab · Science you can trace
BioVantage Lab Request access

Selected by Anthropic · Startup Program 2026

Science you
can trace.

BioVantageLab exists to make life-science research traceable, connected, and worth building on, now that AI agents produce it faster than anyone can check.

The problem

Much of life-science research cannot be checked. The gap is growing.

Science moves forward when results can be checked and built on. In the life sciences, that check already failed often before AI [1][3]. AI agents now run analyses faster than anyone can check them. Newer models still make scientific mistakes [14][15], and the setup around a model, its harness, changes how much of the work holds up [13].

Read the full article →

As AI speeds research up, errors get harder to see

Papers an agent reproduced, same model

42% → 78%Claude Opus 4.5

Only the agent harness changed: one general setup, one built for the task. 45 tasks, December 2025.

HAL, CORE-Bench Hard, 2025

Best score on full computational-biology studies

0.48out of 1

Best of 13 frontier models on 20 analyses, August 2026. Agents struggle with data scale, long analysis chains and error recovery.

Koch et al., BixBench3, 2026

References GPT-5 invented without web search

51%GPT-5

Claude Sonnet 4: 22%. Gemini 2.5 Pro: 59%. 2026 test of generated citations.

GhostCite, arXiv preprint, 2026

Planted research flaws an auditor caught

55% → 82%paper → logs and code

Reading the run logs and code, not only the AI-written paper, found far more of the flaws.

Luo et al., NeurIPS 2025 AI4Science

AI makes research faster. The harness around a model decides how much of that work you can check. Sereh is the harness we are building: every finding keeps its session, file and tool call.

Why we exist

We want a biology lab where any result can be checked against its record, by anyone, in minutes, however fast an agent produced it.

Science has entered the age of AI. An analysis that once took weeks can now run overnight and reach you by morning. That speed could change how fast biology moves.

But a result is only useful if someone can check it. When nobody can say which run, file or model produced a number, speed makes research harder to trust, not easier. We started BioVantageLab to close that gap.

How we solve it

Four commitments that build trust.

Provenance

Keep the record with the result

Each finding stays linked to the session, the file and the tool call that produced it.

Connection

Connect what a project knows

New work builds on checked findings, so each task does not start from zero.

Transparency

Show disagreement

When findings conflict, we flag it. We do not claim to decide what is true.

Control

Keep you in control

Your work stays on your computer. Data leaves it only when you pick a cloud model, and you can see which one.

A commitment, in practice

An agent recommended N49W. The lab found no gain.

Open the full example →
Overnight summary · subagent 3 5 of 5 checked · 06:02

The best streptavidin variant, N49W, scored ΔΔG −1.8 kcal/mol1, about 20× tighter2 binding, with a pose RMSD of 0.9 Å3. A second model rated it 9 / 104. The gain is significant, p < 0.015.

  1. 1−1.8 kcal/molLinkedFound in run_07/ddg.log, from the docking session at 02:14.
  2. 220×Recalculatedexp(1.8 / 0.593) ≈ 21× at 298 K. The arithmetic holds.
  3. 30.9 ÅFlaggedMeasured against the agent's own model, not the crystal structure 1STP.
  4. 49 / 10Does not countThe second model saw the same data. It is not an independent check.
  5. 5p < 0.01No sourceNo statistical test appears in any run or file.

Two of five numbers hold up. The recommendation rested on the other three.

Illustrative example · structure PDB 1STP · values invented

Who we build for

For life-science teams who need to trust their results.

Computational biologists and chemists

Know where each number came from, at agent speed.

Use AI for analysis at full speed, and keep the record you need to defend the result.

Research leads

See which results hold up before you sign off.

Check a claim against its source in minutes, not by asking who ran what.

Bio companies

Keep your data in-house. Show the record behind a claim.

Give partners and investors results with their sources attached.

The tool we are building

Sereh

Sereh is a research workspace that runs on your own computer. It keeps the record behind each finding and shows you where findings disagree. You work with AI agents and choose the model for each step.

Open the full worked example →

  • Runs on your computer local models by default · cloud only when you choose
  • Each record names its source session · file · tool call
  • Findings connect into chapters with links and conflicts
  • Flags disagreement does not judge truth
Sereh, sample project: choose the model. Pick the model for each message, right in the composer. Local models run on your own computer. Claude and OpenAI models are there when you choose them.
Pick the model for each message, right in the composer.

The real app with a sample project · early build, still named Serra inside the app

Try Sereh on your own research.

Mohamed Elrefaiy, founder

I am a chemistry PhD student at UT Austin. I build Sereh so you can stand behind a result, however fast an agent produced it.

Private early access. We will tell you what Sereh records before you install it.