C5R CORP.

SciUniverse

LEVEL 1

Can frontier models carry out scientific work?

SciUniverse is a benchmark of scientific work that spans the entire research process across chemistry, biology, and materials science.

We start with SciUniverse Level 1: tasks that are straightforward for scientists and take no more than a few hours of work. Future releases will test more demanding scientific work.

Models direct work at Facility-0 using our facility harness by controlling machines and giving instructions to human operators. Working with real materials and instruments lets us test whether their plans produce the intended results and how they adapt when experiments fail.

The samples below show recorded footage (left), the equipment involved (center), and each action’s inputs and outputs (right).

RESULTS

Overall benchmark results, averaged across 17 task families. Each task family has equal weight.
ModelPass@1Cost ($)Model/API inference cost only, averaged per task attempt. Lab and labor costs are not reported. Repeats are averaged using the same task and family weights as scores.
Claude Fable 5.1xhigh45.3%$40.61
GPT–6 Astraxhigh32.5%$52.37
Claude Opus 5xhigh30.5%$46.31
Grok 4.6xhigh26.2%$13.41
Gemini 3.8 Flashhigh14.6%$4.55
GPT–5.6 Solxhigh9.4%$16.53
0%50%100%

A task space for scientific work

We think most scientific benchmarks capture too narrow a slice of scientific work, starting after the experiment with clean data ready to analyze. That leaves out much of the work: choosing materials, operating instruments, running experiments, debugging failures, and adapting to real constraints.

To measure model performance across more of that work, we chose tasks covering:

  • Basic sample preparation. We included simple sample-preparation tasks because early tests showed that models struggled with them.
  • Instrument control. We compared how well models controlled instruments using vendor software, Python APIs, and direct firmware commands.
  • Protocol adaptation. We chose standard processes and kits, then tested models’ ability to adapt to limited reagents, equipment, and time.
  • Learning across experiments. We included optimization tasks to see whether models could learn from experiments and choose what to try next with limited materials and budgets.
  • Facility management. We added facility management to see whether models could keep experiments on track through shortages, equipment failures, and deadlines.
  • Interpreting real measurements. We used real NMR, XRD, and chromatography data to test whether models could identify products, quantify mixtures, and recognize unreliable signals.

SciUniverse Level 1 contains 92 tasks across 17 task families. Each task gives a model a specific scientific objective, such as synthesizing a target molecule. Tasks involving the same kind of scientific work form a family. Each family contributes equally to the overall benchmark score.

We map tasks by work horizon and scale of matter to compare chemistry, biology, and materials science in one frame. Select a task on the map to explore its results.

SciUniverse

Real WorldVirtual
Individual operationProtocolCampaignResearch program

Synthesize an amide

Confirm product formation by LCMS when synthesizing N-benzyl-4-methylbenzamide from p-toluic acid and benzylamine.

Reasoning

Loading model transcript…

Model transcript

Video will appear when recorded activity begins.

Loading facility view…

00:0000:01

Model activity

How models fail in the physical world

Models repeatedly miss critical details in laboratory work. We observed them trying to pipette samples that were still frozen, vortexing open well plates, reusing the same pipette tip across DNA-containing wells, and failing to account for evaporating solvents. They struggle with both physical reasoning and protocol design. The table below breaks down each model’s documented failures in physical tasks by category.

Failure modeAstra 6Fable 5.1Gemini 3.8Grok 4.6Opus 5Sol 5.6
Percentages show how each model’s recorded failures in physical tasks are split across categories.

Acknowledgments

We thank the scientists, operators, and engineers who made this benchmark and Facility-0 possible, and the researchers and maintainers who shared the experimental data and software behind some of the simulated tasks in SciUniverse Level 1.

Full credits

We also thank the developers of PyLabRobot, one of our instrument-control interfaces, and the contributors to GSAS-II and the Crystallography Open Database, used to build our XRD reference baselines.

Contact us

Get in touch to evaluate your model on SciUniverse or work with us on new scientific tasks. We welcome collaborations with teams developing AI models, building evaluations, and doing experimental research.

Email contact@c5r.net