The model learns which experiment is worth running next

GOLLuM joins a language encoder to a Gaussian-process surrogate and trains the representation through the probabilistic objective used for Bayesian optimization. Experimental descriptions with similar outcomes move closer together, producing an uncertainty-aware search surface rather than relying on static text similarity.

The framework uses natural-language descriptions across organic synthesis, process chemistry, materials, catalysis, and molecular design without task-specific feature engineering for each domain.

Sample efficiency is the breakthrough and the deployment constraint

Across 23 tasks beginning with ten below-median experiments, the main variant ranked first on average and matched the final performance of traditional Bayesian optimization with a median 41 percent fewer iterations. In one reaction benchmark it found high-performing conditions at a much higher rate than fixed LLM embeddings and specialized baselines.

These are benchmarked experimental-design tasks, not proof of safe autonomous discovery across laboratories. Materials, equipment, hazardous conditions, dataset quality, and unmeasured objectives still require domain experts, hard constraints, replication, and a record of why each physical experiment was selected.

Primary trail

Go to the source

Read the evidence behind this analysis. External links open in a new tab.

Nature Machine Intelligence — Uncertainty-calibrated optimization for experimental discovery