The model learns which experiment is worth running next
GOLLuM joins a language encoder to a Gaussian-process surrogate and trains the representation through the probabilistic objective used for Bayesian optimization. Experimental descriptions with similar outcomes move closer together, producing an uncertainty-aware search surface rather than relying on static text similarity.
The framework uses natural-language descriptions across organic synthesis, process chemistry, materials, catalysis, and molecular design without task-specific feature engineering for each domain.
Sample efficiency is the breakthrough and the deployment constraint
Across 23 tasks beginning with ten below-median experiments, the main variant ranked first on average and matched the final performance of traditional Bayesian optimization with a median 41 percent fewer iterations. In one reaction benchmark it found high-performing conditions at a much higher rate than fixed LLM embeddings and specialized baselines.
These are benchmarked experimental-design tasks, not proof of safe autonomous discovery across laboratories. Materials, equipment, hazardous conditions, dataset quality, and unmeasured objectives still require domain experts, hard constraints, replication, and a record of why each physical experiment was selected.
Go to the source
Read the evidence behind this analysis. External links open in a new tab.
Nature Machine Intelligence — Uncertainty-calibrated optimization for experimental discovery


