Google DeepMind reported 2.2 million crystal structures that its model considered stable. A robotic lab in Berkeley then tried 57 computationally selected targets and made 36 of them in 17 days. The distance between those two numbers is the whole business of AI materials discovery.

Prediction has become cheap enough to flood a laboratory with candidates. Physical proof remains slow, chemistry-specific, and full of awkward details such as powders that clump, volatile precursors, metastable phases, contaminated instruments, and crystals that form only under a narrow heating schedule. AI creates value when it improves decisions across that physical process. A long list of plausible formulas is inventory, not impact.

AI changes the search, not the laws of chemistry

Materials search is combinatorial. Selecting five elements from a practical pool of 60 already yields 5,461,512 unordered sets. Composition ratios multiply that count. Crystal structures, defects, pressure, temperature, processing, and microstructure expand it again. Two samples with the same nominal formula can behave differently because atoms occupy different sites or because grain boundaries dominate the measured property.

Computational materials science narrows this space with density functional theory, or DFT. Instead of solving the full many-electron wavefunction, DFT works with electron density and an approximate exchange-correlation functional. It has produced large databases of calculated formation energies, band gaps, elastic properties, and structures. The Materials Project makes computed information available for well over 100,000 materials and supports screening across chemical families.

DFT has two limitations that matter to product teams. It is still expensive at scale, and its error is structured rather than random. A functional that works well for one class of oxides may miss another class. Common high-throughput calculations represent ordered crystals near zero kelvin. Products operate with defects, finite temperature, interfaces, stress, and impurities. Magnetism, strongly correlated electrons, excited states, and reaction kinetics can require more specialized treatment.

Machine learning can act as a fast surrogate. Graph neural networks encode atoms as nodes and bonds or neighborhoods as edges. Trained on DFT results, they can rank new structures without running a fresh calculation for every candidate. Generative models can propose structures conditioned on target properties. Language models can extract synthesis steps from papers. Active-learning models can pick experiments that reduce uncertainty rather than merely choosing the highest predicted score.

The Asianometry video that prompted this analysis gets the central risk right: a startup cannot win by running more synthetic calculations than a large research lab. Its defensible asset is the measured loop connecting a narrow application, repeatable synthesis, characterization, and decisions.

Stability is one gate among several

GNoME is a good example of both progress and category confusion. The Nature paper says its models found 2.2 million structures stable with respect to the Materials Project, with 381,000 new structures on an updated convex hull. That is an enormous expansion of computational candidates. The paper also names open problems: competing polymorphs, vibrational and configurational effects, and synthesizability.

The convex hull compares calculated energies among compositions. A structure on or near the hull may be thermodynamically stable against decomposition under the modeled conditions. That does not provide a recipe, prove that a reaction reaches it, or show that it has a useful property. Kinetic barriers may block formation. A different phase may nucleate first. Air or moisture may destroy it. The desired property may depend on defects absent from the model.

Teams should separate four gates:

Gate Question Minimum evidence
Computational Is the candidate plausible? Calibrated prediction, domain check, uncertainty
Synthetic Can we make the intended phase? Recipe, yield, structural characterization, repeatability
Functional Does it meet the property target? Independent measurements under use conditions
Industrial Can it become a product? Process window, scale, cost, safety, supply chain, lifetime

Passing one gate does not raise a candidate automatically through the others. A battery cathode may show high computed capacity but cycle poorly. A catalyst may perform on a tiny clean coupon and fail after contaminants enter a production stream. A permanent magnet may have high intrinsic magnetization but need a crystal phase that cannot be made economically in bulk.

This framing improves model evaluation. Precision among the top candidates matters more than average error across a random test set because laboratories only make a small ranked subset. Calibration matters because a scientist needs to know when the model is outside its training distribution. Diversity matters because ten near-duplicate candidates can all fail for the same hidden reason.

Use a held-out time split when possible. Randomly dividing a materials database can place closely related structures in training and test sets, making performance look better than a real search into new chemistry. Matbench Discovery exists to evaluate models in a high-throughput discovery setting, but a company still needs an internal benchmark matched to its instruments, chemistry, and product specifications.

The closed loop is the product

A useful system records each candidate, preparation step, instrument setting, raw measurement, analysis, decision, and failure. After an experiment, the software updates its belief and chooses the next action. This turns failures into training data rather than discarded lab notes.

Berkeley's A-Lab shows what that loop requires. Its published system combined computed phase stability, recipes inferred from literature, robotic powder handling, furnaces, X-ray diffraction, and active learning. It made 36 of 57 targets during 17 days of operation. Seventeen targets failed initially; the researchers traced the failures to slow reaction kinetics, precursor volatility, amorphization, and computational error. Manual regrinding and hotter treatment recovered two more targets.

That last detail is more instructive than the headline success rate. The robot's action space omitted procedures that a materials scientist would routinely try. Autonomy is bounded by available hardware, software representations, and safety rules. A model cannot select intermittent grinding if the workflow has no operation for it.

Build the loop in this order:

  1. Define one measurable product objective, including operating conditions and forbidden inputs.
  2. Standardize sample identifiers, recipes, instrument metadata, and units.
  3. Automate the highest-volume repeatable step, not the most photogenic step.
  4. Capture negative and inconclusive outcomes with the same care as successes.
  5. Add uncertainty-aware experiment selection and explicit stop conditions.
  6. Run periodic blind repeats and human reviews for drift or contamination.

Data plumbing is often the limiting work. Instrument vendors export different formats. A sample can split across preparation and characterization systems. Researchers revise labels after manual inspection. The knowledge graph must preserve provenance rather than overwriting an early interpretation. Without that history, the model trains on conclusions whose experimental basis cannot be checked.

Human review belongs at points where the model proposes a new hazard, a new chemical family, an expensive batch, or an irreversible instrument action. Routine runs can proceed automatically inside an approved envelope. This is a better safety design than asking a scientist to click "approve" on hundreds of ordinary steps until attention disappears.

High-throughput does not mean representative

Automation favors measurements and synthesis methods that fit trays, liquid handlers, deposition tools, and fast instruments. That creates selection bias. Thin films are easier to vary combinatorially than kilogram castings. Optical spectra are easier to collect than ten-year corrosion data. A microgram sample can confirm a crystal phase without showing whether the same phase forms in a thick component.

The experimental proxy must predict the product. For a thermal-interface material, a fast screen might measure composition and conductivity on a coupon. Product qualification also needs contact resistance, pump-out, thermal cycling, manufacturability, and aging. If the early score ignores those properties, the loop gets very good at finding candidates the business cannot use.

Before automating a campaign, write a proxy contract:

  • the property measured in the fast loop;
  • the product outcome it is expected to predict;
  • the conditions where that relationship was established;
  • known confounders such as porosity or sample thickness;
  • the expensive validation performed every fixed number of runs.

Measurement uncertainty must travel with every value. A conductivity of 102 is not meaningfully better than 100 when instrument and sample variation span 8. Replicates help estimate noise, and reference materials catch drift. Acquisition functions should trade predicted improvement against uncertainty and experimental cost. Otherwise the optimizer can chase instrument noise or repeatedly choose exotic precursors that create fragile high scores.

Experimental validation also needs adversarial samples. Give the workflow blanks, known reference materials, near-neighbor phases, and samples prepared by a different operator. These controls expose contamination, label leakage, and analysis software that always finds the requested phase. When a model proposes both the sample and the expected diffraction pattern, an automated analyzer can agree with its own assumptions. Periodic manual refinement or a second measurement method breaks that loop.

Property confirmation should come from an instrument the optimizer did not use for selection whenever the cost permits. If X-ray diffraction determined that the target phase exists, microscopy or spectroscopy can test composition and local structure. If a rapid four-point probe ranked electrical performance, a separately prepared sample and calibrated fixture should confirm it. Independence matters more than adding another model to the same raw signal.

The DOE's program for AI-driven autonomous laboratories explicitly joins robotics, edge analysis, feedback, hypothesis generation, and data curation. That list is a useful procurement warning. Buying a robot arm or a foundation model does not buy a closed-loop laboratory.

The economics favor narrow industrial problems

Commercialization continues after a promising sample. A new material needs a scalable process, qualified suppliers, environmental and safety review, reliability data, customer integration, and often regulatory approval. An incumbent material may be slightly worse on one property but available from three suppliers with decades of field data. The candidate has to beat that installed system by enough to justify switching.

Scale-up should begin before the search ends. For each candidate, record precursor availability, price range, toxicity, geographic concentration, expected energy use, and sensitivity to purity. Add these attributes to ranking rather than filtering them after a year of experiments. A material that depends on a scarce precursor or a one-degree process window deserves a lower score even when its intrinsic property is excellent.

The team also needs a transfer test. Ask a second laboratory or pilot line to make the material from the written recipe without coaching from its inventor. Differences in furnace geometry, atmosphere, milling, or feedstock often expose hidden know-how. The failed transfer is useful evidence: it identifies process variables that the automated lab did not record. A candidate should not enter customer qualification until the recipe survives this test at the intended batch scale.

Qualification economics can be staged. Spend little on broad computational screening, more on repeatable gram-scale proof, and reserve expensive pilot production for candidates with a clear customer specification. Set kill criteria at every gate, including maximum precursor cost, minimum yield, acceptable property variation, and a deadline for independent reproduction. This prevents a scientifically interesting sample from consuming the whole portfolio because the team has already invested in it.

This makes broad "discover any material" positioning weak. A stronger program starts with a costly industrial constraint and a customer who owns the qualification path. Examples include reducing an expensive element in a magnet, extending a coating's life in a known reactor, or lowering cure temperature on an existing production line. The search space becomes smaller, the measurements become specific, and the economic value of improvement is visible.

There are three practical business models:

Model Revenue path Main risk
Proprietary material Sell or license the final material Long qualification and buyer power
Joint development Milestones with an industrial partner Custom work can limit reuse
R&D platform Sell software, automation, and operation Must integrate messy lab systems

The platform route can produce value even when no blockbuster compound appears. An industrial lab pays for fewer dead-end experiments, better provenance, faster formulation, and more consistent transfer between researchers. Its moat comes from workflow and validated proprietary data, not from a generic model trained on public DFT records.

An investment committee should ask for validated candidates per lab-month, cost per confirmed property improvement, time from suggestion to measurement, repeatability, and progress through the four gates. Count a candidate as "discovered" only after the team defines which gate the word represents.

For the next campaign, pick one product specification and run a 90-day closed-loop pilot. Limit the chemistry and equipment, reserve part of the budget for blind repeats, and require every run to produce model-ready provenance. The deliverable is a measured reduction in experiments or calendar time, plus a dataset of failures the next campaign can use.