Catalysis is essential for making processes in the chemical industry, one of the major consumers of energy, more efficient. In recent years, experimental approaches for finding new catalysts were increasingly complemented by theoretical simulation. However, the most widespread computational methods, like density functional theory (DFT), are too computationally costly to apply them to large systems and long timescales, or to screen large numbers of potential catalyst candidates.
In the past decade, machine learning (ML) techniques have reduced the cost of simulations substantially while mostly preserving the accuracy of the parent methods. However, these methods need large amounts of data, which requires the use of black-box parent methods that are mostly of so-called single-reference type (e.g. DFT or coupled cluster theory). Many first-row transition metal complexes, which are promising candidates for cheap homogeneous catalysts, have a complicated electronic structure that requires the use of more sophisticated multireference (MR) methods. These are not easily set up in a black-box fashion for a large number of compounds, which is why they are not yet widespread for labeling machine learning data.
The main objective of ML4Catalysis was the automation of MR calculations in order to train so-called Δ-ML potentials that are more accurate for transition metal complexes. These potentials could then be used for the screening and design of new homogeneous catalysts.