Sabari Defends Dissertation!

Sabari Presents Research Seminar

Sabari Kumar recently presented a research seminar Emergent-Scale Prediction of Molecular Properties with Chemically-Informed Neural Networks at Colorado State University. Sabari discussed his work in molecular representation and chemically-informed machine learning.

Congrats Sabari!

Abstract:

Machine learning (ML) models have emerged as powerful tools for scientific discovery. Leveraging modern computing hardware allows for the generation of large data sets which can be used to train accurate, generalizable models to predict the physico-chemical behavior of molecular systems. However, the predictive power of such models is bounded by how chemical knowledge is encoded within the model. Unlike general-purpose neural networks, which treat molecular data as abstract numerical arrays, chemically-informed architectures incorporate physical and structural constraints directly into their mathematical form. This talk examines the design principles behind such architectures and traces their application across different intra- and intermolecular scales of organization.

At the level of individual molecules, graph neural networks represent atoms as nodes and bonds as edges, propagating information through the molecular graph in a manner that mirrors molecular topology. We will examine the results of different parameterizations of this information in the context of predicting bond dissociation enthalpies and adiabatic singlet-triplet excitation energies, and present novel analytic tools which connect learned model behaviors with chemical intuition.

We then extend this approach to predict the interaction behaviors of molecular species with their surroundings. We show that learned mixing operators mirroring established chemical intuition can reliably predict nonlinear blending behavior, and introduce a new tool inspired by operator symmetrization to analyze this interaction.

Lastly, we will demonstrate how advances in ML algorithm design can enable the accurate prediction of properties which are not possible to simulate through conventional means. By adapting techniques from computational topology and geometric deep learning, we construct a novel neural network architecture capable of predicting protein solubility from predicted protein structures. We show that the model’s internal representations of protein structure align with those obtained through conventional molecular dynamics simulations.

Taken together, these results reflect a common organizing principle: the most transferable and physically meaningful models are those whose structures recapitulate the interactions that govern chemical behavior.

Group Publishes Work on S0-T1 Fragmentation

We’re thrilled to announce that our recent work “A Fragment Based Approach Towards Curating, Comparing and Developing Machine Learning Models Applied in Photochemistry.” was accepted into Chemical Science!

In this work, we developed a novel fragmentation scheme to aid in the prediction of adiabatic singlet-triplet energy gaps.

20250128 fig2 frag scheme

20250128 fig2 frag scheme

Abstract:

The development of Graph Neural Networks for predicting molecular properties has garnered significant attention, as it enables the correlation of quickly computable atomic and bond descriptors with overall molecular properties. With the raising interest in photochemistry and photocatalysis as sustainable alternatives to thermal reactions, curation of virtual databases of computed photophysical properties for training of machine learning models has become popular. Unfortunately, current efforts fail to consider the exciton localization onto different chromophores of the same molecule, leading to potentially large prediction errors. Here we describe a molecular fragmentation strategy that can be used to overcome this limitation, while also providing a way to compare structural diversity between different libraries. Using a newly generated database of 46,432 adiabatic S0-T1 energy gaps (ALFAST-DB), we compare its diversity with two datasets from the literature and demonstrate that a fragment-based delta learning approach improves model generalizability while achieving accuracies comparable to those of traditional message passing graph neural network architectures (MPGNN).

To read more click here

Chris Defends Dissertation!

Chris Presents Research Seminar

Chris Stubbs recently presented a research seminar “Advancing Solubility Prediction Through Machine Learning” at Colorado State University. Chris discusses work from two recent papers from the group “Predicting homopolymer and copolymer solubility through machine learning” and Enhancing Predictive Models for Solubility in Multicomponent Solvent Systems using Semi-Supervised Graph Neural Networks.

Congrats Chris!

Abstract:

Solubility is a fundamental chemical property with wide-ranging applications including reaction optimization, waste recycling, and manufacturing. As measuring solubility can be time or resource-intensive, predicting solubility through computational methods has received significant attention in recent work. In particular, solubility prediction through machine learning (ML) has been heavily studied due to its speed and accessibility advantages over alternative methods using quantum mechanical or semi-empirical formulations. In this talk, we discuss two recent advances in solubility prediction for polymers and small molecules respectively. We first discuss our recent work to predict polymer solubility in single solvents, which is of interest for applications in plastic recycling and polymer design. We found that simple tree-based models with low-dimensional features can achieve over 80% prediction accuracy on homopolymer and copolymer solubility, and that these predictions can be rationalized using explainable AI methods such as Shapley Additive Explanations (SHAP). Following our discussion of polymer solubility prediction, we next examine ML predictions of small molecule solubility in multiple solvents (multicomponent solubility). In comparison to single solvent solubility, multicomponent solubility has increased complexity but allows for greater control over solute separation and processing, leading to uses in biomass upgrading and recycling. To accelerate these applications, we curated a new multicomponent solubility database (MixSolDB) which we used to train two graph neural network (GNN) models to predict solute solubility in up to three solvents. We find that our novel subgraph architecture for solubility prediction outperforms the more common concatenation architecture, achieving a mean absolute error (MAE) of 0.67 kcal/mol on ΔGsolv prediction. In summary, we demonstrate that ML-based predictions of solubility are chemically accurate while remaining useful for sustainable applications.

Dr. Sumin Song Joined as a New Postdoc

Chris presents at STEM Poster Day