Using AI tools to enhance virtual screening for covalent drug candidates

Original article by Sandy Field can be found at the APS Website.

High throughput screening of candidate drug structures to inhibit “druggable” targets in disease seems like a perfect application of artificial intelligence (AI) as it involves grinding through thousands of variables with known or desired physicochemical properties and lots of big data from protein structures. But does AI perform better than the modeling tools we already have? A team from The Weizmann Institute of Science in Rehovot, Israel and Iowa State University in Ames, Iowa recently published work that aimed to answer this question.

Small molecule drugs that bind to their target irreversibly in a covalent interaction have been used for more than 100 years but it is only in the past decade or so that screening tools have been developed to identify and optimize them. Covalent drugs have the potential to be more potent, selective, and longer acting than non-covalent drugs but also have the potential to cause problems related to irreversible off-target interactions. When designed properly, covalent inhibitors show great promise and recent success stories include identification of a long-sought drug for oncogenic K-Ras and the development of nirmatrelvir for COVID-19. This work focused on optimization of screening tools to enhance identification of new covalent drugs by testing whether an AI tool could perform better than current computational docking software tools for enrichment of active covalent compounds to the top of the results list.

The binding of a covalent inhibitor involves two steps, a non-covalent molecular recognition interaction followed by covalent bond formation, but current screening tools ignore the possible reactivity of the reactants and the orientation of reaction intermediates. The research team sought to understand whether the Alpha-Fold 3 (AF3) AI all-atom structure prediction modeling tool from Google Deep Mind, which can predict covalent biomolecular complexes, could improve current screening methods to enhance prediction of the covalent binding step. For this study, they focused on targeting a nucleophilic reaction at a cysteine (Cys) amino acid in the Bruton tyrosine kinase (BTK) by an acrylamide-bearing electrophile.

The first step was to design a curated database of structures and inhibitors to test AF3 against current docking tools and optimize computational parameters. The final COValid database included nine enzyme structures, 874 active compounds, and 37,919 commercially available “decoy” compounds to be able to discern successful enrichment. Testing of three current docking tools on the COValid database against AF3 showed that COValid yielded good results and that the AI tool performed better than current tools. After more testing to understand strengths and weaknesses of the system and to check for bias in the way the AI tool was trained, it was time for the real test, identifying covalent inhibitors of BTK.

AF3 was tasked with enriching covalent binders to Cys481 of BTK from a database of 906,000 acrylamide compounds. AF3 identified 390 possibilities and the team narrowed this down to 13 that they synthesized and tested for activity. Liquid chromatography/mass spectrometry (LC/MS) testing showed that three of the thirteen compounds reached 100% covalent binding to BTK within two hours and these three novel compounds, YS1, YS2, and YS3, underwent more testing. All three compounds stabilized BTK and inhibited its kinase activity, with YS1 being the most potent inhibitor in vitro and in cells. In testing against potential off-target kinases, YS1 was shown to have excellent selectivity for BTK.

Next, the research team wanted to evaluate the binding interactions predicted by AF3. For this, they collected X-ray crystallography data at the Northeastern Collaborative Access Team (NE-CAT) beamline at the Advanced Photon Source, a U.S. Department of Energy (DOE) Office of Science user facility at DOE’s Argonne National Laboratory, for the three candidate compounds in complex with the kinase domain from BTK. YS1 and YS2 resolved to 1.6 Å and 1.27 Å, respectively, while YS3 resolved to 3.5 Å. The structural data for YS1 and YS2 showed that the binding of these inhibitors to BTK was very similar to that predicted by AF3 with YS2 forming a complex in the so-called “back-pocket” of the active site and YS1 forming a complex that was neither back-pocket nor front-pocket (Figure 1).

Interestingly, this finding shows that AF3 provided sub-angstrom details for an inhibitor that extends our current understanding of inhibitor binding to the active site of this well-known kinase. The structure of YS3 was not as well resolved but was sufficient to determine that the AF3-predicted interactions differed from the X-ray diffraction data by a single rotation around a bond.

The next steps for the team will be to validate and extend these results with other protein families beyond kinases and GTPases. However, given that the team was able to identify novel inhibitors for a well-known kinase, the importance of kinases in human disease means that the impact of this work promises to be significant to both medicine and the field of virtual drug screening.

References:

Y. Shamir1, R. Gabizon1, A. Rogel1, D.Y. Lin2, A.H. Andreotti2, N. London1, “Discovery of covalent ligands with AlphaFold3,” Journal American Chemical Society (2026) 148 (12): 13043–13054. DOI: https://doi.org/10.1021/jacs.5c22222

Author affiliations: 1Weizmann Institute of Science; 2Iowa State University.

We thank Sarel Fleishman for critical reading of the manuscript, the Irwin lab in UCSF for access to their cluster for DOCK6 and DOCKovalent ligand generation and in particular Dr. Khanh Tang for technical assistance. We thank Dr. Trent Balius for assistance with DOCK6.12, as well as sharing DUDE-Z enrichment calculation data, and Dr. Alexey Orlov for sharing his ligand generation pipeline. We acknowledge Crelux, a WuXi AppTec company, for performing the kinetic analysis of the covalent hits. Y.S. is funded by the CHE fellowship for data sciences. This research was generously supported by the Knell Family Institute of Artificial Intelligence. Research in the London lab is funded by the Abisch-Frenkel foundation, European Research Council (ERC_CoG 101125683), the Israel Science Foundation (1869/24), the Honey and Dr. Barry Sherman Lab, the Dr. Barry Sherman Institute for Medicinal Chemistry, the Abisch-Frenkel RNA Therapeutics Center, the Moross Integrated Cancer Center, the Goldhirsh-Yellin Foundation and Celia Zwillenberg-Fridman. D.Y.L. and A.H.A. thank the Roy J. Carver Charitable Trust for financial support. This work also utilized the resources at the NE-CAT beamlines (GM124165), a Pilatus detector (RR029205), and an Eiger detector (OD021527), all of which are funded by the NIH.

Share this post