News

Predicting How New Drugs and Proteins Interact

An LAU-led AI model combines accuracy and interpretability to predict interactions involving unfamiliar drugs and proteins.

By Sergio Thoumi

Finding a promising new drug starts with understanding how a chemical compound interacts with a protein linked to a disease. Researchers can test these interactions in the lab, but with so many possible combinations, investigating them one by one is slow and costly.

Furthermore, computer models that narrow the field tend to perform best on compounds and proteins similar to those they have encountered before.  A more useful model should be evaluated on compounds and proteins that were not included in its training data, while also indicating which molecular features influenced its prediction.

That gap between benchmark performance and real-world usefulness motivated Dr. Lina Abou-Abbas, assistant professor at the School of Engineering, and her team to look more closely at how models perform when confronted with compounds and proteins held out from the model’s training data.

“Although many machine-learning models report strong performance, their evaluation is often based on random data splits,” said Dr. Abou-Abbas. Under such splits, chemically or biologically similar compounds and proteins can appear in both the training and testing data, potentially making the prediction task easier than it would be in practice. The researchers therefore wanted to determine whether performance reported under conventional testing conditions accurately reflects a model’s ability to handle new biological data.

In “Protein and ligand novelty in drug–target interaction prediction: a dual-encoder fusion strategy for more interpretable and generalizable modeling,” published in BMC Bioinformatics, Dr. Abou-Abbas and her team built a system with two views of each compound. One reads its chemical notation as a sequence, much as a language model reads a sentence, while the other maps atoms and bonds as a network. Each approach, or branch, predicts in its own way whether the compound will interact with a protein, and the system combines the two results.

The model learned from 1,988,402 drug–protein interaction records and was tested on familiar pairs, new compounds, new proteins, and pairs in which neither partner had appeared during training. The researchers also examined 200 compounds to identify the molecular features that influenced the model’s prediction.

When tested on previously withheld data from BindingDB Dataset, the model performed well even when it encountered unfamiliar compounds and proteins. Its area under the curve, a measure of how well a model separates interacting from non-interacting pairs, was 0.91 for unseen compounds, and 0.85 for unseen proteins and when both the compound and protein were unseen during training. Its F1 score, which reflects the balance between precision and recall, reached up to 0.87.

These results indicate good discrimination and a balanced precision-recall performance under the study’s evaluation conditions. The two branches showed complementary strengths across the different novelty scenarios, suggesting that combining them could produce more reliable predictions across a wider range of cases.

 A stricter analysis offered another encouraging result. Even among compounds built around chemical frameworks that had not appeared in training, the model’s performance stayed around the same. On the external datasets, however, the area under the curve measure fell to roughly 0.60–0.64.  Dr. Abou-Abbas said the decline was expected because independent datasets can differ from the training data in the drugs and proteins they contain, as well as in experimental protocols, data quality, and overall data distribution.

Rather than viewing that decline as a weakness of the model, Dr. Abou-Abbas sees it as an important evaluation lesson. “It shows that evaluating models only on conventional benchmark splits may not fully reflect performance under cross-dataset distribution shift,” she said. The finding reinforces the need to test predictive systems under conditions that more closely resemble their eventual use, particularly when the goal is to identify interactions involving molecules or proteins not previously included in the model’s training data.

The interpretability analysis, conducted on 200 compounds, showed that the model’s predictions were influenced more by polar molecular groups, such as amide nitrogens and carbonyl oxygens, than by aromatic atoms.  These results describe which features influenced the model’s predictions.

Looking beyond the findings reported in the study, Dr. Abou-Abbas, along with researchers from the Department of Electrical and Computer Engineering at LAU’s School of Engineering, is developing a curated drug–target interaction dataset and benchmarking infrastructure with the goal of supporting more standardized and realistic evaluation conditions. These will include tests involving unseen drugs, unseen proteins and cases in which both are new. The aim, she said, is to enable fairer comparisons between models and help researchers develop methods that generalize more reliably to new biological data.

The study offers a practical advance without presenting computation as a substitute for laboratory evidence. By testing the model on compounds and proteins held out from training as well as on external datasets, and by identifying where its predictions become less reliable, the study provides a more informative picture of how performance changes under novelty and cross-dataset distribution shift.

To browse more scholarly output by the LAU community, visit our open-access digital archive, the Lebanese American University Repository (LAUR).