Explainable AI attributions in protein language models do not recover allergen epitopes

Jianzhou Yao1,2, Anxiong Song1,2, Katja Baerenfaller1,3 and Damir Zhakparov1,3

  1. Swiss Institute of Allergy and Asthma Research, Davos, Switzerland
  2. ETH Zurich, Zurich, Switzerland
  3. Swiss Institute of Bioinformatics, Lausanne, Switzerland

Background. Assessing whether a new food protein could cause an allergic reaction is an important part of food safety evaluation. The immune system does not recognize all parts of a protein equally. Instead, it responds to specific regions called epitopes. Artificial intelligence models trained on large numbers of protein sequences, known as protein language models, can be used to build models that predict whether a protein is likely to be allergenic with high accuracy. However, it remains unclear which parts of a protein these models use to make their predictions. Explainable AI methods can highlight individual amino acids that are important for a model's decision. One commonly used method, Integrated Gradients, assigns an importance score to each amino acid residue in the protein. Such importance maps are sometimes interpreted as identifying biologically meaningful regions, but it is unclear how well they match experimentally known allergen epitopes.

Methods. We tested whether amino acids highlighted by allergenicity prediction models overlap with experimentally validated immune-recognized regions collected from the Immune Epitope Database. We focused on MHC class II epitopes, protein regions that contribute to activating immune responses involved in allergy. We evaluated several allergenicity classifiers based on ESM-2, a protein language model trained to learn patterns from millions of protein sequences, as well as the previously published DeepPlantAllergy model. We also trained a model to predict both whole-protein allergenicity and epitope locations simultaneously to test whether explicitly teaching the model about epitopes would make its explanations more biologically meaningful. Finally, we removed highly important amino acids or systematically replaced them with other amino acids computationally to investigate what information the models were sensitive to.

Results. Although the models accurately classified proteins as allergenic or non-allergenic, the amino acids highlighted as important for these predictions did not overlap with known epitopes more strongly than expected by chance. When the model was explicitly trained to identify epitopes, it successfully learned their locations, but this information was still not reflected in the explanation of its overall allergenicity predictions. Removing highly highlighted amino acids reduced the models' predicted allergenicity, showing that these residues were genuinely important for the models’ predictions. However, systematically changing these residues suggested that the models were sensitive to properties such as amino-acid chemistry and local sequence composition rather than to specific epitope patterns.

Conclusion. Strong allergenicity classification performance does not necessarily imply that residue-level explanations obtained with explainable AI recover experimentally known allergen epitopes. Our results show that residues highlighted by Integrated Gradients as important for a model’s prediction can differ from immunologically annotated regions. Quantitative comparison with experimental epitope data should therefore be required before attribution maps are used to support food-safety assessments or the design of less allergenic proteins.