Optimizing ProtBERT Deep Learning Models for Protein Allergen Prediction
The increasing prevalence of food and environmental allergies has intensified the demand for reliable computational approaches for protein allergen prediction. Recent advances in protein language models have enabled end-to-end sequence representation learning without reliance on handcrafted features. In this study, ProtBERT was optimized and calibrated for protein allergen prediction using curated datasets of allergenic and non-allergenic proteins. ProtBERT embeddings were coupled with a lightweight neural classification head and optimized through learning-rate tuning, dropout regularization, partial fine-tuning of upper transformer layers, and class-balanced loss functions. Model calibration and decision-threshold optimization were further applied to enhance probabilistic reliability. Following optimization, the ProtBERT-based model achieved an accuracy of 0.9392, F1-score of 0.9339, and ROC-AUC of 0.9757, demonstrating strong discriminative capability directly from raw amino acid sequences. For contextual evaluation, ProtBERT performance was compared with physicochemical feature-based machine learning approaches, which served as reference baselines. Despite differences in modeling paradigms, ProtBERT offers advantages in scalability, reduced feature engineering, and compatibility with foundation-model frameworks. These findings indicate that optimized and calibrated ProtBERT models provide a robust sequence-based framework for protein allergen prediction and represent a promising foundation for future hybrid and large-scale computational allergology systems.
Authors:
Muhammad Rezki Rasyak, Mahmud Isnan, Bens Pardamean
International Conference on Smart Computing, IoT, and Machine Learning (SIML 2026)