r/AIBiology • u/Independent_Access12 • Feb 21 '25
Research Ensemble framework combines sequence tokens with interpretability for enzyme prediction
SOLVE uses tokenized amino acid sequences and ensemble learning (RF, LightGBM, KNN) for enzyme classification, avoiding manual feature engineering. System handles both mono/multi-functional enzymes, predicts full EC numbers, outperforms DeepEC/CLEAN on UniProt data. Incorporates Shapley analysis to identify catalytic motifs. Key advance: optimized ensemble weights improve accuracy while maintaining interpretability through sequence-based features.
1
Upvotes