PeptiVerse: A unified platform for therapeutic peptide property prediction - Nature
High confidence: full text extraction produced 62124 characters.
Abstract
Therapeutic peptides combine the advantages of small molecules and antibodies, offering target flexibility and low immunogenicity, yet their successful translation requires careful evaluation of multiple developability properties beyond binding alone. As chemically modified peptides become increasingly common in drug design, no unified platform currently supports systematic property assessment across both canonical sequences and SMILES-based representations. Leveraging the generalizability of large foundational models trained on protein and chemical data, we introduce PeptiVerse, a universal therapeutic peptide property prediction platform. PeptiVerse accepts either amino acid sequences or chemically modified peptide SMILES, delivers state-of-the-art performance across diverse property prediction tasks, and provides both a web interface and open-source implementation for rapid, accessible, and scalable peptide developability analysis. By unifying property prediction across representations, PeptiVerse directly supports early-stage peptide therapeutic development campaigns and property-aware generative design workflows.
Similar content being viewed by others
Introduction
Peptide-based therapeutics have gained significant attention in recent years, highlighted by the clinical success of GLP-1 receptor agonists for metabolic diseases1,2. As a therapeutic modality, peptides occupy a unique position between small molecules and antibodies, combining larger interaction surfaces capable of engaging protein-protein interfaces traditionally considered undruggable3,4 with reduced immunogenicity and manufacturing complexity relative to full-length antibodies4,5,6. These features make peptides candidates for a broad spectrum of therapeutic targets, including receptors, enzymes, and intrinsically disordered proteins.
Despite these advantages, native peptides often exhibit suboptimal translational profiles. For example, poor membrane permeability limits cellular uptake and oral bioavailability7,8, rapid proteolytic degradation results in short circulating half-life and frequent dosing4,9, and low solubility promotes aggregation and formulation failure10,11. In addition, certain features of amphipathic sequences can cause hemolysis through nonspecific membrane disruption12,13, while nonspecific protein adsorption (fouling) reduces effective concentration and increases off-target interactions in complex biological environments14,15. These limitations can be partially mitigated through cyclization, terminal modifications, D-amino acids, or other noncanonical residues4,16,17, but such interventions push peptides beyond the assumptions of traditional protein sequence-based predictors18,19. As a result, practical peptide design requires systematic evaluation of multiple experimentally grounded properties beyond binding affinity alone.
Existing computational tools inadequately address this need. Sequence-based predictors such as PeptideBERT and those in Peptipedia v2.0 are confined to natural amino acids and cannot accommodate chemical modifications15,20. Small-molecule ADMET platforms accept “chemical" inputs in the form of SMILES but are trained on drug-like chemical space that differs substantially from peptides and proteins21,22. Peptide-oriented SMILES predictors, including PepLand, PepDoRA, and PeptideDashboard, represent important steps toward chemistry-aware peptide modeling23,24,25, yet cover only a limited subset of relevant properties. Fitness-focused platforms such as ProteinGym benchmark mutational effects but do not target critical biochemical traits relevant for therapeutic peptide design18,19. Together, these gaps highlight the lack of a comprehensive, modality-flexible predictor capable of evaluating the full chemical and sequence diversity of modern peptide therapeutics.
The recent emergence of large protein and chemical language models enables unified property predictors that leverage rich learned representations without requiring explicit structural information26,27. Such predictors have become integral to modern generative peptide design workflows, where classifiers guide sampling, rank candidates, or provide post hoc filtering28,29,30,31,32,33,34,35,36. In this setting, gradient coupling between generator and predictor is not always necessary, allowing effective use of classical methods such as SVM37, Elastic Nets38, and XGBoost39, which perform well on pretrained embeddings while reducing the risk of overfitting40.
To realize this potential in a unified framework, we introduce PeptiVerse (Fig. 1), a universal therapeutic peptide property evaluation platform developed to standardize and accelerate computational peptide design. PeptiVerse integrates state-of-the-art foundation models with carefully curated datasets to deliver fast, accurate, and scalable property predictions, supporting both sequence-based and SMILES-based peptide inputs. Beyond post hoc evaluation, these predictors can be directly used as guidance oracles within generative modeling workflows29,30,32,33,34, enabling generation, ranking, and optimization of peptide candidates across diverse targets. Together, PeptiVerse provides a unified foundation for property-aware peptide discovery, enabling both early-stage candidate prioritization and integration with generative design workflows to accelerate therapeutic translation.
Results
Dataset composition highlights constraints across peptide properties
Accurate peptide property prediction is fundamentally constrained by data availability, coverage, and experimental diversity. As an initial step, we curated and integrated experimentally derived datasets spanning multiple peptide properties from a wide range of public sources15,16,23,41,42,43,44,45,46. These datasets collectively cover both canonical amino acid sequences and chemically modified peptides represented as SMILES, enabling unified analysis across representation modalities. Examination of dataset composition reveals substantial heterogeneity in dataset size, label balance, and value distributions across properties (Fig. 2 and Table 2).
Our data show that several classification tasks are supported by large, well-populated datasets with broad coverage of peptide sequence and chemical space. Hemolysis and non-fouling datasets comprise thousands to tens of thousands of peptides curated from antimicrobial and surface-interaction studies41,47, while solubility datasets aggregate protein expression outcomes from structural genomics pipelines and mutational databases43,48. Permeability datasets include both canonical cell-penetrating peptides and noncanonical cyclic peptides measured using PAMPA and Caco-2 assays16,49,50, yielding relatively balanced class distributions and continuous-valued measurements spanning multiple orders of magnitude. These properties exhibit broad value distributions and sufficient sample sizes to support robust model training and evaluation under similarity-aware splits.
In contrast, regression tasks such as peptide half-life and binding affinity remain comparatively data-limited. Half-life measurements, curated from THPdb2, PEPlife, and PepTherDia44,45,46, are sparse, heterogeneous in experimental protocol, and often reported in coarse or qualitative units, resulting in limited sample sizes for both sequence- and SMILES-based representations. Binding affinity datasets aggregate diverse experimental readouts (Kd, Ki, and IC50) across protein-peptide pairs23,51, but remain modest in scale relative to classification tasks. Together, our observations highlight that achievable predictive performance across peptide properties is frequently constrained by data availability and experimental variability rather than model capacity alone, motivating property-specific modeling strategies and emphasizing the importance of continued dataset expansion.
PeptiVerse deploys state-of-the-art predictors for a broad range of therapeutically relevant peptide properties
Given our heterogeneous data settings, we evaluated a diverse set of predictor architectures within PeptiVerse, including linear models, boosting methods, multilayer perceptrons (MLPs), convolutional neural networks (CNNs), support vector machines (SVMs), and transformer-based models (Fig. 3). Across all classification tasks and both amino acid sequence and SMILES inputs, overall performance differences between architectures were modest when trained on fixed, information-rich embeddings derived from ESM-2, PeptideCLM or ChemBERTa (Supplementary Fig. S1 and Supplementary Tables S1, S2, S3). This consistency suggests that representation quality and dataset characteristics, rather than downstream model capacity, constitute the primary performance bottleneck in peptide property prediction.
Rather than enforcing a single architecture, PeptiVerse identifies the best-performing model for each property and representation, reflecting the observation that no single architecture emerges as universally optimal. Different peptide properties favored different inductive biases, with hemolysis and permeability benefiting from margin-based or kernel methods, and chemically rich SMILES-based favoring convolutional architectures (Fig. 3). Importantly, similar trends were observed for regression tasks, where multiple nonlinear models achieved comparable performance, and no clear architectural dominance was observed (Supplementary Figs. S2, S3). In both settings, properties supported by larger datasets with broader coverage of the underlying value distributions, such as the cell-penetrating peptide permeability task (Permeability CPP) task, consistently yielded stronger performance than properties with limited or sparsely distributed data, such as peptide half-life. In addition, all evaluations were conducted using similarity-based data splits, indicating that simple models can generalize comparably to deep architectures on out-of-distribution peptide sequences when initialized with high-quality representations.
Across all tasks, embedding choice was the dominant factor governing predictive accuracy rather than model architecture. The spread in performance across embeddings consistently exceeded the spread across models within the same embedding. For regression, ChemBERTa consistently outperformed PeptideCLM on both PAMPA (ρ = 0.69 vs. 0.59) and Caco-2 (ρ = 0.80 vs. 0.75; Supplementary Table S2), likely reflecting ChemBERTa’s broader pretraining chemical diversity compared to PeptideCLM’s focus on synthetic cyclic peptides. DNN models have low variance with different random seeds, confirming that performance differences across configurations reflect genuine embedding and task effects rather than optimization noise. Although we find that baseline embeddings (e.g., ECFP, VHSE, or one-hot) can achieve competitive performance at times, embeddings from language models consistently attain the highest performance across properties (Supplementary Tables S4, S5, S6, S7, S8, S9). We find that ESM-C embeddings only improve upon ESM-2 embeddings for solubility (F1 = 0.765 vs. 0.754) and half-life (ρ = 0.636 vs. 0.582; Supplementary Table S10), highlighting ESM-2’s competitive representation ability.
PeptiVerse provides accurate binding affinity predictions
To assess whether PeptiVerse binding affinity predictions reflect physically meaningful interaction strength, we examined their relationship to structure-based confidence metrics derived from state-of-the-art de novo protein-peptide complex predictors. Recent studies have reported correlations between structure prediction confidence scores, such as ipTM, and docking-based interactions metrics for protein-protein interactions, motivating the use of such scores as proxies for binding strength in ab initio modeling pipelines52,53. We therefore asked whether similar relationships hold for peptide-protein interactions, including both wild-type and chemically modified peptides.
Using OpenFold3-predicted complexes54, we compared ipTM scores against experimentally measured binding affinities for 1433 amino acid sequence inputs and 1702 SMILES-based peptide inputs. In contrast to prior observations demonstrating the efficacy of the ipTM score for estimating protein-protein interaction strength53, ipTM showed negligible association with experimental binding affinity across either peptide representation (Supplementary Fig. S4; ∣ρ∣ ≈ 0.1), indicating that structure confidence metrics are insufficient proxies for peptide-protein binding strength. These results suggest that ipTM does not reliably capture the energetic or kinetic determinants governing peptide binding, likely due to the increased flexibility, shallow binding interfaces, and diverse chemistries characteristic of peptide ligands. Moreover, even when accurate complex structures are available, structure prediction is computationally heavier than embedding-based inference and is not well-suited for high-throughput screening across large peptide libraries. PeptiVerse therefore provides a fast affinity surrogate that complements structure prediction rather than replacing it.
PeptiVerse binding affinity predictions show statistically significant agreement with experimental measurements across both representation modalities, achieving Spearman ρ = 0.57 for amino acid sequence inputs and ρ = 0.61 for SMILES-based inputs (all p < 10−3; Supplementary Fig. S3). These results, obtained using a transformer-based cross-attention architecture operating directly on protein and peptide embeddings, demonstrate that PeptiVerse provides experimentally relevant binding affinity estimates that complement structural modeling and enable binding-aware filtering and generative peptide design.
PeptiVerse demonstrates superior predictive performance relative to existing peptide property predictors
PeptiVerse is, to our knowledge, the first unified framework capable of predicting multiple physicochemical and developability properties for both amino acid sequence and SMILES-encoded peptide inputs. In contrast, prior tools have made meaningful progress but remain specialized in either input modality or property scope23,27. For example, PeptideBERT15 operates primarily on canonical amino acid sequences and does not generalize to chemically modified peptides, whereas PepLand23 handles SMILES-based representations of both canonical and non-canonical peptides but covers only a limited set of properties. None of the existing tools provides the breadth of properties, multimodality, or user accessibility offered by PeptiVerse.
To quantitatively assess these differences, we compare PeptiVerse against representative prior methods across multiple regression and classification benchmarks in Table 1. Across most evaluated tasks, PeptiVerse achieves competitive or superior performance under a unified evaluation protocol. We note that PepLand reports higher performance on SMILES-based binding affinity prediction23. However, these results were obtained using random data splits, which permit substantial overlap in sequence or chemical similarity between the training and test sets. In contrast, for binding affinity PeptiVerse employs score-distribution-based splits, which ensure both splits span the full affinity range, as targets and binders recur across many peptide-protein pairs and cannot be cleanly separated by similarity. Performance differences on SMILES-based binding affinity thus reflect differences in the evaluation protocol rather than representational limitations.
Additionally, the PepLand + ESM-2 setting is reported only for canonical peptide tasks, and the mechanism for switching between SMILES and sequence inputs is not explicitly specified. By contrast, PeptiVerse provides an explicit and unified multimodal design, enabling consistent evaluation across both canonical and non-canonical peptide spaces.
PeptiVerse is deployed as a unified web interface for peptide evaluation
Finally, we developed an interactive web interface to make PeptiVerse accessible for practical use by both experimental and computational researchers. PeptiVerse is deployed as a web server (Fig. 4) built with Gradio55 and hosted on HuggingFace Spaces56. The interface allows users to submit peptide inputs as either amino acid sequences or SMILES strings, enabling property evaluation across both canonical and chemically modified peptides. It supports a broad set of therapeutically relevant properties, including binding affinity to a target protein sequence, hemolytic activity, non-fouling behavior, permeability, solubility, toxicity, and peptide half-life.
To improve interpretability and usability, the interface also provides visualizations of the underlying training data distributions, allowing users to contextualize input peptides relative to experimentally characterized datasets. All datasets used to train the deployed models are standardized and distributed in HuggingFace Dataset format, facilitating reproducibility, benchmarking, and integration into downstream peptide design and optimization workflows.
Discussion
PeptiVerse introduces a unified framework for therapeutic peptide property prediction that supports both amino acid sequence and SMILES-encoded inputs, including peptides containing non-canonical amino acids. By building on pretrained protein (ESM-2) and chemical (PeptideCLM) language models26,27, PeptiVerse focuses on training lightweight predictor heads rather than full representation models, enabling efficient, scalable, and easily deployable property evaluation. These results highlight the heterogeneity of peptide property landscapes and motivate a flexible, property-aware modeling strategy rather than reliance on a single predictor.
Compared with prior SMILES-based peptide models such as PepLand23, which rely on complex graph constructions and full-model retraining, PeptiVerse shows that language model embeddings paired with simple, well-regularized classifiers are sufficient (and often superior) for practical peptide property prediction. This design reduces computational cost while improving generalizability, making PeptiVerse compatible with modern peptide design pipelines. The advantage is particularly clear for binding affinity, where structure prediction confidence alone fails to track peptide-protein binding strength (Supplementary Fig. S4), whereas PeptiVerse yields strong, statistically significant agreement with experimental measurements (Supplementary Table S5). Together, these results motivate fast, data-driven affinity predictors that complement rather than replace structural modeling, and position PeptiVerse as an open, extensible benchmark for therapeutic peptide discovery.
Importantly, PeptiVerse is designed as a lightweight and accessible framework that supports direct incorporation of property predictors into iterative and high-throughput design workflows, rather than limiting use to web-based post hoc evaluation57. This design enables a natural application to generative peptide sequence modeling, where efficient property evaluation is required to guide optimization. In this setting, PeptiVerse predictors enable rapid assessment of therapeutically relevant properties, allowing generative models to produce sequences with a higher likelihood of translating in wet-lab settings29,30,32,33,34. Such guidance relies on fast reward evaluation to support gradient-based or iterative refinement over generated sequences58, particularly in multi-objective settings where properties such as binding affinity, solubility, and toxicity must be jointly optimized.
In practice, PeptiVerse predictors have already been integrated with state-of-the-art generative frameworks, including PepTune33, which introduces Monte Carlo Tree Guidance for multi-objective-guided peptide optimization, TR2-D234, which applies reinforcement learning for reward fine-tuning of peptide discrete diffusion models, and moPPIt28, which enables motif-specific peptide generation via discrete flow matching and is experimentally verified across multiple target classes. Together, these examples position PeptiVerse as a modular reward evaluation layer that integrates naturally into modern generative pipelines, supporting both inference-time steering and training-time optimization.
To ensure continued relevance, PeptiVerse will be updated regularly as new data and predictors become available. While several property models already achieve strong performance, others (i.e., chemically-modified peptide half-life in SMILES space) remain constrained by limited and heterogeneous experimental data. Importantly, the modular design of PeptiVerse enables seamless integration of improved representation models as they emerge; for example, retraining with newer embeddings such as ESM-C59 provides performance gains over ESM-2 on properties such as half-life and solubility (Supplementary Tables S8 and S10), further supporting that advances in representation learning directly translate to improved predictive accuracy within this framework. We therefore encourage open-source deposition of peptide property measurements from both academia and industry. In parallel, we are developing new specialized binding predictors, including peptide isoform-specificity36, motif-specificity28, and metal-binding propensity35, which will be incorporated as additional properties within the PeptiVerse framework. By providing a standardized, extensible, and openly accessible platform for integrating new data and models, PeptiVerse establishes a practical mechanism for expanding predictive scope, enabling property-guided generative design, and supporting the translation of next-generation peptide therapeutics.
Methods
Data collection and preparation
Throughout this work, we distinguish peptide inputs by representation modality rather than biological origin. “Amino acid” inputs refer to canonical sequence-based representations processed by protein language models, while “SMILES” inputs refer to chemistry-aware molecular representations, which may include both canonical and non-canonical peptides.
Hemolysis
Hemolysis data were retrieved from PeptideBert and peptide-dashboard15,25, and cross-validated against the original experimental records in DBAASP v3.041. Peptides labeled as 1 are considered hemolytic, whereas 0 denotes non-hemolytic activity. The final dataset comprised 4765 non-hemolytic and 1311 hemolytic peptide entries.
Permeability
All permeability annotations were obtained from PepLand23, which sources experimental measurements from CycPeptMPDB16. The noncanonical dataset contains 7475 non-canonical peptides with reported permeability values measured using either PAMPA49 or Caco-2 assays50. Permeability is reported as \(\log {P}_{\exp }\), the logarithm of the effective permeability coefficient, which reflects the peptide’s lipophilicity and ability to passively diffuse across lipid membranes.
PAMPA and Caco-2 assays quantify different biological processes, with PAMPA measuring passive membrane permeability and Caco-2 assays relating more closely to intestinal absorption potential49,60,61. For this reason, the two assay types were handled separately during data preparation, not following PepLand’s combined training strategy. There were 6,869 PAMPA sequences and 606 Caco-2 sequences collected in the end. Following CycPeptMPDB conventions, peptides with \(\log {P}_{\exp }\ge -6.0\) were labeled as high permeability, and the remaining peptides as weak permeability16. The predictive tasks associated with these datasets are called Permeability (PAMPA) and Permeability (Caco-2).
The canonical permeability dataset contains 1162 cell-penetrating peptides and 1162 non-penetrating peptides, curated from PepLand23. This dataset was constructed to balance peptide length distributions between positive and negative classes. Positive peptides were originally collected from 22 independent cell-penetrating peptide studies, while negative examples were sourced from UniProt. A notable fraction of the positive peptides (~ 8.7%) exhibited low sequence complexity, defined as a ratio of peptide length to the number of unique amino acids greater than 5, compared to only 0.95% among non-penetrating peptides. Although these low-complexity sequences are likely engineered, they were retained, as they may encode informative features relevant to membrane permeability23. The predictive task associated with this dataset is called Permeability_CPP.
Non-fouling
Non-fouling annotations were obtained from PeptideBERT and Peptide-Dashboard15,25, both of which source data from the dataset of Barrett et al.47. In this context, non-fouling peptides are defined as sequences that resist nonspecific protein adsorption, while fouling peptides permit such adsorption, which can lead to functional loss and reduced performance47,62. Peptides labeled as 1 are classified as non-fouling, and those labeled as 0 are considered fouling. The curated dataset comprised 13,580 fouling and 3600 non-fouling peptide entries.
Toxicity
Toxicity data were obtained from ToxinPred3.0, which provides canonical amino acid peptide sequences with experimentally validated toxicity labels42. The dataset contains 5518 toxic and 5518 non-toxic peptides, where label 1 denotes toxic, and 0 denotes non-toxic. Amino acid sequences were converted into SMILES format with the fasta2smi command from p2smi63. Molecular redundancy was reduced by clustering peptides using Morgan fingerprints (radius 2, 2048 bits, including chirality) and RDKit’s Butina clustering algorithm with a similarity threshold of 0.6.
Solubility
Solubility labels follow the protocol established in PROSO II43, as used in PeptideBERT15, with additional sequences incorporated from SoluProtMutDB48. In PROSO II, the soluble class (1) was assigned based on experimental metadata recorded in pepcDB, following the Protein Structure Initiative (PSI) pipelines64. A sequence was labeled soluble once it reached the “Soluble" stage or any downstream stage. Additional soluble entries were derived from PDB records (up to 2010) annotated with expression from E. coli, as crystallographic and NMR structure determination requires a protein that has been purified in solution. The insoluble class (0) consisted of constructs that remained in the “not Soluble" state for at least eight months, based on the different pepcDB releases. This yields a combined dataset of 8,785 soluble and 9668 non-soluble sequences in total.
Binding affinity
Binding affinity prediction was performed using paired protein sequences and peptide SMILES from the PepLand dataset23, which aggregates different experimental measurements together (Kd, Ki, IC50). Their data collection protocol follows CAMP51, which subsets data from RCSB PDB65 and DrugBank entries containing the label of “peptide”66. All scores were transformed based on the negative logarithm of the original affinity data into a unified scale. A shared unit scale was roughly designed to indicate the strength of binding: with 9 indicating strong nM to pM binders, 7–9 indicating nM to μM medium binders, and < 7 indicating weak μM binders. The data split was based on affinity score distribution matching, ensuring both splits contain a similar distribution of data. The final dataset contained 1433 peptide-protein pairs with canonical peptide sequences (amino acid inputs) and 1702 pairs with SMILES-encoded peptides containing non-canonical amino acids.
Thus, higher values indicate stronger binding, which is more intuitive for modeling.
Stability
Stability data for amino acid sequence inputs were collected and organized from TAPE67 and SaProt68 (Huggingface: SaProtHub/Dataset-Stability-TAPE) and used to pretrain stability predictors before half-life modeling. Stability was treated as a continuous score reflecting the ability of proteins to remain folded above concentration thresholds. A total of 68,845 sequences were collected.
Half-life
Half-life data were compiled from THPdb244, PEPlife45, and PepTherDia46. Only human serum measurements were retained. Reported values across datasets vary in units and granularity. All half-life measurements were therefore converted into hours for interpretability. Details of the data construction are provided in the Supplementary Information. The curated dataset comprised 130 amino acid sequence entries and 245 SMILES-based sequences.
Protein-peptide ipTM scores
The interface predicted TM-score (ipTM) is a confidence metric originally designed to assess the accuracy of predicted protein-protein interaction interfaces in multimeric structure prediction models69. All ipTM scores for protein-peptide pairs, including SMILES-encoded pairs, were obtained from OpenFold354 via its API on the NVIDIA NGC platform70. MSA profiles for the target proteins were constructed via MMseqs271 against the Uniref30 database72.
Additional properties
Additional physicochemical features, including isoelectric point, molecular weight, and hydrophobicity, were computed using utilities from the Biopython package73. These properties were calculated dynamically with a tunable pH parameter to account for protonation-state dependence.
Data splitting
All properties defined on amino acid-based peptide representations were split using a shared clustering strategy. Sequences from all amino acid sequence datasets were clustered independently using MMseqs271 with identical parameters (–min-seq-id 0.3 -c 0.8 –cov-mode 0) and split using an 80/20 cluster-level split to prevent sequence leakage between training and evaluation subsets. Amino acid sequences were converted into SMILES format with the fasta2smi command from p2smi63. For properties evaluated using both sequence and SMILES inputs (hemolysis, non-fouling, solubility, permeability_CPP), the original cluster-based train/validation assignments were preserved after conversion, ensuring fair comparison across input modalities. For datasets natively represented in SMILES notation, molecules were clustered using Morgan fingerprints and Tanimoto distance via RDKit74, followed by an 80/20 cluster-level split. Datasets used for binding affinity prediction, which contain paired peptide-protein inputs across two molecular modalities, were instead split by matching affinity score distributions between training and evaluation sets. Detailed dataset sizes, label distributions, and train/validation splits for all properties are summarized in Table 2 and Fig. 2.
Model architecture and training
Sequence and SMILES representations
Protein sequences were represented using embeddings from ESM-2 (esm2_t33_650M_UR50D)26, while peptide sequences were represented using either ESM2, PeptideCLM (PeptideCLM-23M-all)27, which employs a tokenizer of size 586 designed to capture noncanonical peptide chemistry, or ChemBERTa (ChemBERTa-77M-MLM)75, trained across millions of chemical structures curated from PubChem. See Supplementary Methods for details on other embedding approaches, including ESM-C and baseline embeddings. Depending on the downstream model, embeddings were either average-pooled across sequence positions to form fixed-length representations or retained in unpooled, position-resolved form to preserve residue-level and positional information. This distinction allowed different model classes to exploit either global sequence summaries or fine-grained positional structure.
Predictor architectures
For properties including toxicity, solubility, hemolysis, non-fouling, and permeability, lightweight predictor heads were trained on frozen foundational embeddings. We evaluated a broad set of predictor architectures, including multilayer perceptrons (MLPs), convolutional neural networks (CNNs), and transformer-based models, as well as classical statistical learners such as Elastic Nets (ENET)38,76, support vector machines (SVMs)37, epsilon-support vector regression (SVR)77, and XGBoost (XGB)39. ENET and SVM classifiers were implemented using the RAPIDS cuML library78 to enable GPU acceleration, while SVR models were implemented using scikit-learn79. Classification models were trained using standard objective functions appropriate to each method, whereas regression models were optimized using mean squared error. All models were trained using their canonical formulations without architectural modification. For CNN- and transformer-based predictors, unpooled embeddings were used to allow convolutional kernels and self-attention mechanisms to explicitly exploit positional structure, while pooled embeddings were used for MLP and tree-based models operating on fixed-dimensional inputs.
Hyperparameter optimization
Hyperparameters for all models were optimized using the Optuna framework80, with an initial target of 200 optimization trials per configuration (Supplementary Tables S12–S15). For parameters sampled on a logarithmic scale, values were drawn uniformly in log space to efficiently explore multiple orders of magnitude. For regression tasks, both mean squared error (MSE) and Huber loss were evaluated, using the latter to improve robustness to outliers via its δ parameter controlling the transition between L1 and L2 penalties. For computationally intensive configurations exceeding approximately 10 hours of wall-clock training time, the number of Optuna trials was reduced to 50 or 20 while preserving representative coverage of the hyperparameter space. Results of the hyperparameter exploration are reported in Supplementary Figs. S2 and S5.
Binding affinity prediction
Binding affinity prediction was formulated as a regression task using a transformer-based architecture with cross multi-head attention to learn a joint latent representation between peptide and protein modalities. The model architecture was fixed across experiments, with variation introduced only through peptide embedding initialization. Peptide inputs were represented using either ESM-2 or PeptideCLM embeddings26,27, in pooled or unpooled form, while protein targets were consistently represented using ESM-2 embeddings. The model produced a continuous affinity score, which was additionally mapped to three discrete affinity classes, yielding a multitask framework aligned with other property predictors. Hyperparameters were optimized using 200 Optuna trials, and Spearman’s rank correlation coefficient (ρ) was used as the primary model selection criterion (Supplementary Table S15 and Supplementary Fig. S2).
Half-life prediction
Peptide half-life prediction was formulated as a regression task using both amino acid sequence inputs and SMILES-based representations. For amino acid-based inputs, models were initialized from pre-trained stability predictors trained on the wild-type stability dataset and subsequently fine-tuned on half-life data. The same model families used for binary property prediction were retained. Four regression configurations were evaluated: (i) XGBoost predicting half-life in hours, (ii) XGBoost predicting \(\log (1+\,{{\rm{half-life}}})\), (iii) a transformer-based regressor predicting half-life in hours, and (iv) a transformer-based regressor predicting \(\log (1+\,{{\rm{half-life}}})\). Hyperparameters were optimized using Optuna with five-fold cross-validation, using 200 trials for XGB, SVR, and ENET models and 50 trials for CNN, MLP, and Transformer models, with 20 training epochs per trial.
The transformation equation is defined as follows:
where t1/2 is the half-life in hours.
For SMILES-based half-life prediction, models were trained directly on peptide SMILES embeddings without intermediate stability pretraining. Predictor architectures matched those used for amino acid-based half-life prediction and were trained to predict \(\log (1+\,{{\rm{half-life}}})\). Evaluation was performed using five-fold cross-validation. Optuna optimization employed 200 trials for XGB, SVR, and ENET models and 50 trials for CNN, MLP, and Transformer models, with transformer-based models trained for 100 epochs per trial.
Evaluation metrics
Model evaluation metrics were selected according to task type. Classification performance was assessed using F1 score and area under the receiver operating characteristic curve (AUC) with results reported in Supplementary Fig. S1 and Supplementary Table S3. Regression performance was evaluated using Pearson correlation coefficient (r), Spearman’s rank correlation coefficient (ρ), and the coefficient of determination (R2), with results summarized in Supplementary Fig. S3.
Model stability, confidence intervals, and uncertainty estimation
All deep learning models (MLP, CNN, Transformer) were retrained five times with distinct random seeds (1986, 42, 0, 123, 12345) using fixed optimal hyperparameters from OPTUNA. The 95% confidence intervals for classification (F1, AUC) and regression (Spearman ρ, RMSE, R2) metrics were estimated as \(\bar{x}\pm {t}_{0.025,4}\cdot \frac{s}{\sqrt{5}}\). Per-sequence uncertainty was quantified as the standard deviation of predictions across seeds.
For classical ML models (SVM, ElasticNet), optimization is convex and therefore deterministic given fixed data and hyperparameters81, leading to negligible variability across random seeds. For XGBoost, while stochastic subsampling of features and instances introduces minor seed dependence39, empirical variability is small relative to data sampling variability. Confidence intervals for all three model types were therefore derived via non-parametric percentile bootstrap (n = 2000) over held-out validation predictions \(({y}_{i},{\widehat{y}}_{i})\)82 without retraining. The percentile method was preferred over symmetric t-intervals to respect the bounded range of Spearman ρ ∈ [ − 1, 1], with Fisher z-transform intervals computed as a cross-check when ρ > 0.9. The decision threshold was fixed on the full validation set and held constant across bootstrap resamples to avoid optimistic bias. Per-sequence uncertainty was estimated as the probability margin from the decision boundary for probabilistic classifiers, including XGBoost (ci = 2∣pi − 0.5∣), and as empirical residual intervals (\({\widehat{y}}_{i}\pm {z}_{0.975}\widehat{\sigma }\)) for deterministic regression models (SVR, ElasticNet).
For the half-life regression models, 95% confidence intervals were computed by non-parametric percentile bootstrap over out-of-fold prediction pairs \(({y}_{i},{\widehat{y}}_{i})\) obtained from cross-validation without retraining. For each bootstrap resample, RMSE, MAE, R2, and Spearman’s ρ were recomputed, and the 2.5th and 97.5th percentiles of the resulting empirical distributions were reported as the interval bounds. In addition to the point estimate on the full validation set, we also report the bootstrap mean and standard deviation for each metric.
Ablation of full fine-tuning and LoRA for half-life prediction
To further assess the role of parameter-efficient adaptation in this data-limited setting, we performed an additional half-life transfer-learning ablation comparing full fine-tuning and low-rank adaptation (LoRA) using the stability-pretrained transformer83 (Supplementary Table S11). This comparison was restricted to the transformer architecture, as it was the strongest sequence-based stability model and the most appropriate setting for LoRA-based adaptation. In the LoRA setting, the pretrained stability checkpoint was first loaded, selected linear layers in the transformer were replaced with LoRA-augmented modules while the backbone weights were frozen, and only the low-rank update matrices and the final regression head were trained. Performance was evaluated for both raw and log-transformed half-life targets. Full fine-tuning outperformed LoRA under both target formulations, with the best results obtained for full fine-tuning on log-transformed half-life, whereas LoRA remained competitive but consistently lower-performing. These results suggest that transfer from stability prediction to half-life prediction benefits from broader adaptation of the pretrained representation, while log-space prediction provides a more favorable learning target for this task.
Inference-time confidence estimation
To provide confidence scores per-prediction at inference time, different strategies were applied according to the model class, following the framework of Lakshminarayanan et al.84.
For classification tasks, uncertainty is quantified as the predictive entropy of the model’s output probability. For deep learning classifiers (MLP, CNN, Transformer), inference-time uncertainty was quantified using the deep ensemble of five independently seeded models. Given a query input x, each model m ∈ {1, …, 5} produces output probability \({p}_{{\theta }_{m}}(y| {{\bf{x}}})\). The ensemble predictive distribution is the uniformly-weighted mixture \(\bar{p}(y| {{\bf{x}}})=\frac{1}{5}{\sum }_{m=1}^{5}{p}_{{\theta }_{m}}(y| {{\bf{x}}})\). For classical classifiers (XGBoost, SVM), \(\bar{p}(y| {{\bf{x}}})\) is the single model’s output probability. For linear classifiers (ElasticNet) that produce a decision function score f(x) rather than a native probability, \(\bar{p}(y| {{\bf{x}}})=\sigma (f({{\bf{x}}}))\) where σ denotes the sigmoid function. In both cases, per-sample uncertainty is quantified as the binary predictive entropy:
High entropy indicates either ensemble disagreement (epistemic uncertainty) or collectively diffuse predictions (aleatoric uncertainty). The deep ensemble formulation is preferred over MC Dropout85 as deep ensembles explore multiple modes of the loss landscape via independent random initialization, whereas MC Dropout samples within a single basin, leading to systematically overconfident predictions on out-of-distribution inputs84.
For regression tasks, uncertainty estimation depends on the model class. For deep learning regression models (MLP, CNN, Transformer), inference-time uncertainty is reported as the standard deviation of point predictions across the five ensemble members:
where \(\bar{y}({{\bf{x}}})=\frac{1}{5}{\sum }_{m} \, {\widehat{y}}_{m}({{\bf{x}}})\) is the ensemble mean prediction. For deterministic classical regression models (SVR, ElasticNet) where seed-based ensembling is inapplicable due to convex optimization yielding identical solutions, and for XGBoost regression, where retraining seed ensembles was outside the scope of this work, inference-time prediction intervals were constructed using the residual normalized conformity score86,87. Rather than using MAPIE’s sklearn API directly, we adopt a custom implementation since our base models embed raw sequences via GPU-accelerated transformer forward passes which are incompatible with sklearn’s fit(X, y) interface, the binding affinity model requires dual sequence inputs, and held-out predictions were already precomputed during training. An auxiliary XGBoost residual estimator \(\widehat{\sigma }\), that provides stronger nonlinear modeling than the linear default in MAPIE, is fitted on precomputed held-out embeddings and absolute residuals \(| {y}_{i}-{\widehat{y}}_{i}|\). Normalized conformity scores \({s}_{i}=| {y}_{i}-{\widehat{y}}_{i}| /{\hat\sigma}({{{\bf{x}}}}_{i})\) are then computed, and the conformal quantile defined as:
is taken at the finite-sample corrected level, guaranteeing marginal coverage ≥ 1 − α = 0.9 under exchangeability88. At inference, the prediction interval \([\widehat{y}({{\bf{x}}})-q\cdot \hat{\sigma }({{\bf{x}}}),\,\widehat{y}({{\bf{x}}})+q\cdot \hat{\sigma }({{\bf{x}}})]\) has width that varies per input, since \(\hat{\sigma }({{\bf{x}}})\) captures local prediction difficulty.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Data availability
All processed datasets used to train PeptiVerse predictors are available at https://huggingface.co/datasets/ChatterjeeLab/PeptiVerse_data. All benchmarking model weights are available in Zenodo: https://zenodo.org/records/19989009 and the best model weights are available at https://huggingface.co/ChatterjeeLab/PeptiVerse. Additional figures and tables are provided in a Supplementary Information file. Source data for all figure plots and tables, in both the main text and Supplementary Information file, are provided with this paper in the Source Data file. Source data are provided in this paper.
Code availability
The full set of predictors is available through a simple API at https://huggingface.co/ChatterjeeLab/PeptiVerse, which can also be used to download datasets. Code for plotting in R is available at https://github.com/ynuozhang/Peptiverse_R.git. For users who prefer a no-code interface, all predictors can also be easily run via an interactive HuggingFace Space at https://huggingface.co/spaces/ChatterjeeLab/PeptiVerse. All code is open source, with predictors and associated resources updated and maintained regularly.
References
- Zheng, Z. et al. Glucagon-like peptide-1 receptor: mechanisms and advances in therapy. Signal Transduct. Target. Ther. 9, 234 (2024).
- Drucker, D. J. Discovery of glp-1–based drugs for the treatment of obesity. N. Engl. J. Med. 392, 612–615 (2025).
- Chen, L. T. et al. Target sequence-conditioned design of peptide binders using masked language modeling. Nat. Biotechnol. 44, 1002–1010 (2025).
- Wang, L. et al. Therapeutic peptides: current applications and future directions. Signal Transduct. Target. Ther. 7, 48 (2022).
- Tang, S., Han, E. L. & Mitchell, M. J. Peptide-functionalized nanoparticles for brain-targeted therapeutics. Drug Deliv. Transl. Res. 16, 741–760 (2025).
- Chen, T., Hong, L., Yudistyra, V., Vincoff, S. & Chatterjee, P. Generative design of therapeutics that bind and modulate protein states. Curr. Opin. Biomed. Eng. 28, 100496 (2023).
- Storchmannová, K., Balouch, M., Juračka, J., Štěpánek, F. & Berka, K. Meta-analysis of permeability literature data shows possibilities and limitations of popular methods. Mol. Pharmaceut. 22, 1293–1304 (2025).
- Chen, G., Kang, W., Li, W., Chen, S. & Gao, Y. Oral delivery of protein and peptide drugs: From non-specific formulation approaches to intestinal cell targeting strategies. Theranostics 12, 1419 (2022).
- Werle, M. & Bernkop-Schnürch, A. Strategies to improve plasma half life time of peptide and protein drugs. Amino Acids 30, 351–367 (2006).
- Malavolta, L., Pinto, M. R., Cuvero, J. H. & Nakaie, C. R. Interpretation of the dissolution of insoluble peptide sequences based on the acid-base properties of the solvent. Protein Sci. 15, 1476–1488 (2006).
- Zapadka, K. L., Becher, F. J., Gomes dos Santos, A. L. & Jackson, S. E. Factors affecting the physical stability (aggregation) of peptide therapeutics. Interface Focus 7, 20170030 (2017).
- Kellermeyer, R. W. Hemolytic effect of therapeutic drugs: Clinical considerations of the primaquine-type hemolysis. JAMA 180, 388 (1962).
- Timmons, P. B. & Hewage, C. M. Happenn is a novel tool for hemolytic activity prediction for therapeutic peptides which employs neural networks. Sci. Rep. 10, 10869 (2020).
- Jiang, S. & Cao, Z. Ultralow-fouling, functionalizable, and hydrolyzable zwitterionic materials and their derivatives for biological applications. Adv. Mater. 22, 920–932 (2010).
- Guntuboina, C., Das, A., Mollaei, P., Kim, S. & Barati Farimani, A. Peptidebert: A language model based on transformers for peptide property prediction. J. Phys. Chem. Lett. 14, 10427–10434 (2023).
- Li, J. et al. Cycpeptmpdb: a comprehensive database of membrane permeability of cyclic peptides. J. Chem. Inf. Model. 63, 2240–2250 (2023).
- DeGruyter, J. N., Malins, L. R. & Baran, P. S. Residue-specific peptide modification: a chemist’s guide. Biochemistry 56, 3863–3873 (2017).
- Notin, P. et al. ProteinGym: Large-scale benchmarks for protein fitness prediction and design. Thirty-seventh Conference on Neural Information Processing Systems (2023).
- Xie, J., Li, Y. & Fu, T. Deepprotein: deep learning library and benchmark for protein sequence learning. Bioinformatics 41, btaf165 (2025).
- Cabas-Mora, G. et al. Peptipedia v2.0: a peptide sequence database and user-friendly web platform. a major update. Database (2024).
- Fu, L. et al. Admetlab 3.0: an updated comprehensive online admet prediction platform enhanced with broader coverage, improved performance, api functionality and decision support. Nucleic Acids Res. 52, W422–W431 (2024).
- Swanson, K. et al. Admet-ai: a machine learning admet platform for evaluation of large-scale chemical libraries. Bioinformatics 40, btae416 (2024).
- Zhang, R. et al. Pepland: a large-scale pre-trained peptide representation model for a comprehensive landscape of both canonical and non-canonical amino acids. Brief. Bioinform. 26, bbaf367 (2025).
- Wang, L. et al. PepDORA: A unified peptide language model via weight-decomposed low-rank adaptation. NeurIPS Workshop on AI for New Drug Modalities (2024).
- Ansari, M. & White, A. D. Serverless prediction of peptide properties with recurrent neural networks. J. Chem. Inf. Model. 63, 2546–2553 (2023).
- Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023).
- Feller, A. L. & Wilke, C. O. Peptide-aware chemical language model successfully predicts membrane diffusion of cyclic peptides. J. Chem. Inf. Model. 65, 571–579 (2025).
- Chen, T. et al. moPPIt: De novo generation of motif-specific and functionally active peptide binders via discrete flow matching. Preprint at https://doi.org/10.1101/2024.07.31.606098 (2026).
- Chen, T., Zhang, Y., Tang, S. & Chatterjee, P. Multi-objective-guided discrete flow matching for controllable biological sequence design. ICML Generative AI and Biology (GenBio) Workshop (2025).
- Chen, T., Zhang, Y. & Chatterjee, P. Areuredi: Annealed rectified updates for refining discrete flows with multi-objective guidance. ICLR Workshop on Deep Generative Model in Machine Learning: Theory, Principle and Efficacy (2025).
- Goel, S. et al. Token-level guided discrete diffusion for membrane protein design. NeurIPS AI4Science Workshop (2025).
- Tang, S., Zhang, Y., Tong, A. & Chatterjee, P. Gumbel-softmax flow matching with straight-through guidance for controllable biological sequence generation. ICLR 2025 Workshop on AI for Nucleic Acids (2025).
- Tang, S., Zhang, Y. & Chatterjee, P. PepTune: De novo generation of therapeutic peptides with multi-objective-guided discrete diffusion. Forty-second International Conference on Machine Learning (2025).
- Tang, S., Zhu, Y., Tao, M. & Chatterjee, P. TR2-D2: Tree search guided trajectory-aware fine-tuning for discrete diffusion. arXiv preprint arXiv:2509.25171 (2025).
- Zhang, Y., Srijay, D., Quinn, Z. & Chatterjee, P. Metalorian: De novo generation of heavy metal-binding peptides with classifier-guided diffusion sampling. Preprint at https://doi.org/10.1101/2025.07.10.664242 (2025).
- Vincoff, S., Liu, H., Linardic, C. & Chatterjee P. SOAPIA: Specificity-Guided Generation of Off-Target-Avoiding Protein Interactions with High Target Affinity. ICML Workshop on Generative and Agentic AI for Biology (2026).
- Hearst, M. A., Dumais, S. T., Osuna, E., Platt, J. & Scholkopf, B. Support vector machines. IEEE Intell. Syst. their Appl. 13, 18–28 (1998).
- Simoncini, V. Variable accuracy of matrix-vector products in projection methods for eigencomputation. SIAM J. Numer. Anal. 43, 1155–1174 (2005).
- Chen, T. et al. xgboost: Extreme gradient boosting. CRAN: Contributed Packages (2014).
- De Landsheere, J., Zamyatin, A., Karwounopoulos, J. & Heid, E. Chemtorch: A deep learning framework for benchmarking and developing chemical reaction property prediction models. 66, 2434–2442 (2025).
- Pirtskhalava, M. et al. Dbaasp v3: database of antimicrobial/cytotoxic activity and structure of peptides as a resource for development of new therapeutics. Nucleic Acids Res. 49, D288–D297 (2021).
- Rathore, A. S., Choudhury, S., Arora, A., Tijare, P. & Raghava, G. P. Toxinpred 3.0: An improved method for predicting the toxicity of peptides. Comput. Biol. Med. 179, 108926 (2024).
- Smialowski, P., Doose, G., Torkler, P., Kaufmann, S. & Frishman, D. Proso ii–a new method for protein solubility prediction. FEBS J. 279, 2192–2200 (2012).
- Jain, S., Gupta, S., Patiyal, S. & Raghava, G. P. Thpdb2: compilation of fda approved therapeutic peptides and proteins. Drug Discov. Today 29, 104047 (2024).
- Mathur, D. et al. Peplife: a repository of the half-life of peptides. Sci. Rep. 6, 36617 (2016).
- D’Aloisio, V., Dognini, P., Hutcheon, G. A. & Coxon, C. R. Peptherdia: database and structural composition analysis of approved peptide therapeutics and diagnostics. Drug Discov. Today 26, 1409–1419 (2021).
- Barrett, R., Jiang, S. & White, A. D. Classifying antimicrobial and multifunctional peptides with bayesian network models. Pept. Sci. 110, e24079 (2018).
- Velecký, J. et al. Soluprotmutdb: A manually curated database of protein solubility changes upon mutations. Comput. Struct. Biotechnol. J. 20, 6339–6347 (2022).
- Avdeef, A. The rise of pampa. Expert Opin. Drug Metab. Toxicol. 1, 325–342 (2005).
- Artursson, P., Palm, K. & Luthman, K. Caco-2 monolayers in experimental and theoretical predictions of drug transport. Adv. Drug Deliv. Rev. 64, 280–289 (2012).
- Lei, Y. et al. A deep-learning framework for multi-level peptide–protein interaction prediction. Nat. Commun. 12, 5465 (2021).
- Zhu, W., Shenoy, A., Kundrotas, P. & Elofsson, A. Evaluation of alphafold-multimer prediction on multi-chain protein complexes. Bioinformatics 39, btad424 (2023).
- Peng, C., Ni, W., Liu, Q., Hu, G. & Zheng, W. A comprehensive benchmarking of the alphafold3 for predicting biomacromolecules and their interactions. Brief. Bioinform. 26, bbaf616 (2025).
- The OpenFold3 Team. Openfold3-preview https://github.com/aqlaboratory/openfold-3. (2025).
- Abid, A. et al. Gradio: Hassle-free sharing and testing of ml models in the wild. ICML Workshop on Human in the Loop Learning (2019).
- Jain, S. M. Hugging face. In Introduction to transformers for NLP: With the Hugging Face Library and Models to Solve Problems, 51–67 (Springer, 2022).
- Tan, X. et al. pepadmet: A novel computational platform for systematic admet evaluation of peptides. J. Chem. Inf. Model. 66, 936–946 (2026).
- Nisonoff, H., Xiong, J., Allenspach, S. & Listgarten, J. Unlocking guidance for discrete state-space diffusion and flow models. In The Thirteenth International Conference on Learning Representations (2025).
- EvolutionaryScale. Esm cambrian: Revealing the mysteries of proteins with unsupervised learning. https://www.evolutionaryscale.ai/blog/esm-cambrian (2024).
- Ottaviani, G., Martel, S. & Carrupt, P.-A. Parallel artificial membrane permeability assay: a new membrane for the fast prediction of passive human skin permeability. J. Med. Chem. 49, 3948–3954 (2006).
- Van Breemen, R. B. & Li, Y. Caco-2 cell permeability assays to measure drug absorption. Expert Opin. Drug Metab. Toxicol. 1, 175–185 (2005).
- Banerjee, I., Pangule, R. C. & Kane, R. S. Antifouling coatings: recent developments in the design of surfaces that prevent fouling by proteins, bacteria, and marine organisms. Adv. Mater. 23, 690–718 (2011).
- Feller, A. L. & Wilke, C. O. p2smi: A toolkit enabling smiles generation and property analysis for noncanonical and cyclized peptides. J. Open Source Softw. 10, 8319 (2025).
- Berman, H. M. et al. The protein structure initiative structural genomics knowledgebase. Nucleic Acids Res. 37, D365–D368 (2009).
- Kouranov, A. et al. The rcsb pdb information portal for structural genomics. Nucleic Acids Res. 34, D302–D305 (2006).
- Knox, C. et al. Drugbank 6.0: the drugbank knowledgebase for 2024. Nucleic Acids Res. 52, D1265–D1275 (2023).
- Rao, R. et al. Evaluating protein transfer learning with tape. Adv. Neural Inf. Process. Syst. 32, 9689–9701 (2019).
- Su, J. et al. SaProt: Protein language modeling with structure-aware vocabulary. The Twelfth International Conference on Learning Representations (2024).
- Abramson, J. et al. Accurate structure prediction of biomolecular interactions with alphafold 3. Nature 630, 493–500 (2024).
- NVIDIA Corporation. Explore biology models ∣ try NVIDIA NIM APIshttps://build.nvidia.com/explore/biology (2025).
- Steinegger, M. & Söding, J. Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol. 35, 1026–1028 (2017).
- Consortium, U. Uniprot: a worldwide hub of protein knowledge. Nucleic Acids Res. 47, D506–D515 (2019).
- Cock, P. J. A. et al. Biopython: freely available python tools for computational molecular biology and bioinformatics. Bioinformatics 25, 1422–1423 (2009).
- Landrum, G. et al. rdkit/rdkit: 2025_03_1 (q1 2025) release. Zenodo https://doi.org/10.5281/zenodo.15115844 (2025).
- Chithrananda, S., Grand, G. & Ramsundar, B. Chemberta: large-scale self-supervised pretraining for molecular property prediction. Preprint at https://doi.org/10.48550/arXiv.2010.09885 (2020).
- Zou, H. & Hastie, T. Regularization and variable selection via the elastic net. J. R. Stat. Soc. Ser. B: Stat. Methodol. 67, 301–320 (2005).
- Carrasco, M., López, J. & Maldonado, S. Epsilon-nonparallel support vector regression. Appl. Intell. 49, 4223–4236 (2019).
- Raschka, S., Patterson, J. & Nolet, C. Machine learning in python: Main developments and technology trends in data science, machine learning, and artificial intelligence. Information 2020 11, 193 (2020).
- Pedregosa, F. et al. Scikit-learn: Machine learning in python. J. Mach. Learn. Res. 12, 2825–2830 (2011).
- Akiba, T., Sano, S., Yanase, T., Ohta, T. & Koyama, M. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2623–2631 (2019).
- Boyd, S. & Vandenberghe, L. Convex Optimization (Cambridge University Press, 2004).
- Efron, B. & Tibshirani, R. J. An Introduction to the Bootstrap (Chapman and Hall/CRC, 1994).
- Hu, E. J. et al. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations (2022).
- Lakshminarayanan, B., Pritzel, A. & Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. In Proceedings of the Advances in Neural Information Processing Systems (2017).
- Gal, Y. & Ghahramani, Z. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In International Conference on Machine Learning, 1050–1059 (PMLR, 2016).
- Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R. J. & Wasserman, L. Distribution-free predictive inference for regression. J. Am. Stat. Assoc. 113, 1094–1111 (2018).
- Taquet, V., Blot, V., Morzadec, T., Lacombe, L. & Brunel, N. MAPIE: an open-source library for distribution-free uncertainty quantification. ICML Workshop on Distribution-Free Uncertainty Quantification (2022).
- Vovk, V., Gammerman, A. & Shafer, G. Algorithmic Learning in a Random World (Springer, 2005).
Acknowledgements
We thank the entire experimental team of the Chatterjee Lab for productive input during the curation of these models and datasets. We also thank Lauren Hong for designing the PeptiVerse logo and J. Chatterjee for emotional support.
Funding
This research was supported by the Hartwell Foundation, NIH grant R35GM155282, and by a pilot grant from the High-throughput Institute for Discovery (HIT-ID) at the University of Pennsylvania to the lab of P.C. E.H.M. discloses support for the research of this work from NSF PRFB Grant No. 2410554.
Author information
Authors and Affiliations
Contributions
Y.Z., S.T., T.C., E.M. and S.V. trained models, curated datasets, and prepared results. Y.Z. managed dataset and model curation and wrote the paper, with input from S.T. and T.C. P.C. conceived, designed, and directed the study, and reviewed and finalized the manuscript.
Corresponding author
Ethics declarations
Competing interests
P.C. is a co-founder of Gameto, Inc., UbiquiTx, Inc., AtomBioworks, Inc., and Recognition Bio., Inc., and advises companies involved in peptide therapeutics development. P.C.’s interests are reviewed and managed by the University of Pennsylvania in accordance with their competing-of-interest policies. The remaining authors have no competing interests to declare.
Peer review
Peer review information
Nature Communications thanks the anonymous reviewers for their contribution to the peer review of this work. A peer review file is available.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
About this article
Cite this article
Zhang, Y., Tang, S., Chen, T. et al. PeptiVerse: A unified platform for therapeutic peptide property prediction. Nat Commun 17, 6819 (2026). https://doi.org/10.1038/s41467-026-74167-w
- Received:
- Accepted:
- Published:
- Version of record:
- DOI: https://doi.org/10.1038/s41467-026-74167-w