Studies combining MALDI-TOF with AI models to identify and classify viruses remain scarce and challenging in scientific literature. To address this, we developed a comprehensive workflow encompassing viral sample preparation, spectral data acquisition, and model training and evaluation. For this feasibility study, we utilized five viruses representative of various viral groups, including both DNA and RNA viruses, as well as enveloped and non-enveloped viruses. Key steps in this workflow involved viral particle enrichment followed by efficient protein extraction.
In clinical virology, several MALDI-TOF sample preparation approaches have been reported including chemical lysis, solvent-based extraction of viral particles and direct spectral analysis of virus-infected cells compared to an uninfected control44,45,46. The direct analysis method relies on examining spectral changes in the protein profiles of infected cells. This method has been widely used for building viral spectral libraries but often results in spectra contamination by non-specific host-derived proteins. In addition, the method is laborious as it requires culturing viruses and infecting permissive cell lines.
In chemical lysis protocols, detergents, chaotropic agents or solvents are used to extract viral proteins prior to MALDI-TOF MS. These methods may introduce contaminants that can interfere with the ionization process, and their efficiency can be challenged when analyzing low biomass samples. Solvent-based extraction such as acetone precipitation enables selective concentration of viral particles and their proteins. This method also enables sample cleanup from contaminants such as salt, residual detergent and lipids and supports viral protein enrichment. It is particularly relevant for enveloped viruses.
Our protocol integrates an ultrafiltration procedure for virus purification and concentration using 10 kDa filters followed by further solvent-based sample cleanup using acetone precipitation. Our choice of ultrafiltration as the first step was based on our previously published work demonstrating that polyethersulfone membranes (PES) with a 30 kDa cutoff achieved > 4–5 log₁₀ viral retention (> 99.99%) for MS2 and AcNPV47. We have adapted this step by using 10 kDa PES filters, which are expected to be at least as efficient as 30 kDa in terms of particle-size exclusion. The subsequent acetone extraction and precipitation of viral particles retained on 10 kDa membranes enables further protein enrichment and removal of residual salts, lipids and contaminants, thereby improving ionization efficiency and producing cleaner spectral signals.
This two-step integrated approach for MALDI-TOF viral sample preparation offers several advantages compared to previously reported methods. It is culture-independent providing a pre-analytical concentration of viral particles present in the sample followed by sample clean up through organic solvent precipitation, which is particularly advantageous for analyzing low-biomass clinical samples.
Collectively, our approach grounded in well-established quantitative filtration performance on viral surrogates and combined with efficient solvent-based extraction practices provides a justified rationale for its enhanced efficiency, robustness and simplicity for viral samples analysis with MALDI-TOF. Applying this sample preparation workflow to viral surrogate samples, followed by data acquisition and model training using 5-fold cross validation procedure, we demonstrated perfect viral classification, with clear distinction from bacterial spectra. Although this approach was applied to a limited number of viruses, it has proven effective, simple, and warrants further evaluation on a broader collection of viral samples.
In the current study, ten AI models were evaluated and demonstrated strong, reliable classification performance across a variety of tasks using our internal MALDI-TOF experimental spectra dataset. Several models achieved high performance in both binary and multi-class classification tasks. Notably, seven models consistently delivered perfect classification results for distinguishing bacteria from viruses, identifying Gram type, and reliably classifying the panel of 12 bacterial and viral samples studied into their respective species, thereby highlighting their precision and robustness. Interestingly, these models showed strong ability to rapidly and accurately differentiate bacteria from viral agents which has direct implications for clinical triage, antimicrobial stewardship, and outbreak management, particularly when rapid decision-making is critical. In addition, the models demonstrate high accuracy in distinguishing closely related species, including Bacillus simulants of B. anthracis, when assessed on bacterial spectra generated internally underscoring the potential utility of AI-enhanced MALDI-TOF for rapid identification in biodefense scenarios.
To our knowledge, this is among the first studies to demonstrate the feasibility of integrating MALDI-TOF MS with AI models for the classification of both bacterial and viral pathogens, highlighting the potential for broad-spectrum, agent-agnostic detection. The inclusion of viruses in this study, spanning DNA and RNA both enveloped and non-enveloped types, illustrates the adaptability of our workflow to diverse pathogen types, thereby addressing a critical gap in current MALDI-TOF applications.
The usefulness and validity of an AI model depend on its ability to generalize to unseen data, both from external datasets and real-world scenarios. Although our best-performing models Extra Trees Classifier, Support Vector Classifier, and 1-D CNN achieved near perfect performance across various classification tasks on internal datasets, we observed robust accuracy in classifying bacterial species by their Gram type and distinguishing bacterial from viral spectra. However, performance declined notably when cross-validated against spectra from the RKI MALDI-TOF database. This decline was more pronounced with Support Vector Classifier and to a lesser extent with 1-D CNN and Extra Trees Classifier models suggesting potential overfitting. A major and a common source of misclassification was the consistent confusion of B. cereus and B. subtilis spectra. The persistent misclassification between B. cereus and B. subtilis can be attributed to the fact that the internal dataset comprised three classes belonging to the Bacillus genus, which, in combination with an adequate but not extensive number of spectra for model training, reasonably leads to a non‑negligible rate of misclassification during the evaluation with the external database among species within this genus.
Overall, the performance degradation observed in the cross-dataset evaluation using spectra from the external database occurs because these spectra were generated by different research groups, each employing distinct MALDI-TOF protocols, sample concentrations, and sample preparation procedures. More specifically, trifluoroacetic acid (TFA) was used to inactivate bacteria and prepare samples of the RKI external database, whereas we used direct transfer (DT), extended direct transfer (eDT), and standard protein extraction (PE) recommended by Bruker for producing spectra of our internal dataset. These subtle differences in the extraction protocols may have contributed to the decline of model performance. In addition, these groups applied various instrument calibration standards, further increasing inter-laboratory variability. Consequently, the heterogeneity of the spectra in the external dataset constitutes a critical factor contributing to the reduced classification performance of the models. The observed decrease in performance on the external RKI dataset underscores the challenge of model generalization in MALDI-TOF spectral analysis and highlights the importance of training AI models on heterogeneous and diverse datasets to ensure robustness in real-world applications.
Although MALDI-TOF MS is among the dominant benchmarking methods in the field of microbial identification, the present study is limited by the use of spectra acquired from reference strains under controlled conditions rather than from complex real-world clinical or environmental samples. Sample preparation therefore remains a critical factor, and current workflows still need to move beyond simple direct spotting towards more robust extraction approaches to improve spectral quality in real-world applications. Complex matrices such as blood, urine, or environmental samples can substantially influence MALDI-TOF results by introducing contaminants that suppress analyte ionization, increase background noise, and cause peak interference48. As a result, the model performance reported here may not fully capture the variability encountered under routine diagnostic or field conditions.
A further limitation of this study concerns the relatively small size of the dataset, comprising 255 spectra in total, which constrains our ability to draw definitive conclusions regarding the generalizability of the proposed models and increases the risk of overfitting. This consideration is supported by the notably high accuracy and F1-score values observed during model evaluation, although comparable performance has been reported in the literature for analogous classification tasks. Although k-fold cross-validation was implemented to maximize the use of the limited dataset by allocating spectra to both training and validation subsets, this strategy does not fully eliminate the possibility of optimistic performance estimates or overfitting.
Addressing these limitations will require future studies focused on the acquisition of MALDI-TOF spectra from clinical and environmental samples, the use of larger and more heterogeneous datasets, and the systematic evaluation of preprocessing strategies designed to mitigate matrix effects for bacterial and viral pathogens49,50. This includes additional scenarios such as aerosol sample collection and accurate bacterial spore discrimination51,52, which are particularly relevant for biosecurity applications. Meanwhile, advances in MALDI matrix systems have already demonstrated improved matrix-analyte stoichiometric control53 and, together with the rapid development of AI-based tools and expanding databases, provide a strong foundation for the evolution of MALDI-TOF into a high-precision, high-throughput platform in future real-world deployments.
In conclusion, this study involved culturing and preparing bacterial and viral samples to generate an internal MALDI-TOF MS dataset, followed by rigorous spectral pre-processing. Ten models were trained and evaluated for two binary classification tasks and for multi-class classification of seven bacterial species and five viral species. Three top-performing models were further cross-validated using external MALDI-TOF MS spectra from the Robert Koch Institute database. We identified the Extra Trees Classifier and a 1-D CNN as the best-performing approaches for classifying bacterial and viral classes with strong performance metrics within our internal dataset. When these models were assessed and cross-validated against the external MALDI-TOF MS dataset of highly pathogenic bacteria, they demonstrated acceptable performance.
However, further work is required to train these models on broader and more diverse datasets to enhance feature extraction and improve discrimination among closely related species and isolates. The incorporation of a broader range of species and the increase of inter-laboratory data and protocol diversity in datasets reduce the risk of overfitting internal spectral signatures and assists in capturing intra-species variability. Learning curves are considered a reliable tool to investigate the proportion of additional MALDI-TOF spectra required to further improve the performance of trained models. Further work may also include the development of advanced feature extraction methods and deep learning architectures tuned for fine-grained species and strain discrimination.
In addition, we developed a workflow for virus classification that requires evaluation on a larger panel of viral samples to validate its effectiveness, generalizability, and applicability to clinical settings. Overall, these findings illustrate the potential of combining MALDI-TOF MS with AI to create scalable, rapid, and reliable microbial identification platforms for laboratory-based applications. Such platforms could accelerate diagnostics, enhance biodefense preparedness, and provide an adaptable framework for emerging pathogen detection within controlled laboratory environments.
Methods
The test panel consisted of 12 biological agents including seven bacteria and five viral species (Table 6). The bacterial strains E. coli (DSM 500), Enterobacter cloacae (DSM109592), K. aerogenes (DSM 30053), and K. pneumoniae (HUMB 01336), B. subtilis (BR 151), B. atrophaeus (DSM 675) et B. cereus (DSM 2302) were obtained as lyophilized cultures from the German Collection of Microorganisms and Cell cultures (DSMZ) and the Human Microbiome Project culture collection (HUMB; https://eemb.ut.ee). Adeno-associated virus type 2 (AAV2), Moloney murine leukemia virus (MMLV), Lentivirus derived from Human Immunodeficiency Virus 1 (HIV-1) and baculovirus derived from Autographa californica multiple nucleopolyhedrovirus (AcMNPV) were purchased from Vectorbuilder Inc (Chicago, IL, USA). Bacteriophage MS2 was propagated in the E. coli F + MC-4100/pOX38 host strain using standard cultivation protocol47.
Table 6 Description of microbial test panel of seven bacteria and five viruses used in the current study. The number of experimental spectra in the internal MALDI-TOF dataset used for training ML/DL models is provided.Microbial culture and sample preparation
Lyophilized bacterial strains were rehydrated in 500 µL of tryptone soy broth medium (Neogen, USA). A 100 µL aliquot of each suspension was plated on tryptone soy agar (TSA) plates (Neogen, USA) and incubated overnight at 37 °C in 5% CO2 atmosphere.
The following day, seven isolated colonies from each TSA plate were selected for protein extraction. Samples were processed using three methods for MALDI-TOF MS analysis recorded by Bruker Microflex MALDI-TOF Mass Spectrometer (Bruker Daltonik, Bremen, Germany): direct transfer (DT), extended direct transfer (eDT) and standard protein extraction (PE), following the manufacturer’s instructions provided in the MALDI Biotyper® protocol (Bruker Daltonics, Billerica, MA, USA). Briefly, in the direct transfer (DT) method, an isolated colony was smeared onto the target MALDI-TOF plate and overlaid with 1 µL of α-cyano-4-hydroxycinnamic acid (HCCA) matrix. In the extended direct transfer (eDT) method, colonies were pretreated with 70% formic acid before HCCA application. The protein extraction was performed according to the Bruker Biotyper® protocol, with mild modifications. Briefly, a sterile 1 µL inoculation loop was used to transfer isolated colonies into 100 µL of HPLC-grade water and mixed thoroughly. Using a pipette, 300 µL of pure ethanol was added to the suspension. After thoroughly mixing, the specimen was centrifuged for two minutes at 14,000 rpm. The supernatant was removed, and the centrifugation step was repeated. Residual ethanol was removed by pipetting. After allowing the pellet to dry at room temperature for a minimum of 5 min, 5 µL of 70% aqueous formic acid was added to re-suspend the pellet, followed by 5 µL of acetonitrile, which was mixed by pipetting. The specimen was then centrifuged for two minutes at 14,000 rpm and 1 µL of supernatant was loaded onto the target plate and overlaid with 1 µL of α-cyano-4-hydroxycinnamic acid (HCCA) matrix.
Viral samples were prepared using centrifugal filter units with a 10 kDa cut-off (Vivaspin® 500, Sartorius, Gottingen, Germany). Briefly, 200 µL of each viral suspension were filtered through the units by centrifugation at 12,000 x g for 10 min. The filtrate was discarded, and the retained viral particles on the filter membrane were washed twice with 400 µL of Milli-Q water, followed by centrifugation under the same conditions. The viral particles were then recovered in 50 µL of Milli-Q water and extracted by thorough mixing with 100 µL of acetone47. The samples were heat-inactivated at 70 °C for 10 min and stored at −20 °C until further use. For the MALDI-TOF MS analysis, extracted viral samples were briefly centrifuged and 1 µL of the extract was applied on the target plate, dried at room temperature and overlaid with 1 µL of HCCA matrix.
MALDI generates singly charged ions through pulsed-laser irradiation of the analyte. These ions are subsequently separated into a TOF mass spectrometer based on their velocities prior to detection. The mass-to-charge (m/z) ratios of the ions are determined by measuring the time required for each ion to traverse the flight tube. Figure 3 presents a representative raw mass spectrum obtained from the analysis of E. coli. In this plot, the x-axis corresponds to the m/z values of the detected ions, while the y-axis indicates their respective signal intensities. The m/z range spans from 0 to 20,000 Daltons (Da).
Spectra pre-processing
Raw MALDI-TOF mass spectra were subjected to a standardized pre-processing pipeline using the MicrobeMS software, a MATLAB tool specifically designed by Peter Lasch at the Robert Koch Institute for the analysis of MALDI-TOF mass spectra of microbial samples54 as raw data typically contain high levels of noise, baseline drift, and experimental variability that can obscure biologically relevant information and hinder accurate interpretation. The pipeline included baseline subtraction, smoothing, normalization, and spectral cutting. Baseline subtraction was performed using the asymmetric least squares (AsLS) algorithm, which constructs a baseline correction curve through shape-preserving piecewise cubic interpolation over user-defined intervals, effectively removing background noise. Smoothing was applied using the Savitzky-Golay filter to reduce high-frequency noise while preserving peak shape, enhancing signal clarity for downstream analysis. Spectra were then normalized using a modified 1-norm algorithm that computes a scaling factor from intensity differences across selected m/z bins, thereby reducing intensity variation across samples.
Afterwards, spectra were cut to a defined m/z range between 2,000 and 12,000 to remove irrelevant or low-informative regions and reduce memory requirements. All m/z values between 0 and 2,000 Da and the respective intensities were excluded since this region can be crowded with noise and matrix-related peaks which can interfere with the analysis and obscure the signals of interest. Finally, to ensure uniform dimensionality across spectra, intensity measurements were binned based on bin size of 2 Da, resulting in a vector containing 5,000 features that reflects the bins the m/z axis was partitioned. An example of the pre-processed spectrum of E. coli is depicted in Fig. 3. All hyperparameters for the pre-processing functions were set to the default values provided by the MicrobeMS software. This pre-processing pipeline ensured consistent spectral quality and improved the reliability of subsequent classification methods in our dataset.
Fig. 3
Representative MALDI-TOF spectrum from E. coli before (left) and after preprocessing (right).
Classification methods
In this study, we evaluated the performance of various ML and DL models. The ML models included Random Forest, Linear SVC, Ridge Classifier, k-Nearest Neighbors (kNN), Extra Trees Classifier, Support Vector Classifier, Logistic Regression, and XGBoost. Due to their effectiveness in capturing informative patterns within small datasets, we selected these eight ML models for detailed comparison. In addition, two DL baseline methods were also evaluated: a 1-D Convolutional Neural Network (1-D CNN) and a Denoising Autoencoder followed by a Ridge Classifier (DAE-RC). The 1-D CNN was chosen for its ability to extract spatial features from one-dimensional signals, while the DAE-RC was included for its capacity to denoise and reconstruct input data, compressing it into a compact and informative latent representation. Additional details on the models, concise descriptions and representative studies, as well as their respective advantages are summarized in Table 7. The hyperparameters selected for the ML approaches, such as Extra Trees Classifier, SVM, Logistic Regression, Random Forest, kNN, Linear SVC, Ridge Classifier and XGBoost, followed the default recommendations of their respective libraries. In contrast, the DL approaches employed more complex architecture and were fine-tuned prior to the training process. More specifically, the architecture of the CNN-1D model consisted of an initial 1-D convolutional layer with 64 filters and a kernel size of 3 to extract local temporal features from the input sequences, followed by a 1-D max pooling layer with a pool size of 2 to down sample the feature maps and reduce computational complexity. The resulting feature maps were then flattened into a multidimensional vector. Two fully connected Dense layers followed the convolutional block: the first contained 100 units with ReLU activation to capture non-linear relationships, while the final output layer consisted of 10 units, corresponding to 10 classes, with softmax activation for multi-class probability distribution. The model was compiled using categorical cross-entropy loss and the Adam optimizer, with accuracy as the evaluation metric. Regarding the Denoising Autoencoder, it was implemented to reconstruct clean data from noisy inputs while performing dimensionality reduction. The architecture consisted of an input layer accepting 5,000 flattened features, followed by a Gaussian noise layer that corrupted inputs during training to force the model to learn robust features. The encoder phase compressed the data through a dense layer with 256 units with ReLU activation, batch normalization, and 50% dropout regularization to prevent overfitting. A second dense layer further reduced dimensionality to 128 units, forming the bottleneck representation of the decoder phase. Batch normalization was applied to stabilize training and improve gradient flow throughout the network. The model was trained using Mean Squared Error loss to reconstruct original clean data from corrupted inputs, leveraging the combination of Gaussian noise injection and dropout as regularization mechanisms to enhance feature robustness and generalization.
Table 7 Detailed descriptions of types, characteristics and representative studies of the ten ML/DL models evaluated in this study. In the “Training Time” column, plus signs indicate the relative time required to train each model, whereas in the “Interpretability” column, asterisks represent the ease of interpreting the model outputs.Evaluation protocol
Given the fact that each class of the dataset is featured with no more than 24 spectra, a sophisticated method for model evaluation was implemented using 5-fold cross-validation to ensure robust generalization of the relationship between input spectra and target classes. Moreover, 5-fold cross validation substantially mitigates the risk of overfitting by providing a more robust and unbiased estimate of model generalization performance through systematic data partitioning and evaluation. This was achieved by employing a stratified 5-fold split, as illustrated in Fig. 4 (A), wherein 20% of the dataset was allocated for validation in each iteration while maintaining the class distribution. For each experiment and model, training was performed on four folds and validation on the remaining fold, with this process iteratively applied across all five folds.
Classification performance was assessed using the average and standard deviation of accuracy (Eq. 1) and F1-score (Eq. 4) across five iterations, where True Positives (TP), True Negatives (TN), False Positives (FP), and False Negatives (FN) are exploited for the abovementioned evaluation metrics. Regarding accuracy, it represents the proportion of correct predictions (both true positives and true negatives) among the total number of cases examined, whereas F1-score, being the harmonic mean of Precision and Recall, expresses a single metric that reflects how well a classification model identifies positive cases while minimizing both false positives and false negatives. The model that achieves the highest average accuracy with the smallest standard deviation was considered the best-performing model for a given task.
$$\:Accuracy=\frac{TP+TN}{TP+TN+FP+FN}$$
(1)
$$\:Precision=\frac{TP}{TP+FP}$$
(2)
$$\:Precision=\frac{TP}{TP+FP}$$
(3)
$$\:F1-score\:=\:2\:\cdot\:\frac{Precision\cdot\:Recall}{Precision+Recall}$$
(4)
Fig. 4
Overview of evaluation protocol applied in ML/DL models using (A) five-fold cross validation for the internal dataset and (B) validation on the total of the external dataset.
Model cross validation methodology
To provide a more realistic assessment of the model’s ability to generalize to new unseen data we validated our ML/DL models with MALDI-TOF mass spectra from an external database54. More specifically, spectra from the MALDI-TOF mass spectrometry database of the Robert Koch Institute were selected for further assessment of trained models. This database is freely accessible and can be used without technical restrictions. Concerning the protocols adopted for sample preparation in this database, authors at RKI used TFA as an inactivation protocol for microbial sample preparation, contrary to our work where DT, eDT, and PE were selected for sample preparation. In general, this dataset contains 11,055 spectra from altogether 1,601 bacterial strains and 264 species and is primarily intended to improve the identification and classification of highly pathogenic bacteria. For our experiments, we selected 285 MALDI-TOF spectra for 7 bacteria, 132 for Gram positive bacteria and 153 for Gram negative bacteria, respectively as shown in Table 8. Within the Bacillus genus, 67, 35 and 30 mass spectra were selected for B. cereus, B. atropaheus and B. subtilis, respectively: 62 mass spectra for E. coli, 34 for E. cloacae and 31 for K. pneumoniae. For the second species of the Klebsiella genus, K. aerogenes, 26 mass spectra were chosen. The information about the bacteria genus, species, strains and id within the database (MicrobeMS ID) of the selected spectra are provided in supplemental table S1 (Table S1). To better explore the selected spectra, more details about the bacteria genus, species, strains and id within the database (MicrobeMS ID) are provided in the supplementary file. The information provided can direct readers to the metadata of the database where description about the NCBI tax id’s, growth time & temperature, growth medium & air, concentration, sample treatment, calibration standard and the measurement method are listed per spectrum. Finally, to perform the cross-dataset validation, we evaluated the performance of three trained models that achieved the highest accuracy and F1-score by using the same evaluation metrics after applying the same pre-processing pipeline as applied to the internal dataset. As presented in Fig. 4 (B), the selected models for cross-dataset evaluation were re-trained to the entire bacterial internal dataset prior to their validation to take advantage of all available mass spectra.
Table 8 Number of MALDI-TOF mass spectra used from the external dataset to cross-validate trained ML/DL models.Software requirements
To conduct the experiments in this study, a laptop with an AMD Ryzen 5 4600 cpu (6 cores, 12 threads) was used with 16 GB of RAM. Initially, spectral preprocessing was performed using the MicrobeMS software package (Microbe MS 0.90d)54. Afterwards, Python (3.10) was used for the subsequent stages of our work employing the Pandas (1.4.4) and Pyteomics (5.0) packages for spectra import, scikit-learn (1.1.1) for ML models, XGboost (1.7.3) for gradient boosting, and PyTorch (2.0) for DL models.