AppendixAppendix A: Additional performance results for augmentation strategies

Appendix A provides additional results for different augmentation levels, including 5%, 10%, and 15% of the DB1 training data, complementing the main experiments conducted at 20%. These results enable a more detailed examination of how the magnitude of augmentation affects model behavior under limited-data conditions.

Consistent with the observations in the main text, generative augmentation demonstrates relatively stable performance across different augmentation levels, while standard augmentation exhibits greater variability, particularly at intermediate levels. These results further support the conclusion that augmentation strategies that preserve structural characteristics lead to more consistent performance under data constraints (Table 3).

Table 3 Performance comparison of standard and generative augmentation strategies across different augmentation levels (5%, 10%, 15%, and 20%) on the DB1 test set. Metrics include accuracy, AUC, sensitivity, specificity, F1-score, and EER.

Appendix B: Additional results under domain shift and few-shot adaptation for mixed strategy

Appendix B presents detailed results for the mixed augmentation setting, where generative and standard augmentations are applied in equal proportions (10% each) to the DB1 training data. The model is then fine-tuned using varying proportions of DB2 training data to evaluate performance under domain shift.

The results show that the mixed strategy achieves competitive performance across all data regimes and performs particularly well at higher data levels. These findings suggest that combining augmentation strategies may introduce complementary variations that enhance representation learning. However, as only a single mixing ratio is considered, these results should be interpreted as preliminary observations rather than definitive conclusions (Table 4).

Table 4 Performance of the mixed augmentation strategy (generative 10% + standard 10%) under domain shift across different proportions of DB2 training data used for fine-tuning.

Appendix C: Extended results for augmentation type and few-shot adaptation

Appendix C provides extended results for additional model-data combinations not included in the main text. Performance is reported across different augmentation types and augmented ratio (5%, 10%, 15%, and 20%) under both within-domain (DB1) and cross-domain (DB2) settings, and the proportion of the DB2-train set. The results summarize AUC and accuracy, supporting the trends observed in the main analysis (Fig. 9, Table 5).

Table 5 AUC and accuracy values for generative and standard augmentation strategies across different augmentation levels and proportions of target-domain data used for fine-tuning.
Fig. 9

Fig. 9The alternative text for this image may have been generated using AI.

Performance under domain shift for different augmentation levels (5%, 10%, and 15%). Each subplot shows AUC (solid line) and accuracy (dashed line) as a function of the proportion of DB2 training data used for fine-tuning. The results illustrate consistent trends across augmentation levels, with generative augmentation showing more stable performance than standard augmentation.