GenCyberSynth: A Reproducible Comparative Benchmark for Synthetic Malware Images
Submitted conference manuscript
Research summary
Synthetic malware images are widely used to address data scarcity and class imbalance in malware classification. However, existing studies are often conducted with different datasets, preprocessing pipelines, synthetic sample sizes, and evaluation protocols, making it difficult to fairly compare the results of different generation methods. To address this issue, this paper presents GenCyberSynth, a unified and reproducible evaluation framework for synthetic malware images. The framework follows a standardized train–synth–eval workflow, in which each generator is trained only on the real training split, synthetic samples are generated under a fixed per-class budget, and their practical utility is assessed through a unified downstream classification
pipeline. GenCyberSynth further fixes data partitions, adopts a leakage-safe evaluation process, includes real-only baselines, controls random seeds, and organizes the configuration and output of each run into structured summaries. We validate the proposed framework on two malware-image datasets, USTC-TFC2016 and CICMalDroid2020, and compare seven mainstream generator families under the same experimental protocol. Experimental results show that the effect of synthetic augmentation is strongly dataset-dependent, and that proxy generative metrics do not always align with downstream classification utility. These findings indicate that synthetic malware image generation methods should be evaluated under unified and controlled downstream tasks rather than relying only on generative quality metrics. GenCyberSynth provides a benchmark foundation for future studies on synthetic-data scaling, robustness analysis, and synthetic sample selection strategies.
Bruno’s contribution
First-authored development of the GenCyberSynth Phase-1 benchmark and its utility-first evaluation methodology. The work establishes a repaired, reproducible, seed-aware experimental protocol for comparing seven synthetic-data generator families across two malware-image datasets using matched preprocessing, fixed budgets, real-only baselines, held-out real-test evaluation, structured artifacts, and consistent metric aggregation.
Cite this work
B. Fonkeng, R. Khun, J. Zhang, and F. Li, “GenCyberSynth Phase-1: A Repaired, Reproducible, Seed-Aware Comparative Benchmark for Synthetic Malware Images Across Two Datasets,” submitted conference manuscript, 2026.