
Yuval Elovici
Using image conversion techniques to detect adversarial examples in machine learning models for tabular data
Most adversarial detection methods are designed for deep neural networks (DNNs) trained on homogeneous data (e.g., images), which leaves machine learning (ML) models for tabular data particularly vulnerable to adversarial attacks. This paper introduces two novel approaches for the detection of adversarial examples in tabular data that transform tabular data into image representations, enabling the use of convolutional neural networks (CNNs) to improve adversarial robustness. The first is an unsupervised framework that exploits explainability techniques to compare model behavior using two distinct data representations: the original tabular form and its transformed image-based counterpart. By analyzing discrepancies in feature attributions and model predictions, this approach detects adversarial examples in a model- and attack-agnostic manner. The second is a supervised method utilizing transfer learning which leverages a large-scale pretrained image classification model to extract adversarial patterns from converted tabular data, significantly enhancing detection accuracy with minimal training overhead. The evaluation of these approaches, which assessed their ability to detect multiple types of attacks, was performed on seven structured datasets. Both approaches were shown to outperform five state-of-the-art adversarial detection methods, achieving near-perfect detection rates in most cases. This work pioneers the integration of multi-view learning, explainable AI (XAI), and transfer learning for adversarial detection in tabular data, providing a novel and highly effective paradigm for securing ML models in real-world applications.
| Publication language | English |
| Journal | Applied Soft Computing |
| Volume | 181 |
| Publication status | Published - 01.09.2025 |
| 113288 |