What is it about?
Artificial intelligence is increasingly used to detect malicious software. Explainable AI can make these systems more transparent by showing why a file has been classified as malicious or benign. But what happens if an attacker can fool both the prediction and its explanation? This paper introduces GAME4EXE, an adversarial AI approach designed to investigate this risk in Windows malware detection. The method explores whether malicious executable files can be modified so that a deep-learning malware detector classifies them as benign while, at the same time, the explanation of that decision also resembles the explanation normally associated with benign software. In this way, the study considers a more challenging threat than conventional evasion: malware may not only escape detection, but may also produce misleading explanations that make the incorrect decision appear more credible.
Featured Image
Photo by Julio Lopez on Unsplash
Why is it important?
Explainable AI is often introduced to increase confidence in complex AI systems. In cybersecurity, however, explanations themselves can become targets of adversarial manipulation. If malicious software can evade a detector while also receiving an apparently reassuring explanation, security analysts may have fewer warning signs that the AI system has been deceived. The results of this preliminary study show that prediction and explanation can be attacked together in Windows malware detection. They therefore highlight the need to evaluate the security not only of AI predictions, but also of the explanation mechanisms used to make those predictions understandable. Understanding these vulnerabilities is an important step towards developing more robust malware detectors and more trustworthy explainable AI systems.
Perspectives
Our motivation is defensive: before we can protect AI-based cybersecurity systems against adversarial manipulation, we need to understand how they can fail. With GAME4EXE, we investigate a particularly challenging scenario in which an adversarial malware sample attempts to deceive both a deep-learning detector and the mechanism used to explain its decision. The study shows why explainability alone should not automatically be considered a guarantee of trustworthiness. Explanations must themselves be tested for robustness under adversarial conditions. Our broader goal is to contribute to AI systems that are not only accurate and explainable, but also secure against intentional manipulation. Key takeaways - Deep-learning malware detectors can be vulnerable to adversarial manipulation. - Explainable AI mechanisms may themselves become targets of attacks. - GAME4EXE investigates attacks that simultaneously target prediction and explanation. - The study considers Windows Portable Executable malware and two deep-learning detection models. - The evaluation examines prediction evasion, transferability and the ability to produce misleading, goodware-like explanations. - The findings reinforce the need to assess the robustness of both AI decisions and their explanations. - The research is intended to expose vulnerabilities so that stronger defensive countermeasures can be developed. Publication details Luca Lobascio, Giuseppina Andresini, Annalisa Appice and Donato Malerba Adversarial Malware Can Be Both Evasive and Deceiving: a Gradient-based Attack Against Prediction and Explainability in Windows PE Malware Detection 2026 IEEE 11th European Symposium on Security and Privacy Workshops — EuroS&PW 2026 IEEE, 2026 DOI: 10.1109/EuroSPW72509.2026.00038
Prof. Donato Malerba
Universita degli Studi di Bari Aldo Moro
Read the Original
This page is a summary of: Adversarial Malware Can Be Both Evasive and Deceiving: a Gradient-based Attack Against Prediction and Explainability in Windows PE Malware Detection, July 2026, Institute of Electrical & Electronics Engineers (IEEE),
DOI: 10.1109/eurospw72509.2026.00038.
You can read the full text:
Contributors
The following have contributed to this page







