What is it about?
Sarcasm is common on social media, where a post might pair an image with a caption. For AI, detecting this sarcasm is difficult because the real meaning is often the opposite of the literal words. A key clue is the mismatch between the image and the text, especially in their emotional tone. Our research introduces a new AI model called EilMoB, designed specifically to catch this type of multi-modal sarcasm. It uses an advanced Multi-modal Large Language Model (MLLM) to analyze both the image and text from a social media post. The model is prompted to perform three steps: describe the image, analyze the emotion of the post, and check if the emotions of the image and text align. This process generates a description of the "Emotion-aware Incongruity" (EAI). By converting the emotional conflict into text, our model effectively bridges the gap between visual and textual data. It then uses unique modules to mine for incongruity, weigh the importance of different features, and combine all the information to accurately decide if a post is sarcastic
Featured Image
Why is it important?
This work is important because it tackles a major weakness in current AI systems: understanding nuanced human communication. Existing models often fail to detect sarcasm because they can't effectively process the conflicting emotional signals between images and text. Our EilMoB model significantly improves accuracy, outperforming the best previous methods on benchmark datasets for sarcasm detection. This breakthrough has practical applications in areas like: Sentiment Analysis: Businesses can get a more accurate read on customer feedback, understanding what people truly mean despite sarcastic language. Content Moderation: It can help identify subtle forms of online harassment or bullying that are masked by sarcasm. Human-AI Interaction: It helps build AI assistants that can better understand and respond to the subtleties of human expression. By successfully using MLLMs to extract and analyze emotional incongruity, our research provides a novel framework that enhances AI's ability to interpret complex, multi-modal messages.
Perspectives
This research represents a step forward in creating more contextually and emotionally intelligent AI. Figurative language like sarcasm has long been a challenge, and our model's ability to understand it by detecting cross-modal emotional conflict is a significant advance. However, the model is not perfect. Our error analysis shows it can sometimes be misled by very subtle cues or when the sarcasm relies on complex world knowledge not immediately present in the image and text. For example, it might focus too much on a single expressed emotion and miss the broader ironic context. Looking ahead, we believe the core method—using an MLLM to generate a textual description of multi-modal incongruity—is a powerful technique. It can be adapted for other challenging tasks beyond sarcasm, such as detecting fake news, analyzing video content, or identifying propaganda where visual and textual elements might be intentionally mismatched.
Haochen Zhao
Read the Original
This page is a summary of: EilMoB: Emotion-aware Incongruity Learning and Modality Bridging Network for Multi-modal Sarcasm Detection, June 2025, ACM (Association for Computing Machinery),
DOI: 10.1145/3731715.3733321.
You can read the full text:
Contributors
The following have contributed to this page







