Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization
Researchers identify a stylistic inconsistency in MLLMs where visual comprehension remains robust while safety defenses fail under specific stylistic triggers.
The study introduces an adversarial optimization method using GRPO to exploit the gap between an MLLM's ability to process visual content and its safety alignment. By applying stylistic triggers, attackers can bypass safety filters that are otherwise effective against standard content-based jailbreaks.