TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models
Published in arXiv preprint arXiv:2503.10872, 2025
Vision Language Models demonstrate strong inference abilities but face vulnerabilities to jailbreak attacks. We propose TAIJI, a black-box defense framework that employs key phrase-based textual anchoring to enhance the model’s ability to assess and mitigate harmful content embedded within both visual and textual prompts. Unlike existing methods that require model access or multiple queries, TAIJI operates with a single inference query while maintaining performance on legitimate tasks. Extensive evaluations show that TAIJI substantially improves VLM safety and reliability for real-world deployment.
