Tiny Refinements Elicit Resilience: Toward Efficient Prefix-Model Against LLM Red-Teaming
Published in arXiv preprint arXiv:2405.12604, 2024
We present a plug-and-play prefix module that reconstructs input prompts using fewer than 30 additional tokens to mitigate toxic outputs from large language models. The sentinel model addresses parameter inefficiency and limited model accessibility for fine-tuning large target models. We employ interleaved training using Proximal Policy Optimization to jointly optimize both red team and sentinel models, incorporating a value head-sharing mechanism inspired by multi-agent centralized critic approaches. Testing across text-to-text and text-to-image applications demonstrates effectiveness against larger models including Llama-2, GPT-3.5, and Stable Diffusion, positioning the framework as a practical safety enhancement for a wide range of applications.
