FALCON: Fine-Grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language Model
Published in Advances in Neural Information Processing Systems (NeurIPS) 2025, 2025
Large language models have been widely applied but can inadvertently encode sensitive or harmful information, creating significant safety challenges. We propose FALCON, a machine unlearning approach that goes beyond coarse-grained loss combinations. FALCON employs three key mechanisms: information-theoretic guidance for selecting which parameters to modify, contrastive mechanisms to better separate learned representations, and orthogonal projection of conflicting gradients to balance forgetting objectives with model utility. Experiments demonstrate that FALCON achieves effective knowledge removal while preserving overall model performance and resisting knowledge recovery attempts.
