Interpreting Safety: A LLM and STPA Approach

Published in Pacific Rim International Conference on Artificial Intelligence (PRICAI) 2025, 2025

We combine large language model reasoning with Systems-Theoretic Process Analysis (STPA) to improve interpretability and structured safety assessment in AI systems. By leveraging LLMs to automate and augment the identification of unsafe control actions and loss scenarios defined by STPA, our approach produces interpretable safety analyses that bridge formal safety engineering and modern machine learning. Evaluations demonstrate that the LLM-guided STPA workflow surfaces hazards more comprehensively than manual analysis while remaining accessible to practitioners without deep safety engineering expertise.