Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Circuit-breakers are a novel approach to prevent harmful outputs from AI systems by remapping model representations tied to harmful processes into incoherent or refusal states through Representation Rerouting. The method aims to robustly block harmful generation while preserving general capabilities and generalizes across unseen adversarial attacks in text- and multimodal models. Empirical results with Llama-3-8B-Instruct (+ RR) show a large reduction in harmful outputs, and multimodal systems like LLaVA-NeXT-Mistral-7B (+ RR) maintain performance under attacks while improving robustness.
Parse Score
Sources
circuit-breaker.ai shapes more of what AI says about Circuit Breaker than any other source, at 100% of its citations.