Abstract
Explainability is critical for ensuring safe and accountable control in deep reinforcement learning (DRL). Most explanation approaches train explanation models (e.g., mask networks) offline from system trajectories, rendering the training disconnected from environmental feedback on perturbed states. This disconnection frequently leads to feature importance misalignment, i.e., critical features are overlooked, and irrelevant ones are incorrectly highlighted. To address this problem, we propose MaskCtrl, a novel DRL framework for training mask networks as self-explainable controllers. A trained self-explainable controller achieves two key capabilities: (1) decision-making performance comparable to the original neural controller from which it is derived, and (2) accurate identification of state features critical to the decision-making. These dual strengths stem from our DRL-based training mechanism, which leverages system rewards as environmental feedback to correct feature misalignment during training. Empirical results demonstrate that our self-explainable controllers outperform offline-trained explanation models by up to 50.58% in critical feature identification fidelity. Specifically, this enhanced feature alignment between explanations and decision-making enables 25.2% greater effectiveness in adversarial attacks and a more pronounced robustness gain, with only an 8.7% drop in reward against significant adversarial perturbations.
| Original language | English |
|---|---|
| Article number | 109107 |
| Journal | Neural Networks |
| Volume | 203 |
| DOIs | |
| State | Published - Nov 2026 |
Keywords
- Critical features
- Deep reinforcement learning
- Explainability
Fingerprint
Dive into the research topics of 'MaskCtrl: Training mask networks as self-explainable and performant controllers via deep reinforcement learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver