Reliability of neural networks as diagnostic tools for information extraction

Loading...
Thumbnail Image

Date issued

Editors

Journal Title

Journal ISSN

Volume Title

Publisher

Reuse License

Description of rights: CC-BY-SA-4.0
Item type: Item , DissertationAccess status: Open Access ,

Abstract

Deep neural networks have become powerful tools for modeling complex systems in many scientific fields. Beyond their predictive capabilities, they enable the extraction of meaningful information from high-dimensional data and thus support scientific hypothesis generation. In this context, two complementary approaches can be distinguished. First, interpretability methods such as saliency maps attribute model predictions to input features and identify which variables are relevant for the task. Second, information can be probed implicitly by systematically restricting the input space. If predictive performance remains stable under such constraints, this indicates that the underlying information is robustly represented in the data. This thesis investigates under which conditions neural networks can be used as reliable diagnostic tools for information extraction. We analyze the robustness of gradient-based saliency methods with respect to random model initialization, which can be identified as a significant source of variability in attribution maps, even when predictive performance remains unchanged. To quantify the variability, we introduce a signal to noise ratio framework that measures the stability of explanations across independently trained models and propose marginalization over the initialization distribution as a practical strategy to reduce attribution variability. As an application, we study the ability of deep neural networks to reconstruct cloud structures based on atmospheric state variables and analyze how cloud-relevant information is encoded in the data. Our results show that this information is distributed in a redundant and complementary manner across physical variables and remains partially preserved under reduced input representations and generalization across regions, time shifts, and data sources. Applying the robustness framework in this setting further reveals that saliency patterns for the most relevant variables are considerably more stable than in classification tasks, enabling more reliable identification of physically meaningful structures. Overall, this work clarifies under which conditions neural networks can serve as reliable tools for scientific information extraction. It shows that interpretability is inherently non-deterministic and must be evaluated in terms of robustness and uncertainty. Beyond attribution-based methods, it demonstrates that systematically probing the input space enables complementary and often more stable insights into how relevant information is structured within the data. Consequently, insights derived from data-driven models should not be interpreted as causal, but as evidence of structured information in the data that requires careful validation and uncertainty assessment.

Description

Keywords

Citation

Relationships

Endorsement

Review

Supplemented By

Referenced By