Reliability of neural networks as diagnostic tools for information extraction

dc.contributor.advisorWand, Michael
dc.contributor.advisorSpichtinger, Peter
dc.contributor.authorWörl, Ann-Christin
dc.date.accessioned2026-07-29T06:40:44Z
dc.date.issued2026
dc.description.abstractDeep neural networks have become powerful tools for modeling complex systems in many scientific fields. Beyond their predictive capabilities, they enable the extraction of meaningful information from high-dimensional data and thus support scientific hypothesis generation. In this context, two complementary approaches can be distinguished. First, interpretability methods such as saliency maps attribute model predictions to input features and identify which variables are relevant for the task. Second, information can be probed implicitly by systematically restricting the input space. If predictive performance remains stable under such constraints, this indicates that the underlying information is robustly represented in the data. This thesis investigates under which conditions neural networks can be used as reliable diagnostic tools for information extraction. We analyze the robustness of gradient-based saliency methods with respect to random model initialization, which can be identified as a significant source of variability in attribution maps, even when predictive performance remains unchanged. To quantify the variability, we introduce a signal to noise ratio framework that measures the stability of explanations across independently trained models and propose marginalization over the initialization distribution as a practical strategy to reduce attribution variability. As an application, we study the ability of deep neural networks to reconstruct cloud structures based on atmospheric state variables and analyze how cloud-relevant information is encoded in the data. Our results show that this information is distributed in a redundant and complementary manner across physical variables and remains partially preserved under reduced input representations and generalization across regions, time shifts, and data sources. Applying the robustness framework in this setting further reveals that saliency patterns for the most relevant variables are considerably more stable than in classification tasks, enabling more reliable identification of physically meaningful structures. Overall, this work clarifies under which conditions neural networks can serve as reliable tools for scientific information extraction. It shows that interpretability is inherently non-deterministic and must be evaluated in terms of robustness and uncertainty. Beyond attribution-based methods, it demonstrates that systematically probing the input space enables complementary and often more stable insights into how relevant information is structured within the data. Consequently, insights derived from data-driven models should not be interpreted as causal, but as evidence of structured information in the data that requires careful validation and uncertainty assessment.en
dc.description.abstractTiefe neuronale Netze haben sich als leistungsstarkes Werkzeug zur Modellierung komplexer Systeme in vielen wissenschaftlichen Bereichen etabliert. Neben ihrer Vorhersagefähigkeit ermöglichen sie die Extraktion aussagekräftiger Informationen aus komplexen Daten und können so die Hypothesenentwicklung unterstützen. Dabei lassen sich zwei Ansätze unterscheiden: Zum einen untersuchen Methoden der Interpretierbarkeit den Zusammenhang zwischen Modellvorhersage und Eingabemerkmalen, etwa durch Salienzkarten. Zum anderen können Informationen implizit analysiert werden, indem der Eingaberaum systematisch eingeschränkt und die Güte der Vorhersage bewertet wird. Bleibt diese unter solchen Einschränkungen erhalten, deutet dies auf robust in den Daten enthaltene Informationen hin. In dieser Arbeit untersuchen wir, unter welchen Bedingungen neuronale Netze als zuverlässige Diagnosewerkzeuge zur Informationsextraktion genutzt werden können. Dazu analysieren wir die Robustheit gradientbasierter Salienzmethoden gegenüber zufälliger Modellinitialisierung, die sich als wesentliche Quelle für Variabilität herausstellt -- selbst bei gleicher Vorhersagegüte. Zur Quantifizierung nutzen wir das Signal-Rausch-Verhältnis über unabhängig trainierte Modelle hinweg und schlagen die Marginalisierung über die Initialisierungsverteilung als Strategie zur Reduzierung der Attributionsvariabilität vor. Als Anwendungsfall dient die Rekonstruktion von Wolkenstrukturen aus atmosphärischen Zustandsvariablen mithilfe von neuronalen Netzen. Dies ermöglicht eine Analyse, wie Informationen über Wolkenbildung in den Daten kodiert sind. Die Ergebnisse zeigen eine redundante und komplementäre Verteilung dieser Information über verschiedene physikalische Variablen. Selbst reduzierte Eingaben oder eine Generalisierung über Regionen und Datenquellen beeinträchtigen die Vorhersagequalität nur geringfügig. Zudem zeigt sich, dass Salienzmuster in diesem Regressionssetting deutlich stabiler sind als in Klassifikationsaufgaben, wodurch eine zuverlässigere Interpretation physikalischer Zusammenhänge möglich wird. Insgesamt zeigt diese Arbeit, unter welchen Bedingungen neuronale Netze als zuverlässige Werkzeuge für die wissenschaftliche Informationsextraktion dienen können. Sie macht deutlich, dass Interpretierbarkeit nicht deterministisch ist und hinsichtlich Robustheit und Unsicherheit bewertet werden muss. Über attributionsbasierte Methoden hinaus liefert die systematische Untersuchung des Eingaberaums ergänzende und oft stabilere Einblicke. Erkenntnisse aus datengesteuerten Modellen sollten nicht kausal interpretiert werden, sondern als Hinweise auf strukturierte Informationen, die sorgfältiger Validierung und Unsicherheitsbewertung bedürfen.de
dc.identifier.doihttps://doi.org/10.25358/openscience-15770
dc.identifier.urihttps://openscience.ub.uni-mainz.de/handle/20.500.12030/15791
dc.identifier.urnurn:nbn:de:hebis:77-b16450d6-4fcd-4e83-932c-e7716a2c716d9
dc.language.isoeng
dc.rightsCC-BY-SA-4.0
dc.rights.urihttps://creativecommons.org/licenses/by-sa/4.0/
dc.subject.ddc004 Informatikde
dc.subject.ddc004 Data processingen
dc.titleReliability of neural networks as diagnostic tools for information extractionen_US
dc.typeDissertation
jgu.date.accepted2026-07-02
jgu.description.extentxi, 125 Seiten ; Illustrationen, Diagramme
jgu.identifier.uuidb16450d6-4fcd-4e83-932c-e7716a2c716d
jgu.organisation.departmentFB 08 Physik, Mathematik u. Informatik
jgu.organisation.nameJohannes Gutenberg-Universität Mainz
jgu.organisation.number7940
jgu.organisation.placeMainz
jgu.organisation.rorhttps://ror.org/023b0x485
jgu.rights.accessrightsopenAccess
jgu.subject.ddccode004
jgu.type.dinitypePhDThesisen_GB
jgu.type.resourceText
jgu.type.versionOriginal work

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
reliability_of_neural_network-20260729084044920656.pdf
Size:
15.49 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
5.14 KB
Format:
Item-specific license agreed upon to submission
Description: