<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-18T23:11:35Z</responseDate><request verb="GetRecord" identifier="oai:helda.helsinki.fi:10138/601034" metadataPrefix="dim">https://helda.helsinki.fi/server/oai/request</request><GetRecord><record><header><identifier>oai:helda.helsinki.fi:10138/601034</identifier><datestamp>2026-07-23T15:16:04Z</datestamp><setSpec>com_10138_18086</setSpec><setSpec>com_10138_17738</setSpec><setSpec>col_10138_18093</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="author">Subramanian, Janani</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="organization" lang="fi">Helsingin yliopisto, Matemaattis-luonnontieteellinen tiedekunta</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="organization" lang="en">University of Helsinki, Faculty of Science</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="organization" lang="sv">Helsingfors universitet, Matematisk-naturvetenskapliga fakulteten</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2025-09-11T09:06:57Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="available">2025-09-11T09:06:57Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued">2025-09-11</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="licenseGranted">2025-06-09</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">http://hdl.handle.net/10138/601034</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="urn">URN:NBN:fi:hulib-202509113896</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract" lang="en">Developing Machine learning models that provides a stable performance across domains unseen
during training, remains a persistent challenge. This generalizability of models is particularly
challenging when the relationship between input features and target labels varies across domains.
This study is motivated by recent work suggesting that causal features generalize better across
domains as opposed to non-causal features. However, the evaluations comprised of varying sizes of
input information across the models being compared.
This study addresses this limitation by balancing the number of input features across all models,
thereby improving the fairness of the comparison. Experiments were conducted on four real-world
tabular datasets with different application domains - ANES(Voting), ASSISTments, BRFSS (Diabetes),
and NHANES (Blood Lead). The evaluations utilized tree-based classifiers, MLPs and
domain generalization approaches, including GroupDRO and REx. Each model was trained on
data from one domain and evaluated on another unseen test domain to assess robustness under a
strictly binary domain shift.
Results show that models trained on causal features achieved more stable generalization, particularly
in datasets with strong causal relationships to the target label. Tree-based models consistently
outperformed MLPs.
Despite the strictly binary experimental setup, domain generalization methods that are designed for
group settings were implemented for the purpose of exploration. As expected, Rex and GroupDRO
performed poorly in this setting. Interpretability analyses using SHAP and LIME were done on all
datasets, in an XGBoost model, which offered insights on the level of influence the features have to
the target, which supported the results.
Overall, this work investigates the role of causal features in enhancing robustness under domain
shift, by implementing a fair, interpretable experiment pipeline across real-world datasets.</dim:field>
   <dim:field mdschema="dc" element="language" qualifier="iso">eng</dim:field>
   <dim:field mdschema="dc" element="publisher" lang="fi">Helsingin yliopisto</dim:field>
   <dim:field mdschema="dc" element="publisher" lang="en">University of Helsinki</dim:field>
   <dim:field mdschema="dc" element="publisher" lang="sv">Helsingfors universitet</dim:field>
   <dim:field mdschema="dc" element="rights">In Copyright 1.0</dim:field>
   <dim:field mdschema="dc" element="subject">machine learning</dim:field>
   <dim:field mdschema="dc" element="subject">causal inference</dim:field>
   <dim:field mdschema="dc" element="subject">domain generalization</dim:field>
   <dim:field mdschema="dc" element="subject">interpretability</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="degreeprogram" lang="fi">Datatieteen maisteriohjelma</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="degreeprogram" lang="en">Master&amp;apos;s Programme in Data Science</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="degreeprogram" lang="sv">Magisterprogrammet i data science</dim:field>
   <dim:field mdschema="dc" element="title" lang="en">Evaluating the Role of Balanced Causal and Non-Causal Features in Predictive Modeling</dim:field>
   <dim:field mdschema="dc" element="type" qualifier="ontasot" lang="fi">pro gradu -tutkielma</dim:field>
   <dim:field mdschema="dc" element="type" qualifier="ontasot" lang="en">master&amp;apos;s thesis</dim:field>
   <dim:field mdschema="dc" element="type" qualifier="ontasot" lang="sv">pro gradu -avhandling</dim:field>
   <dim:field mdschema="others" element="access-status">open.access</dim:field>
</dim:dim></metadata></record></GetRecord></OAI-PMH>