You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

泰坦尼克号案例:R与Python生成的决策树结构相反问题排查

Titanic Decision Tree Discrepancy Between R and Python (Graphviz Output)

I’ve built a validated decision tree for the Titanic dataset using R, and I’m trying to replicate the exact same tree in Python using Graphviz. Since I couldn’t import Graphviz in Spyder, I exported the model to a DOT file and used the WebGraphviz tool to generate the visualization with this code:

import sklearn.tree as tree
tree.export_graphviz(rpart, out_file="tree.dot", filled=True, feature_names=list(titanic_dmy.drop(['survived'], axis=1).columns), impurity=False, label=None, proportion=True, class_names=['Survived', 'Died'])

While the numerical values are close, the core issue is that the branch directions are completely reversed between the two trees. For example, in R, the male branch leads to the 'age' node and the female branch leads to the 'Third class' node—but in Python, these paths are swapped. This reversal even leads to contradictory conclusions: R’s tree shows females have higher survival rates, while Python’s tree appears to show males do.

I’m using the exact same dataset and feature columns for both models. Here are the key areas to check for this discrepancy:

  • Categorical variable encoding order
    R and Python often handle categorical variable levels differently by default. For example, R’s factor() might order sex as female → male, while Python’s pd.get_dummies() or LabelEncoder could reverse that. This flip would cause the decision tree’s split conditions to invert. Verify the level order of all categorical features (like sex, Pclass) in both environments.

  • Model parameter mismatches
    R’s rpart and Python’s sklearn.tree.DecisionTreeClassifier have different default settings for CART trees. Check parameters like:

    • Splitting criterion (e.g., gini vs entropy—confirm consistency between tools)
    • Pruning parameters (cp in R vs ccp_alpha in Python)
    • Minimum samples per split/leaf, maximum depth
      Even small differences here can lead to different tree structures, including reversed branches.
  • Class label mapping
    Your class_names parameter in export_graphviz is set to ['Survived', 'Died']. Confirm that R’s model uses the same class order: if R’s survived variable treats 0 as 'Survived' and 1 as 'Died' (or vice versa), this would flip the final class conclusions in the visualization.

  • Preprocessing inconsistencies
    Double-check that all preprocessing steps match exactly:

    • Missing value handling (imputation methods, which values are dropped)
    • Feature engineering (e.g., whether Pclass is treated as categorical or numerical, encoding type)
    • Training data selection (if you’re using a subset of the dataset, ensure both models train on identical rows)

内容的提问来源于stack exchange,提问作者Fish1996

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:38:33