R语言:在泰坦尼克数据集中创建新变量及生存统计问题
解决泰坦尼克号数据集儿童幸存统计问题
嘿,我来帮你搞定这个儿童幸存指标的统计问题!先回顾下你已经完成的步骤:复制了原始数据集到titanic_age_groups,成功创建了child_or_adult分组变量,接下来咱们就一步步搞定幸存儿童数、儿童总数和幸存占比的计算。
假设你用SAS处理数据集
首先先确认下你创建child_or_adult的代码(如果和下面不一样也没关系,核心是筛选儿童组):
data titanic_age_groups; set titanic1; /* 这里假设儿童定义为年龄小于18岁,你可以根据需求调整阈值 */ if age < 18 then child_or_adult = 'Child'; else if age = . then child_or_adult = 'Unknown'; /* 处理年龄缺失值 */ else child_or_adult = 'Adult'; run;
方法1:用PROC SQL直接计算指标
这是最灵活的方式,能直接生成包含所有指标的统计数据集:
proc sql; create table child_survival_summary as select count(*) as total_children format=comma10.0, /* 儿童总数 */ sum(survived) as survived_children format=comma10.0, /* 幸存儿童数(假设survived是1=幸存,0=遇难) */ (sum(survived)/count(*)) as child_survival_rate format=percent7.1 /* 幸存占比,百分比格式显示 */ from titanic_age_groups where child_or_adult = 'Child'; /* 只筛选儿童组 */ quit;
运行完后,child_survival_summary数据集里就有你需要的所有指标了。
方法2:用PROC FREQ快速查看交叉统计
如果你只是想快速查看结果,不需要生成数据集,用proc freq的交叉表更直观:
proc freq data=titanic_age_groups; tables child_or_adult*survived / row percent; /* 行百分比就是每组的幸存率 */ where child_or_adult = 'Child'; /* 只看儿童组 */ run;
输出结果里的行百分比就是儿童的幸存占比,同时也能看到儿童总数和幸存数。
如果你用Python Pandas处理数据集
如果你的工作环境是Python,那代码可以这么写:
先确认你已经完成的步骤:
import pandas as pd # 读取原始数据集(假设是csv格式,你可以根据实际格式调整) titanic1 = pd.read_csv('titanic.csv') # 复制为新数据集 titanic_age_groups = titanic1.copy() # 创建分组变量,同时处理年龄缺失值 titanic_age_groups['child_or_adult'] = titanic_age_groups['age'].apply( lambda x: 'Child' if x < 18 else 'Unknown' if pd.isna(x) else 'Adult' )
计算所需统计指标
# 筛选出所有儿童数据 children_subset = titanic_age_groups[titanic_age_groups['child_or_adult'] == 'Child'] # 计算各项指标 total_children = len(children_subset) survived_children = children_subset['survived'].sum() # 假设survived是1/0格式 child_survival_rate = survived_children / total_children # 打印结果 print(f"👉 儿童总数: {total_children}") print(f"👉 幸存儿童数: {survived_children}") print(f"👉 幸存儿童占比: {child_survival_rate:.1%}")
注意事项
- 一定要处理年龄缺失值!如果你的数据集里有缺失的年龄,直接分到成人组会导致统计结果偏差,建议标记为
Unknown单独处理。 - 确认
survived变量的编码:如果是字符串格式(比如'Yes'/'No'),需要先转换成数值(1/0)再求和,比如在SAS里用if survived = 'Yes' then survived_num = 1; else 0;,在Pandas里用children_subset['survived'] = children_subset['survived'].map({'Yes':1, 'No':0})。
内容的提问来源于stack exchange,提问作者Reeza
相关产品推荐
相关产品推荐

