You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:在泰坦尼克数据集中创建新变量及生存统计问题

解决泰坦尼克号数据集儿童幸存统计问题

嘿,我来帮你搞定这个儿童幸存指标的统计问题!先回顾下你已经完成的步骤:复制了原始数据集到titanic_age_groups,成功创建了child_or_adult分组变量,接下来咱们就一步步搞定幸存儿童数、儿童总数和幸存占比的计算。

假设你用SAS处理数据集

首先先确认下你创建child_or_adult的代码(如果和下面不一样也没关系,核心是筛选儿童组):

data titanic_age_groups;
    set titanic1;
    /* 这里假设儿童定义为年龄小于18岁,你可以根据需求调整阈值 */
    if age < 18 then child_or_adult = 'Child';
    else if age = . then child_or_adult = 'Unknown'; /* 处理年龄缺失值 */
    else child_or_adult = 'Adult';
run;

方法1:用PROC SQL直接计算指标

这是最灵活的方式,能直接生成包含所有指标的统计数据集:

proc sql;
    create table child_survival_summary as
    select
        count(*) as total_children format=comma10.0, /* 儿童总数 */
        sum(survived) as survived_children format=comma10.0, /* 幸存儿童数(假设survived是1=幸存,0=遇难) */
        (sum(survived)/count(*)) as child_survival_rate format=percent7.1 /* 幸存占比,百分比格式显示 */
    from titanic_age_groups
    where child_or_adult = 'Child'; /* 只筛选儿童组 */
quit;

运行完后,child_survival_summary数据集里就有你需要的所有指标了。

方法2:用PROC FREQ快速查看交叉统计

如果你只是想快速查看结果,不需要生成数据集,用proc freq的交叉表更直观:

proc freq data=titanic_age_groups;
    tables child_or_adult*survived / row percent; /* 行百分比就是每组的幸存率 */
    where child_or_adult = 'Child'; /* 只看儿童组 */
run;

输出结果里的行百分比就是儿童的幸存占比,同时也能看到儿童总数和幸存数。

如果你用Python Pandas处理数据集

如果你的工作环境是Python,那代码可以这么写:
先确认你已经完成的步骤:

import pandas as pd

# 读取原始数据集(假设是csv格式,你可以根据实际格式调整)
titanic1 = pd.read_csv('titanic.csv')
# 复制为新数据集
titanic_age_groups = titanic1.copy()
# 创建分组变量,同时处理年龄缺失值
titanic_age_groups['child_or_adult'] = titanic_age_groups['age'].apply(
    lambda x: 'Child' if x < 18 else 'Unknown' if pd.isna(x) else 'Adult'
)

计算所需统计指标

# 筛选出所有儿童数据
children_subset = titanic_age_groups[titanic_age_groups['child_or_adult'] == 'Child']

# 计算各项指标
total_children = len(children_subset)
survived_children = children_subset['survived'].sum()  # 假设survived是1/0格式
child_survival_rate = survived_children / total_children

# 打印结果
print(f"👉 儿童总数: {total_children}")
print(f"👉 幸存儿童数: {survived_children}")
print(f"👉 幸存儿童占比: {child_survival_rate:.1%}")

注意事项

  • 一定要处理年龄缺失值!如果你的数据集里有缺失的年龄,直接分到成人组会导致统计结果偏差,建议标记为Unknown单独处理。
  • 确认survived变量的编码:如果是字符串格式(比如'Yes'/'No'),需要先转换成数值(1/0)再求和,比如在SAS里用if survived = 'Yes' then survived_num = 1; else 0;,在Pandas里用children_subset['survived'] = children_subset['survived'].map({'Yes':1, 'No':0})。

内容的提问来源于stack exchange,提问作者Reeza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:20:08