如何识别有序列中的首个零值?生存分析数据集处理问询
解决生存分析数据集的消亡时间标记问题
嘿,这个需求很明确——给每所学校标记首次招生人数为0的“消亡年份”对吧?我给你准备了两种常用工具的实现方法,R和Python Pandas,你可以根据自己的工作环境选:
方法一:使用R语言(基础函数或Tidyverse)
步骤1:构造示例数据集
首先先还原你的示例数据:
school_data <- data.frame( Name = c("a", "b", "c", "d"), total.89 = c(8, 1, 7, 2), total.90 = c(6, 2, 9, 0), total.91 = c(4, 4, 0, 0), total.92 = c(0, 9, 0, 0) )
步骤2:用基础函数快速处理
我们可以用apply逐行扫描,找到首次出现0的列并提取年份:
# 定义函数:找到每行首次出现0的年份 find_first_zero_year <- function(row) { # 排除Name列,筛选值为0的列位置 zero_cols <- which(row[-1] == 0) if (length(zero_cols) == 0) { return(NA) # 没有0值的学校返回NA } else { # 提取第一个0对应的列名,去掉"total."前缀得到年份 first_zero_col <- names(row[-1])[min(zero_cols)] return(sub("total\\.", "", first_zero_col)) } } # 新增消亡时间列 school_data$death_year <- apply(school_data, 1, find_first_zero_year)
步骤3:查看结果
运行后你的数据集会新增death_year列:
print(school_data) # Name total.89 total.90 total.91 total.92 death_year # 1 a 8 6 4 0 92 # 2 b 1 2 4 9 <NA> # 3 c 7 9 0 0 91 # 4 d 2 0 0 0 90
可选:用Tidyverse的tidy data风格处理
如果你习惯用tidyverse工具链(dplyr+tidyr),可以把宽数据转长后再分组处理:
library(tidyverse) school_data_tidy <- school_data %>% # 把年份列转成长格式 pivot_longer(cols = starts_with("total."), names_to = "year", values_to = "total") %>% # 提取年份数字 mutate(year = str_remove(year, "total\\.")) %>% # 按学校分组 group_by(Name) %>% # 筛选出招生为0的行,取第一行(首次消亡) filter(total == 0) %>% slice_head(n = 1) %>% # 只保留学校名和消亡年份 select(Name, death_year = year) %>% # 合并回原数据集,保留所有学校 right_join(school_data, by = "Name") %>% # 把消亡年份列移到Name后面,更直观 relocate(death_year, .after = Name) print(school_data_tidy)
方法二:使用Python Pandas
步骤1:构造示例数据集
先还原你的数据到DataFrame:
import pandas as pd data = { "Name": ["a", "b", "c", "d"], "total.89": [8, 1, 7, 2], "total.90": [6, 2, 9, 0], "total.91": [4, 4, 0, 0], "total.92": [0, 9, 0, 0] } school_df = pd.DataFrame(data)
步骤2:用apply逐行处理
和R的思路类似,写一个函数逐行找首次0的年份:
# 定义函数:返回每行首次出现0的年份 def get_death_year(row): # 排除Name列,只看招生人数列 total_cols = row.drop("Name") # 找到所有值为0的列索引 zero_indices = total_cols[total_cols == 0].index if len(zero_indices) == 0: return None # 无0值的学校返回None else: # 提取第一个0列的年份部分 first_zero_col = zero_indices[0] return first_zero_col.split(".")[1] # 新增消亡时间列 school_df["death_year"] = school_df.apply(get_death_year, axis=1)
步骤3:查看结果
运行后数据集会新增death_year列:
print(school_df) # Name total.89 total.90 total.91 total.92 death_year # 0 a 8 6 4 0 92 # 1 b 1 2 4 9 None # 2 c 7 9 0 0 91 # 3 d 2 0 0 0 90
可选:用melt转长数据处理
如果你更喜欢tidy data的方式,可以转长后分组处理:
# 把宽数据转成长格式 melted_df = school_df.melt(id_vars="Name", var_name="year", value_name="total") # 从列名提取年份 melted_df["year"] = melted_df["year"].str.split(".").str[1] # 按学校分组,取首次出现0的年份 death_years = melted_df[melted_df["total"] == 0].groupby("Name").first()["year"] # 合并回原数据集 school_df = school_df.merge(death_years, on="Name", how="left").rename(columns={"year": "death_year"}) print(school_df)
内容的提问来源于stack exchange,提问作者Ryan Parsons
相关产品推荐
相关产品推荐

