You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何识别有序列中的首个零值?生存分析数据集处理问询

解决生存分析数据集的消亡时间标记问题

嘿,这个需求很明确——给每所学校标记首次招生人数为0的“消亡年份”对吧?我给你准备了两种常用工具的实现方法,R和Python Pandas,你可以根据自己的工作环境选:

方法一:使用R语言(基础函数或Tidyverse)

步骤1:构造示例数据集

首先先还原你的示例数据:

school_data <- data.frame(
  Name = c("a", "b", "c", "d"),
  total.89 = c(8, 1, 7, 2),
  total.90 = c(6, 2, 9, 0),
  total.91 = c(4, 4, 0, 0),
  total.92 = c(0, 9, 0, 0)
)

步骤2:用基础函数快速处理

我们可以用apply逐行扫描,找到首次出现0的列并提取年份:

# 定义函数:找到每行首次出现0的年份
find_first_zero_year <- function(row) {
  # 排除Name列,筛选值为0的列位置
  zero_cols <- which(row[-1] == 0)
  if (length(zero_cols) == 0) {
    return(NA)  # 没有0值的学校返回NA
  } else {
    # 提取第一个0对应的列名,去掉"total."前缀得到年份
    first_zero_col <- names(row[-1])[min(zero_cols)]
    return(sub("total\\.", "", first_zero_col))
  }
}

# 新增消亡时间列
school_data$death_year <- apply(school_data, 1, find_first_zero_year)

步骤3:查看结果

运行后你的数据集会新增death_year列:

print(school_data)
#   Name total.89 total.90 total.91 total.92 death_year
# 1    a        8        6        4        0         92
# 2    b        1        2        4        9       <NA>
# 3    c        7        9        0        0         91
# 4    d        2        0        0        0         90

可选:用Tidyverse的tidy data风格处理

如果你习惯用tidyverse工具链(dplyr+tidyr),可以把宽数据转长后再分组处理:

library(tidyverse)

school_data_tidy <- school_data %>%
  # 把年份列转成长格式
  pivot_longer(cols = starts_with("total."), 
               names_to = "year", 
               values_to = "total") %>%
  # 提取年份数字
  mutate(year = str_remove(year, "total\\.")) %>%
  # 按学校分组
  group_by(Name) %>%
  # 筛选出招生为0的行,取第一行(首次消亡)
  filter(total == 0) %>%
  slice_head(n = 1) %>%
  # 只保留学校名和消亡年份
  select(Name, death_year = year) %>%
  # 合并回原数据集,保留所有学校
  right_join(school_data, by = "Name") %>%
  # 把消亡年份列移到Name后面,更直观
  relocate(death_year, .after = Name)

print(school_data_tidy)

方法二:使用Python Pandas

步骤1:构造示例数据集

先还原你的数据到DataFrame:

import pandas as pd

data = {
    "Name": ["a", "b", "c", "d"],
    "total.89": [8, 1, 7, 2],
    "total.90": [6, 2, 9, 0],
    "total.91": [4, 4, 0, 0],
    "total.92": [0, 9, 0, 0]
}
school_df = pd.DataFrame(data)

步骤2:用apply逐行处理

和R的思路类似,写一个函数逐行找首次0的年份:

# 定义函数:返回每行首次出现0的年份
def get_death_year(row):
    # 排除Name列,只看招生人数列
    total_cols = row.drop("Name")
    # 找到所有值为0的列索引
    zero_indices = total_cols[total_cols == 0].index
    if len(zero_indices) == 0:
        return None  # 无0值的学校返回None
    else:
        # 提取第一个0列的年份部分
        first_zero_col = zero_indices[0]
        return first_zero_col.split(".")[1]

# 新增消亡时间列
school_df["death_year"] = school_df.apply(get_death_year, axis=1)

步骤3:查看结果

运行后数据集会新增death_year列:

print(school_df)
#   Name  total.89  total.90  total.91  total.92 death_year
# 0    a         8         6         4         0         92
# 1    b         1         2         4         9       None
# 2    c         7         9         0         0         91
# 3    d         2         0         0         0         90

可选:用melt转长数据处理

如果你更喜欢tidy data的方式,可以转长后分组处理:

# 把宽数据转成长格式
melted_df = school_df.melt(id_vars="Name", 
                           var_name="year", 
                           value_name="total")
# 从列名提取年份
melted_df["year"] = melted_df["year"].str.split(".").str[1]
# 按学校分组,取首次出现0的年份
death_years = melted_df[melted_df["total"] == 0].groupby("Name").first()["year"]
# 合并回原数据集
school_df = school_df.merge(death_years, on="Name", how="left").rename(columns={"year": "death_year"})

print(school_df)

内容的提问来源于stack exchange,提问作者Ryan Parsons

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:40:43