You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:DataFrame每行定位最后一个1并添加汇总列及报错排查

问题:为0-1构成的DataFrame添加行汇总列并解决大数据集报错问题

需求

需要给由0和1构成的DataFrame每行添加三列汇总信息:

  • SUM:该行0和1的总和
  • LAST:该行最右侧1所在的列名
  • LAST2:该列名的最后两个字符

示例数据

dat<- structure(list(BIRD = c("aaab", "aaam", "aawm", "abum", "aema" ), `2005` = c(0, 0, 1, 1, 1), `2006` = c(0, 0, 1, 1, 0), `2007` = c(0, 0, 0, 1, 0), `2008` = c(0, 0, 0, 1, 0), `2009` = c(0, 0, 1, 1, 0), `2010` = c(0, 0, 0, 1, 0), `2011` = c(1, 1, 1, 1, 1), `2012` = c(0, 0, 1, 1, 0)), row.names = c(NA, -5L), class = c("tbl_df", "tbl", "data.frame")) 

小数据集可行代码

针对上述小数据集,使用以下代码可提取每行最后一个1的列名:

colnames(dat[apply(dat, 1, function(x) max(which(x == 1)))])

返回结果:[1] "2011" "2011" "2012" "2012" "2011"

大数据集报错问题

将相同代码应用到217×64的大数据框data2months时,出现错误:

Error in `data2months[apply(data2months, 1, function(x) max(which(x == 1)))]`: ! Can't subset columns with `apply(data2months, 1, function(x) max(which(x == 1)))`. ✖ Can't convert from `j` <double> to <integer> due to loss of precision. 

通过str()查看,两个数据框结构相似,仅规模不同。

解决方案

方案1:修正类型转换问题

报错核心原因是apply返回的是double类型索引,大数据集下精度丢失导致无法转为列索引需要的integer类型,强制转换即可解决:

# 提取每行最后一个1的列索引并转为整数
last_col_idx <- apply(data2months, 1, function(x) as.integer(max(which(x == 1))))
# 获取对应列名
last_col_names <- colnames(data2months)[last_col_idx]

方案2:dplyr+tidyr向量化操作(更高效)

对于大数据集,向量化操作比apply的逐行循环更高效,同时避免类型问题:

library(dplyr)
library(tidyr)

result <- data2months %>%
  # 计算SUM列(排除非数值列如BIRD)
  mutate(SUM = rowSums(across(where(is.numeric)))) %>%
  # 宽表转长表,方便按行筛选最后一个1
  pivot_longer(cols = -BIRD, names_to = "year", values_to = "value") %>%
  group_by(BIRD) %>%
  filter(value == 1) %>%
  slice_tail(n = 1) %>%
  select(BIRD, LAST = year) %>%
  # 合并回原表
  right_join(data2months, by = "BIRD") %>%
  # 提取列名最后两位
  mutate(LAST2 = substr(LAST, nchar(LAST)-1, nchar(LAST))) %>%
  # 调整列顺序(可选)
  select(BIRD, starts_with("20"), SUM, LAST, LAST2)

方案3:data.table操作(大数据最优解)

如果数据集规模极大,data.table的操作速度优势明显:

library(data.table)

setDT(data2months)
# 计算每行数值列的和
data2months[, SUM := rowSums(.SD), .SDcols = is.numeric]
# 按行提取最后一个1的列名
data2months[, LAST := {
  cols <- names(.SD)[.SD == 1]
  if(length(cols) > 0) tail(cols, 1) else NA_character_
}, by = BIRD, .SDcols = is.numeric]
# 提取列名最后两位
data2months[, LAST2 := substr(LAST, nchar(LAST)-1, nchar(LAST))]

内容的提问来源于stack exchange,提问作者chill

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 09:48:27