如何通过[row,col]距离索引对指定列执行行求和循环?
问题描述
处理如下结构的数据(dput()输出):
structure(list(type3a1 = c(2L, 6L, 5L, NA, 1L, 3L, NA), type3b1 = c(NA, 3L, 1L, 5L, 6L, 3L, NA), type3a1_arc = c(1L, 2L, 5L, 4L, 5L, 4L, NA), type3b1_arc = c(2L, 2L, 3L, 4L, 1L, 1L, NA), testing = c("Yes", NA, "No", "No", NA, "Yes", NA), cars = c(5L, 12L, 1L, 6L, NA, 2L, NA), house = c(5L, 4L, 0L, 5L, 0L, 10L, NA), type3a2 = c(50L, NA, 20L, 4L, 5L, NA, NA), type3b2 = c(10L, 10L, 15L, 1L, 3L, 1L, NA), type3a2_arc = c(50L, 25L, 30L, 10L, NA, 10L, NA), type3b2_arc = c(NA, 20L, 10L, 50L, 5L, 1L, NA), X = c(NA, NA, NA, NA, NA, NA, NA)), class = "data.frame", row.names = c(NA, -7L))
需求:
- 检查每个
type变量的值是否在(1,2,3,4,5)范围内; - 若符合条件,从对应位置偏移7列获取
total值; - 将所有符合条件的
total值求和,存入新的总计列(例如第一行总计为100,因type3b1为NA不纳入计算)。
尝试的循环代码报错:Error in col + 7 : non-numeric argument to binary operator,代码如下:
# Matching pairing variables (i.e. type_vars:"type3a1" with total_vars:"type3a2") type_vars <- c("type3a1", "type3b1", "type3a1_arc", "type3b1_arc") total_vars <- c("type3a2", "type3b2", "type3a2_arc", "type3b2_arc") valid_list <- c(1,2,3,4,5) totals = list() for(row in 1:nrow(df)) { sum = 0 for(col in type_vars) { if (df[row,col] %in% valid_list) { sum <- sum + (df[row,col+7]) } } totals <- sum }
解决方案
错误原因
代码中col是列名(字符串类型),无法直接与数字7进行加法运算,这是报错的核心原因。另外,原代码中totals <- sum会每次覆盖之前的结果,最终仅保留最后一行的求和值,无法存储所有行的计算结果。
修正后的循环实现
方法1:利用预定义的变量配对关系
这种方式比硬编码偏移列更可靠,避免列顺序变动导致错误:
type_vars <- c("type3a1", "type3b1", "type3a1_arc", "type3b1_arc") total_vars <- c("type3a2", "type3b2", "type3a2_arc", "type3b2_arc") valid_list <- c(1,2,3,4,5) # 初始化存储结果的向量 totals <- numeric(nrow(df)) for(row in 1:nrow(df)) { row_sum <- 0 # 遍历变量对的索引 for(i in seq_along(type_vars)) { type_val <- df[row, type_vars[i]] # 检查值是否有效且非NA if(!is.na(type_val) && type_val %in% valid_list) { total_val <- df[row, total_vars[i]] # 避免NA值影响求和 row_sum <- row_sum + ifelse(is.na(total_val), 0, total_val) } } totals[row] <- row_sum } # 将结果添加为新列 df$total_sum <- totals
方法2:基于列索引偏移实现
如果确认列偏移关系固定为7,可以先将列名转换为索引再计算:
type_vars <- c("type3a1", "type3b1", "type3a1_arc", "type3b1_arc") valid_list <- c(1,2,3,4,5) # 获取type变量的列索引 type_cols <- match(type_vars, names(df)) # 计算对应的total列索引 total_cols <- type_cols + 7 totals <- numeric(nrow(df)) for(row in 1:nrow(df)) { row_sum <- 0 for(i in seq_along(type_cols)) { type_val <- df[row, type_cols[i]] if(!is.na(type_val) && type_val %in% valid_list) { total_val <- df[row, total_cols[i]] row_sum <- row_sum + ifelse(is.na(total_val), 0, total_val) } } totals[row] <- row_sum } df$total_sum <- totals
更高效的向量化实现(推荐)
R中循环效率较低,推荐使用向量化操作提升性能:
type_vars <- c("type3a1", "type3b1", "type3a1_arc", "type3b1_arc") total_vars <- c("type3a2", "type3b2", "type3a2_arc", "type3b2_arc") valid_list <- c(1,2,3,4,5) # 构建逻辑矩阵:标记哪些type值有效 valid_matrix <- sapply(type_vars, function(col) { !is.na(df[[col]]) & df[[col]] %in% valid_list }) # 提取对应的total值,无效位置替换为0 total_matrix <- mapply(function(t_col, v_col) { ifelse(v_col, df[[t_col]], 0) }, total_vars, as.data.frame(valid_matrix)) # 按行求和 df$total_sum <- rowSums(total_matrix, na.rm = TRUE)
运行后,第一行的total_sum会得到预期的100,其他行也会正确计算求和结果。
内容的提问来源于stack exchange,提问作者Monarch
相关产品推荐
相关产品推荐

