R语言数据框处理:按前缀匹配及Period值提取行并补全NA问题
R语言数据匹配提取解决方案
问题1:按前缀匹配提取指定范围行
实现逻辑
- 遍历
df1每一行,提取对应的周期值与三字符前缀 - 匹配
df2中Index列前3位与前缀一致的行位置 - 对每个匹配位置,生成「前N行(N为周期值)+ 当前匹配行」的行号序列
- 过滤有效行号后提取对应数据,最终输出长度为5的结果列表,每个元素对应
df1一行的输出数据框
实现代码
# 基于示例数据运行即可 result_lst <- apply(Df1, 1, function(row) { n_period <- as.numeric(row[["Period"]]) prefix <- row[["Index_1"]] # 匹配前缀一致的行位置 match_pos <- which(substr(Df2$Index, 1, 3) == prefix) # 生成所有目标行号 target_pos <- unlist(lapply(match_pos, function(p) seq(p - n_period, p))) # 过滤合法行号 target_pos <- target_pos[target_pos > 0 & target_pos <= nrow(Df2)] # 返回提取结果 return(Df2[target_pos, c("Index", "Growth")]) }) # 单独提取5个结果数据框 res1 <- result_lst[[1]] res2 <- result_lst[[2]] res3 <- result_lst[[3]] res4 <- result_lst[[4]] res5 <- result_lst[[5]]
问题2:不足行自动填充NA
实现逻辑
不对小于0的行号做过滤,对非法行号(小于1/大于df2总行数)的位置生成填充NA的行,与有效行拼接即可实现补全需求。
实现代码
result_lst_fill <- apply(Df1, 1, function(row) { n_period <- as.numeric(row[["Period"]]) prefix <- row[["Index_1"]] match_pos <- which(substr(Df2$Index, 1, 3) == prefix) # 处理每个匹配位置的补全 all_res <- lapply(match_pos, function(p) { pos_seq <- seq(p - n_period, p) # 逐行判断,非法位置补NA per_pos_res <- lapply(pos_seq, function(x) { if(x < 1 || x > nrow(Df2)) { data.frame(Index = NA_character_, Growth = NA_real_) } else { Df2[x, c("Index", "Growth")] } }) do.call(rbind, per_pos_res) }) # 合并同一前缀所有匹配位置的结果 do.call(rbind, all_res) }) # 查看问题2示例的预期结果Result3 result3 <- result_lst_fill[[1]]
内容的提问来源于stack exchange,提问作者Selina1
相关产品推荐
相关产品推荐

