如何基于samp列数值间断在R数据框中插入多行补全数据
补全R数据框samp列间断行的问题解决
原始数据
我有如下R数据框:
patient_ID CBC CBN totindex samp index 120 1007 BLOQ BLOQ 7 8 1 121 1007 BLOQ BLOQ 8 9 1 122 1007 BLOQ BLOQ 9 10 1 123 1007 BLOQ BLOQ 10 11 1 124 1007 BLOQ BLOQ 11 12 1 125 1007 BLOQ BLOQ 12 15 4 126 1007 BLOQ BLOQ 13 16 1 127 1007 BLOQ BLOQ 14 17 1 128 1007 BLOQ BLOQ 15 18 1 129 1007 BLOQ BLOQ 16 19 1 130 1007 BLOQ BLOQ 17 20 1
需求说明
我用index列标记samp列和上一行不连续的行:差值为1标记3,差值为2标记4,差值3/4分别标记5/6。需要在这些间断处插入1-4行补全数据:
- 保持
patient_ID不变(比如示例中的1007) - 其他列填0(或指定值)
- 补全
samp列的缺失值(示例中补13、14) - 所有行的
index改为1
最终要得到这样的输出:
patient_ID CBC CBN totindex samp index 120 1007 BLOQ BLOQ 7 8 1 121 1007 BLOQ BLOQ 8 9 1 122 1007 BLOQ BLOQ 9 10 1 123 1007 BLOQ BLOQ 10 11 1 124 1007 BLOQ BLOQ 11 12 1 125 1007 0 0 0 13 1 126 1007 0 0 0 14 1 127 1007 BLOQ BLOQ 12 15 1 128 1007 BLOQ BLOQ 13 16 1 129 1007 BLOQ BLOQ 14 17 1 130 1007 BLOQ BLOQ 15 18 1 131 1007 BLOQ BLOQ 16 19 1 132 1007 BLOQ BLOQ 17 20 1
尝试的for循环方案(存在问题)
一开始写了for循环处理,但代码只能处理前几个标记行,覆盖不了最后一个标记为4的行,而且会在最后一个间断处停止,丢失后续数据:
nidx4 <- as.numeric(rownames(df[grep("4", df$index), ])) dfnew <- data.frame() for (idx in 1:length(nidx4)) { if (idx==1){ df1 = df[1:(nidx4[idx]-1),] } else if (idx == length(nidx4)) { df1 = df[nidx4[idx-1]:nrow(df),] } else { df1 = df[(nidx4[idx-1]):(nidx4[idx]-1),] } df1[nrow(df1)+1,] = 0 df1[nrow(df1)+1,] = 0 df1[nrow(df1)-1,21] = df1[nrow(df1)-2,21]+1 df1[nrow(df1),21] = df1[nrow(df1)-1,21]+1 dfnew = rbind(dfnew,df1) } for (row in 1:nrow(dfnew)){ if (dfnew[row,"index"] == 0) {dfnew[row,"index"] = 1} if (dfnew[row,"index"] == 4) {dfnew[row,"index"] = 1} } rownames(dfnew) <- NULL df <- dfnew
最终解决代码
用tidyr的complete函数可以完美解决,自动补全samp列的连续序列,指定缺失行的填充值,最后统一把index设为1:
library(tidyr) library(dplyr) dfnew <- complete(df, patient_ID, samp = full_seq(samp, period = 1), fill = list("Sample_Name_(run_ID)" = "no_sample", Sample_Name = "no_sample", THC = "0", OH_THC = "0", THC_COOH = "0", THC_COO_gluc = "0", THC_gluc = "0", CBD = "0", "6aOH_CBD" = "0", "7OH_CBD" = "0", "6bOH_CBD" = "0", CBD_COOH = "0", CBD_gluc = "0", CBC = "0", CBN = "0", CBG = "0", THCV = "0", CBDV = "0", totindex = 0, index = 1)) %>% mutate(index = 1)
内容的提问来源于stack exchange,提问作者jc2525
相关产品推荐
相关产品推荐

