You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于samp列数值间断在R数据框中插入多行补全数据

补全R数据框samp列间断行的问题解决

原始数据

我有如下R数据框:

patient_ID  CBC  CBN totindex samp index
120       1007 BLOQ BLOQ        7    8     1
121       1007 BLOQ BLOQ        8    9     1
122       1007 BLOQ BLOQ        9   10     1
123       1007 BLOQ BLOQ       10   11     1
124       1007 BLOQ BLOQ       11   12     1
125       1007 BLOQ BLOQ       12   15     4
126       1007 BLOQ BLOQ       13   16     1
127       1007 BLOQ BLOQ       14   17     1
128       1007 BLOQ BLOQ       15   18     1
129       1007 BLOQ BLOQ       16   19     1
130       1007 BLOQ BLOQ       17   20     1

需求说明

我用index列标记samp列和上一行不连续的行:差值为1标记3,差值为2标记4,差值3/4分别标记5/6。需要在这些间断处插入1-4行补全数据:

  • 保持patient_ID不变(比如示例中的1007)
  • 其他列填0(或指定值)
  • 补全samp列的缺失值(示例中补13、14)
  • 所有行的index改为1

最终要得到这样的输出:

patient_ID  CBC  CBN totindex samp index
120       1007 BLOQ BLOQ        7    8     1
121       1007 BLOQ BLOQ        8    9     1
122       1007 BLOQ BLOQ        9   10     1
123       1007 BLOQ BLOQ       10   11     1
124       1007 BLOQ BLOQ       11   12     1
125       1007    0    0        0   13     1
126       1007    0    0        0   14     1
127       1007 BLOQ BLOQ       12   15     1
128       1007 BLOQ BLOQ       13   16     1
129       1007 BLOQ BLOQ       14   17     1
130       1007 BLOQ BLOQ       15   18     1
131       1007 BLOQ BLOQ       16   19     1
132       1007 BLOQ BLOQ       17   20     1

尝试的for循环方案(存在问题)

一开始写了for循环处理,但代码只能处理前几个标记行,覆盖不了最后一个标记为4的行,而且会在最后一个间断处停止,丢失后续数据:

nidx4 <- as.numeric(rownames(df[grep("4", df$index), ]))
dfnew <- data.frame()

for (idx in 1:length(nidx4)) {
  if (idx==1){
    df1 = df[1:(nidx4[idx]-1),]
  }
  else if (idx == length(nidx4)) {
    df1 = df[nidx4[idx-1]:nrow(df),]
  }
  else {
    df1 = df[(nidx4[idx-1]):(nidx4[idx]-1),]
  }
  df1[nrow(df1)+1,] = 0
  df1[nrow(df1)+1,] = 0
  df1[nrow(df1)-1,21] = df1[nrow(df1)-2,21]+1
  df1[nrow(df1),21] = df1[nrow(df1)-1,21]+1
  dfnew = rbind(dfnew,df1)
}

for (row in 1:nrow(dfnew)){
  if (dfnew[row,"index"] == 0) {dfnew[row,"index"] = 1}
  if (dfnew[row,"index"] == 4) {dfnew[row,"index"] = 1}
}

rownames(dfnew) <- NULL
df <- dfnew 

最终解决代码

用tidyr的complete函数可以完美解决,自动补全samp列的连续序列,指定缺失行的填充值,最后统一把index设为1:

library(tidyr)
library(dplyr)

dfnew <- complete(df, patient_ID, samp = full_seq(samp, period = 1),
         fill = list("Sample_Name_(run_ID)" = "no_sample",
                     Sample_Name = "no_sample", THC = "0", OH_THC = "0",
                     THC_COOH = "0", THC_COO_gluc = "0", THC_gluc = "0", 
                     CBD = "0", "6aOH_CBD" = "0", "7OH_CBD" = "0", 
                     "6bOH_CBD" = "0", CBD_COOH = "0", CBD_gluc = "0", 
                     CBC = "0", CBN = "0", CBG = "0", THCV = "0", CBDV = "0",
                     totindex = 0, 
                     index = 1)) %>%
  mutate(index = 1)

内容的提问来源于stack exchange,提问作者jc2525

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 11:15:38