You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从数据框两列提取CR/PD/PR/SD标签并新增列(保留原数据)

问题

我有一个仅关注两列的数据框,这两列中包含需要提取的4个标签:CR, PD, PR, SD,但标签附带io.response或pfs等无关字符串。数据框示例如下:

structure(list(`!Sample_characteristics_ch1.22` = c("duration.of.io.tx: 174", 
"io.response: PD", "io.response: PD", "duration.of.io.tx: 21", 
"io.response: PD", "duration.of.io.tx: 21", "io.response: PD", 
"io.response: PD", "io.response: PR", "duration.of.io.tx: 157", 
"io.response: PD"), `!Sample_characteristics_ch1.23` = c("io.response: PD", 
"pfs: 106", "pfs: 57", "io.response: PD", "pfs: 30", "io.response: PD", 
"pfs: 25", "pfs: 17", "pfs: 338", "io.response: SD", "pfs: 41"
)), row.names = c("Patient sample BACI139", "Patient sample BACI140", 
"Patient sample BACI142", "Patient sample BACI143", "Patient sample BACI144", 
"Patient sample BACI148", "Patient sample BACI149", "Patient sample BACI150", 
"Patient sample BACI151", "Patient sample BACI152", "Patient sample BACI153"
), class = "data.frame")

需求:新增一列(命名不限),仅包含上述4个标签,且不修改或删除原列以保留原始数据。
示例:第一行第二列是io.response: PD,新列对应值为PD;第二行第一列是io.response: PD,新列对应值也为PD。

解决方案

方法一:使用tidyverse工具链

通过dplyr做数据处理,stringr做字符串匹配,步骤清晰易读:

  1. 合并两列文本,统一匹配逻辑;
  2. 用正则精准定位CR/PD/PR/SD标签;
  3. 新增列存储提取结果,可选择删除临时合并列。

代码示例:

library(tidyverse)

# 加载数据
df <- structure(list(`!Sample_characteristics_ch1.22` = c("duration.of.io.tx: 174", 
"io.response: PD", "io.response: PD", "duration.of.io.tx: 21", 
"io.response: PD", "duration.of.io.tx: 21", "io.response: PD", 
"io.response: PD", "io.response: PR", "duration.of.io.tx: 157", 
"io.response: PD"), `!Sample_characteristics_ch1.23` = c("io.response: PD", 
"pfs: 106", "pfs: 57", "io.response: PD", "pfs: 30", "io.response: PD", 
"pfs: 25", "pfs: 17", "pfs: 338", "io.response: SD", "pfs: 41"
)), row.names = c("Patient sample BACI139", "Patient sample BACI140", 
"Patient sample BACI142", "Patient sample BACI143", "Patient sample BACI144", 
"Patient sample BACI148", "Patient sample BACI149", "Patient sample BACI150", 
"Patient sample BACI151", "Patient sample BACI152", "Patient sample BACI153"
), class = "data.frame")

# 新增response_label列
df <- df %>%
  mutate(
    combined_text = paste(`!Sample_characteristics_ch1.22`, `!Sample_characteristics_ch1.23`),
    response_label = str_extract(combined_text, "(CR|PD|PR|SD)")
  ) %>%
  select(-combined_text) # 可选:移除临时合并列

# 查看提取结果
print(df$response_label)

输出结果:

[1] "PD" "PD" "PD" "PD" "PD" "PD" "PD" "PD" "PR" "SD" "PD"

方法二:使用Base R

无需额外加载包,用原生函数实现:

# 加载数据(同上)
df <- structure(list(`!Sample_characteristics_ch1.22` = c("duration.of.io.tx: 174", 
"io.response: PD", "io.response: PD", "duration.of.io.tx: 21", 
"io.response: PD", "duration.of.io.tx: 21", "io.response: PD", 
"io.response: PD", "io.response: PR", "duration.of.io.tx: 157", 
"io.response: PD"), `!Sample_characteristics_ch1.23` = c("io.response: PD", 
"pfs: 106", "pfs: 57", "io.response: PD", "pfs: 30", "io.response: PD", 
"pfs: 25", "pfs: 17", "pfs: 338", "io.response: SD", "pfs: 41"
)), row.names = c("Patient sample BACI139", "Patient sample BACI140", 
"Patient sample BACI142", "Patient sample BACI143", "Patient sample BACI144", 
"Patient sample BACI148", "Patient sample BACI149", "Patient sample BACI150", 
"Patient sample BACI151", "Patient sample BACI152", "Patient sample BACI153"
), class = "data.frame")

# 定义提取函数
extract_label <- function(row) {
  text <- paste(row[1], row[2])
  regmatches(text, regexpr("CR|PD|PR|SD", text))
}

# 新增response_label列
df$response_label <- apply(df[, c("!Sample_characteristics_ch1.22", "!Sample_characteristics_ch1.23")], 1, extract_label)

# 查看提取结果
print(df$response_label)

输出结果与方法一完全一致。

内容的提问来源于stack exchange,提问作者Kev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 05:15:35