You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中提取数据框相邻字符串并按答案及条件拆分存储?

需求与数据说明

数据结构

有一个包含二进制答案(1/2)和对应理由的数据框,结构如下:

A_answer1reasoning1B_answer2reasoning2A_answer3reasoning3
1"some statement"2"some reasoning"1"yes"
2"another statement"1"some sentence"1"because of x"

已完成任务

已实现将所有答案为1的相邻理由存入df1,答案为2的存入df2,代码如下:

drive <- list() #create an empty list
yield <- list()
c <- 0

for(i in seq(3,ncol(df),2)) {
  c<- c+1
  temp <- df[i+1]
  drive[[c]] <- temp[df[i]==1]
  yield[[c]] <- temp[df[i]==2]
}


lapply(drive, function(x) write.table( data.frame(x), 'drive.csv'  , append= T, sep=',' ))
lapply(yield, function(x) write.table( data.frame(x), 'yield.csv'  , append= T, sep=',' ))

待解决任务

需要按条件拆分:A_answer1和A_answer3属于条件"A",提取条件A下:

  • 答案为1的理由到df1_A,期望结果:
    "some statement"
    "yes"
    "because of x"
    
  • 答案为2的理由到df2_A,期望结果:
    "another statement"
    

实现方法

代码实现

# 筛选条件A的答案列(匹配以A_answer开头的列名)
a_answer_cols <- grep("^A_answer", colnames(df), value = TRUE)
# 获取对应理由列(每个答案列的下一列)
a_reason_cols <- sapply(a_answer_cols, function(col) {
  col_idx <- which(colnames(df) == col)
  colnames(df)[col_idx + 1]
})

# 初始化结果容器
df1_A <- c()
df2_A <- c()

# 遍历每一组A类答案-理由列对
for (i in seq_along(a_answer_cols)) {
  ans_col <- a_answer_cols[i]
  reason_col <- a_reason_cols[i]
  
  # 提取答案为1的理由并合并
  df1_A <- c(df1_A, df[[reason_col]][df[[ans_col]] == 1])
  # 提取答案为2的理由并合并
  df2_A <- c(df2_A, df[[reason_col]][df[[ans_col]] == 2])
}

# 转为数据框(可选,按需调整)
df1_A <- data.frame(reason = df1_A)
df2_A <- data.frame(reason = df2_A)

# 写入CSV(可选)
write.table(df1_A, "df1_A.csv", sep = ",", row.names = FALSE)
write.table(df2_A, "df2_A.csv", sep = ",", row.names = FALSE)

代码说明

  • grep("^A_answer", colnames(df), value = TRUE):精准筛选条件A对应的答案列,避免误选其他前缀的列
  • 通过遍历每一组答案-理由列对,针对性提取符合条件的理由并合并
  • 最后可根据需求将结果转为数据框或直接写入文件

内容的提问来源于stack exchange,提问作者Hatice Sahin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 02:10:34