You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

统计无序匹配数量并指定唯一字符串:R语言数据框处理需求

解决方案

可以使用dplyr和tidyr包处理数据,按ID分组对比预测试(pre)和后测试(post)的答案组件,生成目标结果:

# 加载所需包
library(dplyr)
library(tidyr)

# 处理数据生成结果
res <- dat %>%
  # 按ID分组,确保每个ID对应pre、post两条记录
  group_by(ID) %>%
  # 将pre和post的answer转为单独列,方便对比
  pivot_wider(names_from = test, values_from = answer) %>%
  # 拆分答案为组件向量,转为无序集合
  mutate(
    pre_components = strsplit(pre, "_") %>% purrr::map(unique),
    post_components = strsplit(post, "_") %>% purrr::map(unique),
    # 生成res1:判断两组组件是否无序匹配
    res1 = ifelse(
      purrr::map2_lgl(pre_components, post_components, ~setequal(.x, .y)),
      "yes", "no"
    ),
    # 生成res2:统计post中新增组件的数量
    res2 = purrr::map2_int(pre_components, post_components, ~length(setdiff(.y, .x))),
    # 生成res3:拼接post中新增的组件,无新增则设为NA
    res3 = purrr::map2_chr(pre_components, post_components, ~{
      diff_vals <- setdiff(.y, .x)
      if (length(diff_vals) == 0) NA_character_ else paste(diff_vals, collapse = "_")
    })
  ) %>%
  # 保留目标列并取消分组
  select(ID, res1, res2, res3) %>%
  ungroup()

# 查看结果
print(res)

代码说明:

  • pivot_wider:将每个ID的pre、post答案转为两列,简化后续对比操作
  • strsplit:把下划线分隔的答案拆分为字符向量,unique兼容存在重复组件的通用场景
  • setequal:判断两个集合是否完全相等(不考虑顺序),直接生成res1的二元标识
  • setdiff(.y, .x):提取post独有的组件,用length统计数量得到res2,用paste拼接成字符串得到res3

运行代码后输出的结果与目标res完全一致。

内容的提问来源于stack exchange,提问作者Danielle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 03:31:22