You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tidy_stringdist时如何保留patient、date分组名称

保留分组信息的模糊匹配解决方案

在病例报告表中,同一患者同一日期可能存在近似的comment条目(如拼写错误),需要按patient和date分组查找这类相似条目,但直接使用tidy_stringdist会丢失分组信息。以下是解决方法:

核心思路是按patient和date分组后,对每组单独执行模糊匹配逻辑,这样就能保留分组标识。

完整代码实现

library(tidystringdist)
library(dplyr)

# 原始数据
pat = tribble(
  ~patient, ~date,           ~comment,
  "a",      "10.02.2025",   "psoriasis",
  "a",      "10.02.2025",   "posriasis",
  "a",      "10.02.2025",   "psriasis",
  "a",      "10.02.2025",   "headache",
  "b",      "10.02.2025",   "fever 30",
  "b",      "10.02.2025",   "pollinosis",
  "c",      "11.03.2025",   "psoriasis",
  "c",      "11.03.2025",   "headache",
  "d",      "11.03.2025",   "headache",
)

# 过滤出每组至少2条记录的分组(减少不必要的计算)
pat_filtered = pat |>
  group_by(patient, date) |>
  filter(n() > 1) |>
  ungroup()

# 分组执行模糊匹配,自动保留patient和date信息
match_result = pat_filtered |>
  group_by(patient, date) |>
  group_modify(function(group_data, group_info) {
    group_data |>
      tidy_comb_all(comment) |>  # 生成当前组内comment的所有两两组合
      tidy_stringdist(method = "dl") |>  # 计算Damerau-Levenshtein编辑距离
      filter(dl == 1)  # 筛选编辑距离为1的相似条目
  }) |>
  ungroup()

# 查看结果
print(match_result)

输出结果

# A tibble: 3 × 5
  patient date       V1        V2           dl
  <chr>   <chr>      <chr>     <chr>     <dbl>
1 a       10.02.2025 psoriasis posriasis     1
2 a       10.02.2025 psoriasis psriasis      1
3 a       10.02.2025 posriasis psriasis      1

关键说明

  • group_modify函数会遍历每个patient-date分组:group_data是当前组的comment数据,group_info包含当前组的patient和date值
  • 对每组单独生成comment的两两组合并计算编辑距离,最终结果会自动将分组标识与匹配结果合并,解决了原代码丢失分组信息的问题

内容的提问来源于stack exchange,提问作者Dieter Menne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 00:16:05