You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按id分组,从text列移除同组所有snippet字符串?

按组移除多行字符串的解决方案

问题描述

需要按id分组,将同组内所有行的snippet列字符串,从该组每一行的text列字符串中移除。例如:

  • 当id == "p1"时,要把"apple"和"orange"从该组两行的text中全部移除,最终得到" and "
  • 当id == "p2"时,因为snippet是"kiwi",text中没有该字符串,所以保持原内容不变

之前尝试用dplyr::group_by + stringr::str_remove_all只移除了当前行的snippet,未处理同组其他行的内容,达不到预期效果。

解决方案

核心思路是:分组后,先把同组内所有snippet合并成一个正则备选模式(用|分隔多个匹配项),再用这个统一的模式去移除每一行text中匹配的内容。

代码示例

library(dplyr)
library(stringr)

df_in <- tibble::tribble(
  ~id, ~snippet, ~text,
  "p1", "apple", " and orange",
  "p1", "orange", "apple and ",
  "p2", "kiwi", "fruits with apple"
)

# 正确实现代码
df_result <- df_in |>
  group_by(id) |>
  mutate(
    # 把同组所有snippet合并成正则匹配模式,比如p1组得到"apple|orange"
    remove_pattern = str_c(snippet, collapse = "|"),
    # 用合并后的模式移除text中所有匹配的内容
    text = str_remove_all(text, remove_pattern)
  ) |>
  # 可选:移除临时的remove_pattern列
  select(-remove_pattern) |>
  ungroup()

df_result
#> # A tibble: 3 × 3
#>   id    snippet text               
#>   <chr> <chr>   <chr>              
#> 1 p1    apple   " and "            
#> 2 p1    orange  " and "            
#> 3 p2    kiwi    "fruits with apple"

说明

  • str_c(snippet, collapse = "|"):将同组的所有snippet拼接成用|分隔的字符串,正则中|表示“或”,这样就能匹配同组内任意一个snippet内容
  • str_remove_all(text, remove_pattern):一次性移除text中所有匹配该模式的内容
  • 最后用select(-remove_pattern)可以删掉临时生成的模式列,保持结果和预期的df_out结构一致

内容的提问来源于stack exchange,提问作者maia-sh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 18:05:07