You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整str_extract_all代码避免纤维标识重复计数?

问题描述

我有如下测试数据框:

df <- data.frame(
  product = c("apple", "banana", "cherry", "durian", "eggplant", "fuyu"),
  ingredients = c("flour|fibre|500", "sugar|500", "505|wheat|flavouring", "fibre(500)|eggs", "wholegrainrice|sesameoil", "500|fibre|500"),
  stringsAsFactors = FALSE
)

我的目标是检测产品成分中是否包含纤维、统计出现次数并提取记录纤维的标识。本次分析中,纤维可表示为"fibre"、"500"或"fibre(500)"。

当前使用的代码如下:

library(tidyverse)

fibre_strings_to_check <- c("fibre", "500", "fibre\\(500\\)")

df2 <- df %>%
  mutate(
    fibre_present = str_detect(ingredients, paste(fibre_strings_to_check, collapse = "|")),
    fibre_count = str_count(ingredients, paste(fibre_strings_to_check, collapse = "|")),
    fibre_used = str_extract_all(ingredients, paste(fibre_strings_to_check, collapse = "|"))
  )

运行后得到的输出:

productingredientsfibre_presentfibre_countfibre_used
appleflourfibre500TRUE
bananasugar500TRUE1
cherry505wheatflavouringFALSE
durianfibre(500)eggsTRUE2
eggplantwholegrainricesesameoilFALSE0
fuyu500fibre500TRUE

遇到的问题:durian产品的"fibre(500)"被拆分为"fibre"和"500"两次计数,我希望它被计为1次纤维标识,期望输出如下:

productingredientsfibre_presentfibre_countfibre_used
appleflourfibre500TRUE
bananasugar500TRUE1
cherry505wheatflavouringFALSE
durianfibre(500)eggsTRUE1
eggplantwholegrainricesesameoilFALSE0
fuyu500fibre500TRUE
解决方案

问题核心是正则匹配的优先级:原代码把短模式放在前面,导致"fibre(500)"被拆分为两个子串匹配。以下两种方法可解决该问题:

方法1:调整正则匹配顺序

将更具体、更长的模式放在匹配列表最前面,让正则引擎优先匹配完整的"fibre(500)",再匹配其他短模式:

library(tidyverse)

# 优先匹配完整的"fibre(500)",再匹配其他模式
fibre_strings_to_check <- c("fibre\\(500\\)", "fibre", "500")

df2 <- df %>%
  mutate(
    fibre_present = str_detect(ingredients, paste(fibre_strings_to_check, collapse = "|")),
    fibre_count = str_count(ingredients, paste(fibre_strings_to_check, collapse = "|")),
    fibre_used = str_extract_all(ingredients, paste(fibre_strings_to_check, collapse = "|"))
  )

方法2:拆分成分后匹配(更精准)

由于成分是用|分隔的独立项,先拆分成分列表再逐一匹配,可彻底避免部分匹配的问题:

library(tidyverse)

fibre_patterns <- c("fibre", "500", "fibre\\(500\\)")

df2 <- df %>%
  # 按|拆分成分列表
  mutate(ingredient_list = str_split(ingredients, "\\|")) %>%
  # 筛选出属于纤维标识的成分
  mutate(
    fibre_used = map(ingredient_list, ~keep(.x, .x %in% fibre_patterns)),
    fibre_count = map_int(fibre_used, length),
    fibre_present = fibre_count > 0
  ) %>%
  # 移除中间辅助列
  select(-ingredient_list)

运行任意一种方法,都能得到期望的输出:durian的fibre_count变为1,fibre_used为"fibre(500)"。

内容的提问来源于stack exchange,提问作者Jay Bee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 15:03:19