如何调整str_extract_all代码避免纤维标识重复计数?
问题描述
我有如下测试数据框:
df <- data.frame( product = c("apple", "banana", "cherry", "durian", "eggplant", "fuyu"), ingredients = c("flour|fibre|500", "sugar|500", "505|wheat|flavouring", "fibre(500)|eggs", "wholegrainrice|sesameoil", "500|fibre|500"), stringsAsFactors = FALSE )
我的目标是检测产品成分中是否包含纤维、统计出现次数并提取记录纤维的标识。本次分析中,纤维可表示为"fibre"、"500"或"fibre(500)"。
当前使用的代码如下:
library(tidyverse) fibre_strings_to_check <- c("fibre", "500", "fibre\\(500\\)") df2 <- df %>% mutate( fibre_present = str_detect(ingredients, paste(fibre_strings_to_check, collapse = "|")), fibre_count = str_count(ingredients, paste(fibre_strings_to_check, collapse = "|")), fibre_used = str_extract_all(ingredients, paste(fibre_strings_to_check, collapse = "|")) )
运行后得到的输出:
| product | ingredients | fibre_present | fibre_count | fibre_used |
|---|---|---|---|---|
| apple | flour | fibre | 500 | TRUE |
| banana | sugar | 500 | TRUE | 1 |
| cherry | 505 | wheat | flavouring | FALSE |
| durian | fibre(500) | eggs | TRUE | 2 |
| eggplant | wholegrainrice | sesameoil | FALSE | 0 |
| fuyu | 500 | fibre | 500 | TRUE |
遇到的问题:durian产品的"fibre(500)"被拆分为"fibre"和"500"两次计数,我希望它被计为1次纤维标识,期望输出如下:
| product | ingredients | fibre_present | fibre_count | fibre_used |
|---|---|---|---|---|
| apple | flour | fibre | 500 | TRUE |
| banana | sugar | 500 | TRUE | 1 |
| cherry | 505 | wheat | flavouring | FALSE |
| durian | fibre(500) | eggs | TRUE | 1 |
| eggplant | wholegrainrice | sesameoil | FALSE | 0 |
| fuyu | 500 | fibre | 500 | TRUE |
解决方案
问题核心是正则匹配的优先级:原代码把短模式放在前面,导致"fibre(500)"被拆分为两个子串匹配。以下两种方法可解决该问题:
方法1:调整正则匹配顺序
将更具体、更长的模式放在匹配列表最前面,让正则引擎优先匹配完整的"fibre(500)",再匹配其他短模式:
library(tidyverse) # 优先匹配完整的"fibre(500)",再匹配其他模式 fibre_strings_to_check <- c("fibre\\(500\\)", "fibre", "500") df2 <- df %>% mutate( fibre_present = str_detect(ingredients, paste(fibre_strings_to_check, collapse = "|")), fibre_count = str_count(ingredients, paste(fibre_strings_to_check, collapse = "|")), fibre_used = str_extract_all(ingredients, paste(fibre_strings_to_check, collapse = "|")) )
方法2:拆分成分后匹配(更精准)
由于成分是用|分隔的独立项,先拆分成分列表再逐一匹配,可彻底避免部分匹配的问题:
library(tidyverse) fibre_patterns <- c("fibre", "500", "fibre\\(500\\)") df2 <- df %>% # 按|拆分成分列表 mutate(ingredient_list = str_split(ingredients, "\\|")) %>% # 筛选出属于纤维标识的成分 mutate( fibre_used = map(ingredient_list, ~keep(.x, .x %in% fibre_patterns)), fibre_count = map_int(fibre_used, length), fibre_present = fibre_count > 0 ) %>% # 移除中间辅助列 select(-ingredient_list)
运行任意一种方法,都能得到期望的输出:durian的fibre_count变为1,fibre_used为"fibre(500)"。
内容的提问来源于stack exchange,提问作者Jay Bee
相关产品推荐
相关产品推荐

