如何在R中基于关键词为数据框新增Group列(支持忽略大小写)
在R数据框中按关键词(忽略大小写)新增分组列
需求:给tibble创建的df数据框新增名为Group的列,根据指定关键词为每行标记分组:
- 匹配
keyword1中的任意关键词 → 标记为Politics - 匹配
keyword2中的任意关键词 → 标记为Food - 匹配
keyword3中的任意关键词 → 标记为Sport - 匹配
keyword4中的任意关键词 → 标记为Travel - 匹配
keyword5中的任意关键词 → 标记为LGBT
要求匹配时忽略大小写,此前尝试用%in%未得到预期结果(原因:%in%是精确匹配整个字符串,而非检测文本中是否包含关键词)。
示例数据与关键词
library(tibble) df <- tibble(Title = c("Iran: How we are uncovering the protests and crackdowns", "Deepak Nirula: The man who brought burgers and pizzas to India", "Phil Foden: Manchester City midfielder signs new deal with club until 2027", "The Danish tradition we all need now", "Slovakia LGBT attack"), Text = c("Iranian authorities have been disrupting the internet service in order to limit the flow of information and control the narrative, but Iranians are still sending BBC Persian videos of protests happening across the country via messaging apps. Videos are also being posted frequently on social media. Before a video can be used in any reports, journalists need to establish where and when it was filmed.They can pinpoint the location by looking for landmarks and signs in the footage and checking them against satellite images, street-level photos and previous footage. Weather reports, the position of the sun and the angles of shadows it creates can be used to confirm the timing.", "For anyone who grew up in capital Delhi during the 1970s and 1980s, Nirula's - run by the family of Deepak Nirula who died last week - is more than a restaurant. It's an emotion. The restaurant transformed the eating-out culture in the city and introduced an entire generation to fast food, American style, before McDonald's and KFC came into the country. For many it was synonymous with its hot chocolate fudge.", "Stockport-born Foden, who has scored two goals in 18 caps for England, has won 11 trophies with City, including four Premier League titles, four EFL Cups and the FA Cup.He has also won the Premier League Young Player of the Season and PFA Young Player of the Year awards in each of the last two seasons. City boss Pep Guardiola handed him his debut as a 17-year-old and Foden credited the Spaniard for his impressive development over the last five years.", "Norwegian playwright and poet Henrik Ibsen popularised the term /friluftsliv/ in the 1850s to describe the value of spending time in remote locations for spiritual and physical wellbeing. It literally translates to /open-air living/, and today, Scandinavians value connecting to nature in different ways – something we all need right now as we emerge from an era of lockdowns and inactivity.", "The men were shot dead in the capital Bratislava on Wednesday, in a suspected hate crime.Organisers estimated that 20,000 people took part in the vigil, mourning the men's deaths and demanding action on LGBT rights.Slovak President Zuzana Caputova, who has raised the rainbow flag over her office, spoke at the event.") ) keyword1 <- c("authorities", "Iranian", "Iraq", "control", "Riots") keyword2 <- c("McDonald's","KFC", "McCafé", "fast food") keyword3 <- c("caps", "trophies", "season", "seasons") keyword4 <- c("travel", "landscape", "living", "spiritual") keyword5 <- c("LGBT", "lesbian", "les", "rainbow", "Gay", "Bisexual","Transgender")
解决方案
使用dplyr的mutate+case_when,结合stringr的str_detect实现关键词检测(支持忽略大小写):
library(dplyr) library(stringr) # 给每行标记分组 df <- df %>% mutate(Group = case_when( # 合并标题与文本,检测是否包含对应关键词组的任意词,忽略大小写 str_detect(paste(Title, Text), regex(paste(keyword1, collapse = "|"), ignore_case = TRUE)) ~ "Politics", str_detect(paste(Title, Text), regex(paste(keyword2, collapse = "|"), ignore_case = TRUE)) ~ "Food", str_detect(paste(Title, Text), regex(paste(keyword3, collapse = "|"), ignore_case = TRUE)) ~ "Sport", str_detect(paste(Title, Text), regex(paste(keyword4, collapse = "|"), ignore_case = TRUE)) ~ "Travel", str_detect(paste(Title, Text), regex(paste(keyword5, collapse = "|"), ignore_case = TRUE)) ~ "LGBT", # 无匹配时的默认分组(可选) TRUE ~ "Unclassified" ))
结果验证
查看新增的Group列:
df$Group # 输出:[1] "Politics" "Food" "Sport" "Travel" "LGBT"
完全符合预期输出。
关键说明
paste(Title, Text):合并标题和文本字段,统一进行关键词检测regex(paste(keywords, collapse = "|"), ignore_case = TRUE):将关键词组拼接成正则表达式(|表示“或”逻辑),ignore_case=TRUE实现大小写不敏感匹配case_when按顺序判断分组,若某行同时匹配多个关键词组,会优先匹配排在前面的规则
内容的提问来源于stack exchange,提问作者Alegría
相关产品推荐
相关产品推荐

