如何在R中从板球解说文本匹配多词关键词
问题描述
我需要从板球解说文本中提取特定的多词关键词(2-3个单词的词组),指定的关键词列表如下:
region <- c("third man", "deep fine leg", "long leg", "deep square leg", "Deep mid wicket", "cow corner", "long on", "Deep extra cover", "Deep Cover", "Deep point", "Deep backword point", "fly slip", "backword point", "point", "cover", "Extra covers", "mid off", "mid on", "mid wicket", "square leg", "backword square leg", "fine leg", "slips", "gully", "silly point", "silly mid off", "silly mid on", "short leg", "leg gully", "leg slip")
同时有三段板球解说示例文本:
Pretorius to Umesh Yadav, 1 run, pitched up by Pretorius, touch slower as it has been driven along the ground to long-off
Pretorius to Chahar, SIX, that's a great shot. Pitched up by Pretorius outside off, a slower one and Chahar goes down on his knee and plays a fantastic lofted shot to clear the boundary at deep extra cover
Pretorius to Umesh Yadav, 1 run, touch fuller on off, Umesh Yadav drills it to long-off for a single
我使用的是R 4.2.1版本及RStudio,请问如何实现从解说文本中匹配这些多词关键词,并提取出匹配到的内容?
解决方案
解说文本存在大小写不一致、连字符替换空格的格式差异,先统一格式再匹配,推荐使用stringr包处理文本:
1. 安装并加载依赖包
install.packages("stringr") library(stringr)
2. 准备原始数据
# 原始解说文本 commentary <- c( "Pretorius to Umesh Yadav, 1 run, pitched up by Pretorius, touch slower as it has been driven along the ground to long-off", "Pretorius to Chahar, SIX, that's a great shot. Pitched up by Pretorius outside off, a slower one and Chahar goes down on his knee and plays a fantastic lofted shot to clear the boundary at deep extra cover", "Pretorius to Umesh Yadav, 1 run, touch fuller on off, Umesh Yadav drills it to long-off for a single" ) # 原始关键词列表 region <- c("third man", "deep fine leg", "long leg", "deep square leg", "Deep mid wicket", "cow corner", "long on", "Deep extra cover", "Deep Cover", "Deep point", "Deep backword point", "fly slip", "backword point", "point", "cover", "Extra covers", "mid off", "mid on", "mid wicket", "square leg", "backword square leg", "fine leg", "slips", "gully", "silly point", "silly mid off", "silly mid on", "short leg", "leg gully", "leg slip")
3. 统一文本格式
消除大小写、连字符的格式差异:
# 处理关键词:转小写,生成正则匹配模式(用|分隔多个关键词) region_clean <- str_to_lower(region) pattern <- str_c(region_clean, collapse = "|") # 处理解说文本:替换连字符为空格,转小写 commentary_clean <- str_replace_all(str_to_lower(commentary), "-", " ")
4. 提取匹配的关键词
用str_extract_all提取所有匹配内容:
# 提取每个文本中匹配的关键词 matches <- str_extract_all(commentary_clean, pattern) # 查看结果 matches
运行结果
返回一个列表,每个元素对应一段文本的匹配结果:
[[1]] [1] "long on" [[2]] [1] "deep extra cover" [[3]] [1] "long on"
可选:整理为数据框
如果需要结构化输出,可转为数据框:
library(tibble) tibble( 解说文本 = commentary, 匹配到的区域 = matches )
内容的提问来源于stack exchange,提问作者Chandrshekar N

