使用set.intersection()提取关键词句无法匹配多词关键词如何解决
问题原因
你当前代码匹配失败的核心原因是sentence.words返回的是经过分词处理的单个单词集合,集合内的元素都是拆分后的独立词汇,比如"off"、"the"、"road"会被拆成三个独立元素存储,不存在"off the road"、"blue tinge"这类连续多词组成的字符串元素。
当你用set.intersection()做交集计算时,只有单个词形式的关键词"van"能在单词集合里找到对应元素,多词关键词永远无法匹配到集合内的单个词元素,自然不会被识别。
修改方法
不需要把句子拆分为单个单词集合做交集运算,直接遍历关键词判断是否存在于句子文本中即可,支持任意长度的多词匹配。如果需要避免大小写导致的匹配失败,可以提前把句子和关键词统一转为小写再判断;如果需要避免子串误匹配(比如关键词"van"误匹配到"vanguard"的片段),可以搭配正则的单词边界做精确匹配。
修改后的可运行代码如下:
from textblob import TextBlob import nltk nltk.download('punkt') search_words = {"off the road", "blue tinge", "van"} blob = TextBlob("That is the off the road vehicle I had in mind for my adventure. Which one? The one with the blue tinge. Oh, I'd use the money for a van.") matches = [] for sentence in blob.sentences: sentence_text = str(sentence).lower() # 只要句子包含任意一个目标关键词就加入结果列表 if any(keyword.lower() in sentence_text for keyword in search_words): matches.append(str(sentence)) print(matches)
运行后输出结果:
['That is the off the road vehicle I had in mind for my adventure.', 'The one with the blue tinge.', "Oh, I'd use the money for a van."]
内容的提问来源于stack exchange,提问作者CTan
相关产品推荐
相关产品推荐

