如何解决文本尺寸提取中单双词尺寸的匹配冲突问题?
解决文本中尺寸提取的遗漏问题
问题背景
现有固定尺寸列表:
sizes = ['extra small', 'small', 'medium', 'large', 'extra large']
需求是从文本中提取所有被提及的尺寸项。现有代码采用先匹配双词尺寸、再匹配单词尺寸的思路,能正确处理text1 = 'she wears a small size and he wears an extra large size',提取出['extra large', 'small'],但处理text2 = 'she wears a large size and he wears an extra large size'时,仅提取出['extra large'],漏掉了单独提及的large。
问题原因
原有代码判断单字词尺寸是否加入结果时,用了x not in [item for sublist in [x.split() for x in mentioned_sizes] for item in sublist]这个条件——只要单字词出现在已提取的多字词拆分结果里,就直接排除。但在text2场景中,large既作为extra large的一部分出现,也单独被提及,这个条件错误地把单独的large排除了。
修改方案
核心思路:先提取所有多字词尺寸,再将这些多字词从文本中临时移除,最后在剩余文本中提取单字词尺寸,避免多字词里的单字词干扰单独匹配。
修改后的代码:
import re sizes = ['extra small', 'small', 'medium', 'large', 'extra large'] text2 = 'she wears a large size and he wears an extra large size' mentioned_sizes = [] # 按词数从多到少排序,优先匹配完整多字词 sizes_sorted = sorted(sizes, key=lambda x: len(x.split()), reverse=True) # 第一步:提取所有多字词尺寸 multi_word_sizes = [s for s in sizes_sorted if len(s.split()) > 1] for size in multi_word_sizes: # 用单词边界\b确保匹配完整的多字词,避免部分匹配 if re.search(r'\b' + re.escape(size) + r'\b', text2): mentioned_sizes.append(size) # 第二步:创建临时文本,移除已匹配的多字词,消除干扰 temp_text = text2 for size in mentioned_sizes: temp_text = re.sub(r'\b' + re.escape(size) + r'\b', '', temp_text) # 第三步:提取剩余文本中的单字词尺寸 single_word_sizes = [s for s in sizes_sorted if len(s.split()) == 1] for size in single_word_sizes: if re.search(r'\b' + re.escape(size) + r'\b', temp_text): mentioned_sizes.append(size) # 去重并保留原有顺序(可选,根据需求调整) mentioned_sizes = list(dict.fromkeys(mentioned_sizes)) print(mentioned_sizes) # 输出: ['extra large', 'large']
代码说明
- 排序逻辑:保持按词数从多到少排序,确保先匹配
extra large这类多字词,避免先匹配large导致多字词拆分失效。 - 单词边界匹配:用
\b正则语法确保匹配完整的尺寸短语,避免出现“small”匹配“smaller”这类错误。 - 临时文本处理:移除已匹配的多字词后,剩余文本中仅保留单独提及的单字词实例,避免误判。
- 去重处理:用
dict.fromkeys实现去重并保留原有顺序,防止同一尺寸被重复添加。
内容的提问来源于stack exchange,提问作者wanderingcatto
相关产品推荐
相关产品推荐

