基于Excel的Cliche Phrases匹配与类型返回技术问题求助
问题解决方案
一、解决匹配结果缺失问题
匹配结果缺失大概率是因为Matcher的token模式匹配不适配短语精确匹配场景,改用PhraseMatcher能更稳定实现精确匹配,避免因token属性设置或标点干扰导致漏匹配:
import spacy from spacy.matcher import PhraseMatcher nlp = spacy.load('en_core_web_sm') # 初始化PhraseMatcher,基于小写匹配忽略大小写 matcher = PhraseMatcher(nlp.vocab, attr="LOWER") cliches = ['Abandon ship', 'About face', 'Above board', 'All ears'] # 将短语转为Doc对象加入匹配器 cliche_docs = [nlp(cliche) for cliche in cliches] matcher.add("Cliche", None, *cliche_docs) # 测试文本 example_2 = nlp("We must abandon ship! It's the only way to stay above board. And I'm all ears.") matches_2 = matcher(example_2) for match_id, start, end in matches_2: print(example_2[start:end])
二、Excel关联与结果输出、短语标红
1. 读取Excel并建立短语-Type映射
用pandas读取Excel,生成短语到对应type的映射字典:
import pandas as pd # 读取Excel文件,替换为你的实际文件名 df = pd.read_excel("cliches_db.xlsx") # 构建映射字典,统一小写键避免大小写匹配误差 cliche_type_map = {row["Cliche Phrases"].lower(): row["type"] for _, row in df.iterrows()}
2. 完整整合:匹配结果保存+短语标红
以下代码同时实现匹配结果保存、原文档短语标红(支持HTML格式标红和终端标红两种方式):
import spacy from spacy.matcher import PhraseMatcher import pandas as pd # 初始化nlp和匹配器 nlp = spacy.load('en_core_web_sm') matcher = PhraseMatcher(nlp.vocab, attr="LOWER") # 读取Excel数据库 df = pd.read_excel("cliches_db.xlsx") cliche_list = df["Cliche Phrases"].tolist() cliche_type_map = {phrase.lower(): df.loc[df["Cliche Phrases"] == phrase, "type"].values[0] for phrase in cliche_list} # 添加短语到匹配器 cliche_docs = [nlp(phrase) for phrase in cliche_list] matcher.add("Cliche", None, *cliche_docs) # 目标处理文本,替换为你的实际文档内容 target_text = "We must abandon ship! It's the only way to stay above board. About face, I'm all ears." doc = nlp(target_text) # 收集匹配结果 match_results = [] # 准备标红后的HTML文本 reddened_html_text = list(target_text) # 倒序处理匹配结果,避免插入标签后索引偏移 matches_sorted = sorted(matcher(doc), key=lambda x: x[2], reverse=True) for match_id, start, end in matches_sorted: matched_phrase = doc[start:end].text matched_type = cliche_type_map.get(matched_phrase.lower()) match_results.append(f"短语: {matched_phrase}, 类型: {matched_type}") # 插入HTML红标签 reddened_html_text.insert(end, "</span>") reddened_html_text.insert(start, "<span style='color:red'>") # 保存匹配结果到文本文件 with open("cliche_matches.txt", "w", encoding="utf-8") as f: f.write("\n".join(match_results)) # 保存标红后的HTML文件(可直接用浏览器打开查看) with open("reddened_text.html", "w", encoding="utf-8") as f: f.write("".join(reddened_html_text)) # 控制台打印匹配结果 for result in match_results: print(result)
关键说明
- 用
PhraseMatcher替代Matcher:直接基于短语的Doc对象匹配,避免手动构建token模式的误差,更适合精确短语检测 - 统一小写映射键:确保匹配不受原文本大小写影响
- 倒序处理匹配:避免插入标红标签后,后续匹配的索引位置偏移导致标红错误
内容的提问来源于stack exchange,提问作者Programmer_nltk
相关产品推荐
相关产品推荐

