You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Excel的Cliche Phrases匹配与类型返回技术问题求助

问题解决方案

一、解决匹配结果缺失问题

匹配结果缺失大概率是因为Matcher的token模式匹配不适配短语精确匹配场景,改用PhraseMatcher能更稳定实现精确匹配,避免因token属性设置或标点干扰导致漏匹配:

import spacy
from spacy.matcher import PhraseMatcher

nlp = spacy.load('en_core_web_sm')
# 初始化PhraseMatcher,基于小写匹配忽略大小写
matcher = PhraseMatcher(nlp.vocab, attr="LOWER")

cliches = ['Abandon ship', 'About face', 'Above board', 'All ears']
# 将短语转为Doc对象加入匹配器
cliche_docs = [nlp(cliche) for cliche in cliches]
matcher.add("Cliche", None, *cliche_docs)

# 测试文本
example_2 = nlp("We must abandon ship! It's the only way to stay above board. And I'm all ears.")
matches_2 = matcher(example_2)
for match_id, start, end in matches_2:
    print(example_2[start:end])

二、Excel关联与结果输出、短语标红

1. 读取Excel并建立短语-Type映射

用pandas读取Excel,生成短语到对应type的映射字典:

import pandas as pd

# 读取Excel文件,替换为你的实际文件名
df = pd.read_excel("cliches_db.xlsx")
# 构建映射字典,统一小写键避免大小写匹配误差
cliche_type_map = {row["Cliche Phrases"].lower(): row["type"] for _, row in df.iterrows()}

2. 完整整合:匹配结果保存+短语标红

以下代码同时实现匹配结果保存、原文档短语标红(支持HTML格式标红和终端标红两种方式):

import spacy
from spacy.matcher import PhraseMatcher
import pandas as pd

# 初始化nlp和匹配器
nlp = spacy.load('en_core_web_sm')
matcher = PhraseMatcher(nlp.vocab, attr="LOWER")

# 读取Excel数据库
df = pd.read_excel("cliches_db.xlsx")
cliche_list = df["Cliche Phrases"].tolist()
cliche_type_map = {phrase.lower(): df.loc[df["Cliche Phrases"] == phrase, "type"].values[0] for phrase in cliche_list}

# 添加短语到匹配器
cliche_docs = [nlp(phrase) for phrase in cliche_list]
matcher.add("Cliche", None, *cliche_docs)

# 目标处理文本,替换为你的实际文档内容
target_text = "We must abandon ship! It's the only way to stay above board. About face, I'm all ears."
doc = nlp(target_text)

# 收集匹配结果
match_results = []
# 准备标红后的HTML文本
reddened_html_text = list(target_text)
# 倒序处理匹配结果,避免插入标签后索引偏移
matches_sorted = sorted(matcher(doc), key=lambda x: x[2], reverse=True)

for match_id, start, end in matches_sorted:
    matched_phrase = doc[start:end].text
    matched_type = cliche_type_map.get(matched_phrase.lower())
    match_results.append(f"短语: {matched_phrase}, 类型: {matched_type}")
    # 插入HTML红标签
    reddened_html_text.insert(end, "</span>")
    reddened_html_text.insert(start, "<span style='color:red'>")

# 保存匹配结果到文本文件
with open("cliche_matches.txt", "w", encoding="utf-8") as f:
    f.write("\n".join(match_results))

# 保存标红后的HTML文件(可直接用浏览器打开查看)
with open("reddened_text.html", "w", encoding="utf-8") as f:
    f.write("".join(reddened_html_text))

# 控制台打印匹配结果
for result in match_results:
    print(result)

关键说明

  • 用PhraseMatcher替代Matcher:直接基于短语的Doc对象匹配,避免手动构建token模式的误差,更适合精确短语检测
  • 统一小写映射键:确保匹配不受原文本大小写影响
  • 倒序处理匹配:避免插入标红标签后,后续匹配的索引位置偏移导致标红错误

内容的提问来源于stack exchange,提问作者Programmer_nltk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 15:11:27