如何用Python的NLP技术从报告中提取指定人物的行为?
修复提取指定人物行为动作的Spacy代码
原代码无法正确提取John的行为,核心问题是依赖关系判断错误,以及未覆盖被动语态等常见句式,以下是修正方案:
原代码的问题
- 依赖关系搞反:当John是句子主语时,根动词(ROOT)是John的父节点,而非子节点。原代码遍历John的子节点找ROOT,自然找不到正确动作。
- 未处理被动语态:比如"John was assigned a task"这类被动句,原代码无法识别John作为动作相关方的行为。
- 仅提取单个动词:丢失了动作的完整上下文(比如只取"submitted",而不是"submitted the quarterly report")。
修正后的代码
import re import spacy # 加载语言模型 nlp = spacy.load("en_core_web_sm") # 指定要搜索的人物 person = "John" # 正则匹配人物名称(支持大小写不敏感) pattern = re.compile(fr"\b{person}\b", re.IGNORECASE) # 替换为你的大型文本文件内容 text = """ John submitted the quarterly report on Monday. The project was approved by John last week. John is preparing for the upcoming meeting. Mary and John reviewed the client feedback together. """ # 处理文本 doc = nlp(text) actions = [] for sent in doc.sents: for token in sent: # 匹配目标人物 if pattern.match(token.text): # 情况1:主动语态,人物是主语(nsubj) if token.dep_ == "nsubj": root_verb = token.head # 提取完整动作短语(包含动词、宾语等) action_phrase = " ".join([t.text for t in root_verb.subtree]) actions.append(action_phrase) # 情况2:被动语态,人物是被动主语(nsubjpass) elif token.dep_ == "nsubjpass": root_verb = token.head action_phrase = " ".join([t.text for t in root_verb.subtree]) actions.append(action_phrase) # 情况3:人物在"by"引导的施动者短语中 elif token.dep_ == "pobj" and token.head.text.lower() == "by": root_verb = token.head.head action_phrase = " ".join([t.text for t in root_verb.subtree]) actions.append(action_phrase) # 去重并输出结果 unique_actions = list(set(actions)) if unique_actions: print(f"{person} 的行为动作:") for action in unique_actions: print(f"- {action}") else: print(f"未找到 {person} 的相关行为。")
关键改进说明
- 修正依赖关系:主动语态下,主语的父节点就是句子的核心动词,直接通过
token.head获取动作根。 - 覆盖被动场景:处理被动主语(
nsubjpass)和"by"引导的施动者两种被动结构,确保不遗漏相关动作。 - 保留完整动作上下文:通过
root_verb.subtree获取整个动词相关的短语,让动作描述更完整。 - 结果去重:用集合去除重复的动作记录,避免冗余输出。
内容的提问来源于stack exchange,提问作者Fateme Nazari
相关产品推荐
相关产品推荐

