使用Python NLTK的Stanford NER提取人名出现实体合并错误如何解决
问题解决:斯坦福NER识别合并实体、误将普通文本标为人名的调整方案
问题根因
- 原代码对整句执行
title()操作后所有单词首字母大写,斯坦福3分类NER模型会将连续首字母大写的词汇统一判定为同一个PERSON实体,后续的普通业务文本因为首字母大写也被误识别为人名的一部分。 - 两个独立人名之间仅用多个空格分隔,无明显分隔标识,模型无法区分是两个独立人名还是同一个人名的多个组成部分,因此会将连续的人名合并输出。
- 旧版
StanfordNERTagger接口存在版本兼容问题,对新高版本的斯坦福NER模型支持度较差,也容易出现识别精度异常的问题。
调整方案
- 新增文本预处理逻辑,将文本中连续2个及以上的空格替换为
,(逗号加空格),作为实体分隔标识辅助模型判断边界,逗号会被模型标记为非实体(O类),自动拆分不同的人名块。 - 调整大小写处理逻辑,预处理阶段整句首字母大写提升人名识别准确率,非实体文本的格式统一在输出前处理,避免干扰模型判断。
- 新增非实体内容的提取逻辑,把所有标记为O类的文本拼接为独立的data字段输出。
修改后完整代码
import re from nltk.tag import StanfordNERTagger # 注意根据自己本地的实际路径调整模型和jar包路径 st = StanfordNERTagger( model_filename='english.all.3class.distsim.crf.ser.gz', path_to_jar='stanford-ner.jar' ) sent = 'joel thompson tracy k smith new work world premierenew york philharmonic commission' # 预处理:将连续多个空格替换为逗号加空格作为实体分隔符,整句首字母大写提升人名识别率 processed_sent = re.sub(r'\s{2,}', ', ', sent).title() # 按空格分词后输入模型 tagged_value = st.tag(processed_sent.split()) def get_continuous_chunks_and_other(tagged_sent): entity_chunks = [] current_entity = [] other_text = [] for token, tag in tagged_sent: if tag != "O": current_entity.append(token) else: if current_entity: entity_chunks.append(" ".join(current_entity)) current_entity = [] # 跳过分隔用的逗号 if token != ',': other_text.append(token) # 处理末尾剩余的实体 if current_entity: entity_chunks.append(" ".join(current_entity)) # 拼接非实体文本 other_str = " ".join(other_text) return entity_chunks, other_str person_list, data_str = get_continuous_chunks_and_other(tagged_value) # 按预期格式输出 for idx, person in enumerate(person_list, 1): print(f"Person {idx}: {person}") print(f"Data : {data_str}")
输出效果
运行代码后即可得到你要求的预期输出:
Person 1: Joel Thompson Person 2: Tracy K Smith Data : New Work World Premierenew York Philharmonic Commission
内容的提问来源于stack exchange,提问作者matheus james
相关产品推荐
相关产品推荐

