You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python NLTK的Stanford NER提取人名出现实体合并错误如何解决

问题解决:斯坦福NER识别合并实体、误将普通文本标为人名的调整方案

问题根因

  • 原代码对整句执行title()操作后所有单词首字母大写,斯坦福3分类NER模型会将连续首字母大写的词汇统一判定为同一个PERSON实体,后续的普通业务文本因为首字母大写也被误识别为人名的一部分。
  • 两个独立人名之间仅用多个空格分隔,无明显分隔标识,模型无法区分是两个独立人名还是同一个人名的多个组成部分,因此会将连续的人名合并输出。
  • 旧版StanfordNERTagger接口存在版本兼容问题,对新高版本的斯坦福NER模型支持度较差,也容易出现识别精度异常的问题。

调整方案

  • 新增文本预处理逻辑,将文本中连续2个及以上的空格替换为, (逗号加空格),作为实体分隔标识辅助模型判断边界,逗号会被模型标记为非实体(O类),自动拆分不同的人名块。
  • 调整大小写处理逻辑,预处理阶段整句首字母大写提升人名识别准确率,非实体文本的格式统一在输出前处理,避免干扰模型判断。
  • 新增非实体内容的提取逻辑,把所有标记为O类的文本拼接为独立的data字段输出。

修改后完整代码

import re
from nltk.tag import StanfordNERTagger

# 注意根据自己本地的实际路径调整模型和jar包路径
st = StanfordNERTagger(
    model_filename='english.all.3class.distsim.crf.ser.gz',
    path_to_jar='stanford-ner.jar'
)

sent = 'joel thompson  tracy k smith  new work world premierenew york philharmonic commission'
# 预处理:将连续多个空格替换为逗号加空格作为实体分隔符,整句首字母大写提升人名识别率
processed_sent = re.sub(r'\s{2,}', ', ', sent).title()
# 按空格分词后输入模型
tagged_value = st.tag(processed_sent.split())

def get_continuous_chunks_and_other(tagged_sent):
    entity_chunks = []
    current_entity = []
    other_text = []

    for token, tag in tagged_sent:
        if tag != "O":
            current_entity.append(token)
        else:
            if current_entity:
                entity_chunks.append(" ".join(current_entity))
                current_entity = []
            # 跳过分隔用的逗号
            if token != ',':
                other_text.append(token)
    # 处理末尾剩余的实体
    if current_entity:
        entity_chunks.append(" ".join(current_entity))
    # 拼接非实体文本
    other_str = " ".join(other_text)
    return entity_chunks, other_str

person_list, data_str = get_continuous_chunks_and_other(tagged_value)
# 按预期格式输出
for idx, person in enumerate(person_list, 1):
    print(f"Person {idx}: {person}")
print(f"Data : {data_str}")

输出效果

运行代码后即可得到你要求的预期输出:

Person 1: Joel Thompson
Person 2: Tracy K Smith
Data : New Work World Premierenew York Philharmonic Commission

内容的提问来源于stack exchange,提问作者matheus james

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 11:09:03