You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取解析CSV后调用spaCy模型输出异常问题排查

解决spaCy实体识别CSV数据的三类异常问题

问题背景

需求:用Python读取Dog_Breed.csv文件,调用训练好的spaCy模型(models/output/model-best)识别BREED(犬种)和ORIGIN(原产地)实体,按行输出格式:序号,犬种 is a BREED and has an ORIGIN of 原产地

遇到三类异常:

  • 表头行,Dog Breed,Origin被错误识别为实体
  • 实体“South Carolina”被截断,但“Boykin Spaniel”识别正常
  • 第1、2行出现无效的has an ORIGIN of ,输出行

对应解决方案

1. 跳过表头避免错误识别

读取CSV时直接跳过表头行,不将表头传入spaCy模型处理:

import csv
import spacy

# 加载模型
nlp = spacy.load("models/output/model-best")

with open("Dog_Breed.csv", "r", encoding="utf-8") as f:
    csv_reader = csv.reader(f)
    next(csv_reader)  # 跳过表头
    for line_num, row in enumerate(csv_reader, start=1):
        # 将行数据转为字符串传入模型
        doc = nlp(",".join(row))
        # 后续实体提取逻辑

2. 修复多词实体截断问题

多词实体截断是模型训练数据不足导致的,优先方案是补充训练数据:

  • 新增包含完整“South Carolina”作为ORIGIN实体的标注样本
  • 重新训练模型,确保实体边界标注准确

临时应急修复(无需重新训练):
检查相邻的同类型实体,合并连续的ORIGIN实体:

def get_full_origins(doc):
    origins = []
    i = 0
    while i < len(doc.ents):
        ent = doc.ents[i]
        if ent.label_ == "ORIGIN":
            # 检查下一个实体是否为连续的ORIGIN
            if i + 1 < len(doc.ents) and doc.ents[i+1].label_ == "ORIGIN" and doc.ents[i+1].start == ent.end + 1:
                origins.append(f"{ent.text} {doc.ents[i+1].text}")
                i += 2
                continue
            origins.append(ent.text)
        i += 1
    return origins

3. 过滤无效实体输出

添加实体存在性判断,确保只有同时识别到BREED和ORIGIN时才输出有效内容:

# 提取实体
breeds = [ent.text for ent in doc.ents if ent.label_ == "BREED"]
origins = get_full_origins(doc)

# 输出有效内容
if breeds and origins:
    print(f"{line_num},{breeds[0]} is a BREED and has an ORIGIN of {origins[0]}")
else:
    # 可选:跳过或标记无效行
    print(f"{line_num},无有效实体匹配")

内容的提问来源于stack exchange,提问作者kameron cole

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 07:35:26