You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NER模型数据转spaCy格式函数问题求助:模型测试无输出

NER模型数据转spaCy格式函数问题求助:模型测试无输出

嘿,我仔细看了你写的convert_to_spacy_format函数,模型测试没输出大概率是数据转换过程中实体没被正确标记出来,咱们来一步步揪出问题所在:

几个关键问题点

  • I-标签处理逻辑缺失:目前遇到I-标签直接pass,既没验证当前实体是否和I-后缀匹配,也没更新实体的结束位置,导致连续的实体片段(比如B-PER+I-PER)无法被完整拼接,最终实体范围错误。
  • O标签的实体结束位置错误:你现在用token.idx作为实体结束位置,但这是当前O标签token的起始位置,正确的实体结束位置应该是前一个实体token的结束位置(也就是前一个token的idx + len(token))。
  • Tokenization不匹配:用sentence.split()拆分单词和spaCy自带的分词逻辑不一致(比如标点会被拆成单独token),这会导致zip(doc, ner_tags)时token和标签对应错位,很多数据被continue跳过,训练数据量不足甚至无效。
  • 最后一个实体的结束位置不准确:用len(sentence)作为最后一个实体的结束位置,会包含句子末尾可能存在的空格,应该用最后一个实体token的idx + len(token)来准确定位。

修改后的完整函数

# 记得先导入需要的模块
from spacy.tokens import DocBin
from spacy.util import filter_spans
from tqdm import tqdm
import spacy

def convert_to_spacy_format(data):
    nlp = spacy.blank("en")  
    db = DocBin()  
    
    for _, row in tqdm(data.iterrows(), total=len(data)):
        sentence = row["CleanSentence"]
        pos_tags = row["POS"]
        ner_tags = row["Tag"]
        
        # 用spaCy的分词器得到标准token列表,避免split()的差异
        doc = nlp.make_doc(sentence)
        words = [token.text for token in doc]
        
        # 检查长度匹配(用spaCy分词后的words,而不是split()的结果)
        if len(words) != len(ner_tags) or len(words) != len(pos_tags):
            print(f"Warning: Length mismatch: Words: {len(words)}, NER tags: {len(ner_tags)}, POS tags: {len(pos_tags)}")
            continue
            
        ents = []
        current_ent = None
        current_ent_start = None
        prev_token = None  # 新增变量,跟踪实体的最后一个token
        
        # 遍历每个token和对应的标签
        for idx, (token, tag) in enumerate(zip(doc, ner_tags)):
            prev_token = token  # 更新当前token为前一个token
            if tag.startswith("B-"):
                # 如果之前有未闭合的实体,先添加到列表
                if current_ent is not None:
                    ents.append((current_ent_start, prev_token.idx + len(prev_token), current_ent))
                
                # 开始新实体的跟踪
                current_ent = tag[2:]
                current_ent_start = token.idx
                
            elif tag.startswith("I-"):
                # 验证当前I-标签是否和正在跟踪的实体一致,避免标签错误
                if current_ent is not None and tag[2:] == current_ent:
                    # 继续跟踪当前实体,更新最后一个token
                    pass
                else:
                    # 如果标签不匹配,视为新的实体起始(或者根据你的数据集规则调整)
                    if current_ent is not None:
                        ents.append((current_ent_start, prev_token.idx + len(prev_token), current_ent))
                    current_ent = tag[2:]
                    current_ent_start = token.idx
                    
            elif tag == "O":
                # 闭合之前的实体
                if current_ent is not None:
                    ents.append((current_ent_start, prev_token.idx + len(prev_token), current_ent))
                    current_ent = None
                    current_ent_start = None
        
        # 处理循环结束后剩余的未闭合实体
        if current_ent is not None and prev_token is not None:
            ents.append((current_ent_start, prev_token.idx + len(prev_token), current_ent))
        
        # 创建实体span并过滤重叠
        spans = []
        for start, end, label in ents:
            span = doc.char_span(start, end, label=label)
            if span is not None:
                spans.append(span)
        
        filtered_spans = filter_spans(spans)
        doc.ents = filtered_spans
        
        db.add(doc)
    
    return db

额外小建议

  • 转换完成后,可以取出几个Doc对象打印doc.ents,确认实体是否被正确标记,比如:
    # 测试转换结果
    db = convert_to_spacy_format(your_data)
    docs = list(db.get_docs(nlp.vocab))
    print(docs[0].ents)
    
  • 如果数据集里存在标签格式不规范的情况(比如I-开头但前面没有对应的B-),可以在预处理阶段先清洗标签,避免实体识别混乱。

备注:内容来源于stack exchange,提问作者Rohit Gupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 19:35:27