You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Doccano导出的JSONL格式转换为spaCy格式用于NER训练?

将Doccano JSONL格式NER数据集转换为spaCy兼容格式

给你两种直接可用的转换方案,适配spaCy的训练需求:

方案1:转换为spaCy训练元组格式(适合小数据集)

这种格式是(文本内容, {"entities": [(起始偏移, 结束偏移, 实体标签)]})的元组列表,可直接用于自定义训练循环,或者导出为JSON文件备用。

import json

def convert_doccano_to_spacy_tuple(input_file, output_file):
    training_data = []
    with open(input_file, 'r', encoding='utf-8') as f:
        for line in f:
            # 逐行解析JSONL数据
            item = json.loads(line.strip())
            text = item['text']
            entities = []
            # 提取实体的偏移量和标签
            for ent in item['entities']:
                entities.append((ent['start_offset'], ent['end_offset'], ent['label']))
            training_data.append((text, {"entities": entities}))
    
    # 保存为JSON格式,方便后续加载
    with open(output_file, 'w', encoding='utf-8') as f:
        json.dump(training_data, f, ensure_ascii=False, indent=2)

# 替换成你的文件路径
convert_doccano_to_spacy_tuple("your_doccano_data.jsonl", "spacy_train_data.json")

方案2:转换为DocBin格式(推荐,适合大数据集)

这是spaCy官方推荐的高效格式,加载速度快,占用空间小,适合大规模数据集训练。

import json
import spacy
from spacy.tokens import DocBin

def convert_doccano_to_docbin(input_file, output_file):
    # 对应你的数据集语言,中文用"zh",英文用"en"
    nlp = spacy.blank("en")
    doc_bin = DocBin()
    
    with open(input_file, 'r', encoding='utf-8') as f:
        for line in f:
            item = json.loads(line.strip())
            text = item['text']
            # 创建spaCy的Doc对象
            doc = nlp.make_doc(text)
            ents = []
            for ent in item['entities']:
                # 从字符偏移量创建实体span
                span = doc.char_span(ent['start_offset'], ent['end_offset'], label=ent['label'])
                if span:
                    ents.append(span)
                else:
                    # 如果偏移量不匹配,打印警告(比如文本有转义问题时)
                    print(f"警告:文本位置 {ent['start_offset']}-{ent['end_offset']} 在当前文本中无效")
            doc.ents = ents
            doc_bin.add(doc)
    
    # 保存为.spacy文件
    doc_bin.to_disk(output_file)

# 替换成你的文件路径
convert_doccano_to_docbin("your_doccano_data.jsonl", "spacy_train_data.spacy")

注意事项

  • 运行代码前确保安装了spaCy:pip install spacy
  • 如果是中文数据集,把代码里的spacy.blank("en")改成spacy.blank("zh")
  • 若出现偏移量不匹配的警告,检查Doccano导出时的文本是否有特殊转义问题(比如换行符、斜杠),代码会自动跳过无效实体,不影响其他数据转换

内容的提问来源于stack exchange,提问作者Blue

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 23:25:55