You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Spacy NER标注的实体格式转换为CONLL 2003格式

格式转换实现方法

你当前使用的是spaCy生态通用的NER标注格式,可以通过以下步骤转换为CONLL 2003格式:

首先明确CONLL 2003格式规范

CONLL 2003采用BIO标注体系,每行对应1个分词结果,标准结构为4列,用空格分隔:

  • 第1列:分词得到的单词
  • 第2列:词性标注(如果不需要可以统一填占位符-)
  • 第3列:句法块标注(如果不需要可以统一填占位符-)
  • 第4列:NER标签,B-<实体类型>表示实体起始词,I-<实体类型>表示实体后续词,O表示非实体
    不同句子之间用空行分隔。

转换代码示例

这里提供可直接运行的Python实现,依赖nltk做分词,提前运行pip install nltk安装依赖后即可使用:

import nltk
nltk.download('punkt')
from nltk.tokenize import word_tokenize

# 你的原始标注数据
raw_data = [
    ('The F15 aircraft uses a lot of fuel', {'entities': [(4, 7, 'aircraft')]}),
    ('did you see the F16 landing?', {'entities': [(16, 19, 'aircraft')]}),
    ('how many missiles can a F35 carry', {'entities': [(24, 27, 'aircraft')]}),
    ('is the F15 outdated', {'entities': [(7, 10, 'aircraft')]}),
    ('how long does it take to train a F16 pilot',{'entities': [(33, 36, 'aircraft')]}),
    ('how much does a F35 cost', {'entities': [(16, 19, 'aircraft')]})
]

conll_lines = []
for text, annot in raw_data:
    entities = annot['entities']
    # 先获取每个token的偏移位置
    tokens = word_tokenize(text)
    # 记录每个token的起始、结束下标
    token_spans = []
    current_idx = 0
    for token in tokens:
        start = text.find(token, current_idx)
        end = start + len(token)
        token_spans.append((start, end, token))
        current_idx = end
    # 给每个token打标签
    for start, end, token in token_spans:
        label = 'O'
        for ent_start, ent_end, ent_type in entities:
            # 判断当前token是否在实体范围内
            if start >= ent_start and end <= ent_end:
                if start == ent_start:
                    label = f'B-{ent_type}'
                else:
                    label = f'I-{ent_type}'
                break
        # 词性和块标签用-占位
        conll_lines.append(f"{token} - - {label}")
    # 句子之间加空行
    conll_lines.append("")

# 输出结果或者写入文件
conll_result = "\n".join(conll_lines)
print(conll_result)
# 如果要保存成文件就运行下面这行
# with open("output.conll", "w", encoding="utf-8") as f:
#     f.write(conll_result)

转换后的效果示例

以上代码运行后输出的CONLL格式内容示例如下:

The - - O
F15 - - B-aircraft
aircraft - - O
uses - - O
a - - O
lot - - O
of - - O
fuel - - O

did - - O
you - - O
see - - O
the - - O
F16 - - B-aircraft
landing - - O
? - - O

你可以根据自己的需求调整分词工具、是否保留词性/块标签列即可。

内容的提问来源于stack exchange,提问作者imhans33

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 06:06:03