如何将Spacy NER标注的实体格式转换为CONLL 2003格式
格式转换实现方法
你当前使用的是spaCy生态通用的NER标注格式,可以通过以下步骤转换为CONLL 2003格式:
首先明确CONLL 2003格式规范
CONLL 2003采用BIO标注体系,每行对应1个分词结果,标准结构为4列,用空格分隔:
- 第1列:分词得到的单词
- 第2列:词性标注(如果不需要可以统一填占位符
-) - 第3列:句法块标注(如果不需要可以统一填占位符
-) - 第4列:NER标签,
B-<实体类型>表示实体起始词,I-<实体类型>表示实体后续词,O表示非实体
不同句子之间用空行分隔。
转换代码示例
这里提供可直接运行的Python实现,依赖nltk做分词,提前运行pip install nltk安装依赖后即可使用:
import nltk nltk.download('punkt') from nltk.tokenize import word_tokenize # 你的原始标注数据 raw_data = [ ('The F15 aircraft uses a lot of fuel', {'entities': [(4, 7, 'aircraft')]}), ('did you see the F16 landing?', {'entities': [(16, 19, 'aircraft')]}), ('how many missiles can a F35 carry', {'entities': [(24, 27, 'aircraft')]}), ('is the F15 outdated', {'entities': [(7, 10, 'aircraft')]}), ('how long does it take to train a F16 pilot',{'entities': [(33, 36, 'aircraft')]}), ('how much does a F35 cost', {'entities': [(16, 19, 'aircraft')]}) ] conll_lines = [] for text, annot in raw_data: entities = annot['entities'] # 先获取每个token的偏移位置 tokens = word_tokenize(text) # 记录每个token的起始、结束下标 token_spans = [] current_idx = 0 for token in tokens: start = text.find(token, current_idx) end = start + len(token) token_spans.append((start, end, token)) current_idx = end # 给每个token打标签 for start, end, token in token_spans: label = 'O' for ent_start, ent_end, ent_type in entities: # 判断当前token是否在实体范围内 if start >= ent_start and end <= ent_end: if start == ent_start: label = f'B-{ent_type}' else: label = f'I-{ent_type}' break # 词性和块标签用-占位 conll_lines.append(f"{token} - - {label}") # 句子之间加空行 conll_lines.append("") # 输出结果或者写入文件 conll_result = "\n".join(conll_lines) print(conll_result) # 如果要保存成文件就运行下面这行 # with open("output.conll", "w", encoding="utf-8") as f: # f.write(conll_result)
转换后的效果示例
以上代码运行后输出的CONLL格式内容示例如下:
The - - O F15 - - B-aircraft aircraft - - O uses - - O a - - O lot - - O of - - O fuel - - O did - - O you - - O see - - O the - - O F16 - - B-aircraft landing - - O ? - - O
你可以根据自己的需求调整分词工具、是否保留词性/块标签列即可。
内容的提问来源于stack exchange,提问作者imhans33
相关产品推荐
相关产品推荐

