将TSV格式标注数据转换为spaCy NER所需训练格式的问题咨询
解决方案
原代码存在的核心问题
- 未将非
O标签的实体起止偏移和标签存入entities列表,导致实体数据完全缺失 - 未处理文件末尾无空行时最后一个句子的入库逻辑,会丢失最后一条标注数据
- 声明了
unique_labels变量但没有追加标签数据的逻辑,返回值为空
修正后可直接运行的代码
def load_data_spacy(file_path): ''' 转换TSV标注数据为spaCy NER要求的输入格式 TSV格式:word \t label \n word \t label \n \n word \t label 输出格式:[["sentence", {'entities': [(start, end, 'tags')]}]] ''' training_data, entities, sentence, unique_labels = [], [], [], [] start = 0 end = 0 with open(file_path, 'r', encoding='utf-8') as file: for line in file: line = line.strip("\n").split("\t") # 处理单词+标签行 if len(line) > 1: word = line[0] label = line[1] sentence.append(word) start = end end += len(word) + 1 # 单词长度+末尾空格 # 非O标签加入实体列表 if label != 'O': # 实体结束位置要减去后面多算的1个空格,匹配spaCy左闭右开的偏移规则 entities.append((start, end - 1, label)) if label not in unique_labels: unique_labels.append(label) # 处理句子分隔空行 if len(line) == 1: if sentence: full_sentence = " ".join(sentence) training_data.append([full_sentence, {'entities': entities}]) # 重置临时变量 end = 0 start = 0 entities, sentence = [], [] # 处理文件末尾没有空行的最后一个句子 if sentence: full_sentence = " ".join(sentence) training_data.append([full_sentence, {'entities': entities}]) return training_data, unique_labels
注意事项
- 代码默认使用utf-8编码读取文件,如果你标注文件是其他编码可自行修改
open方法的encoding参数 - 实体偏移计算已经和spaCy的字符偏移规则完全匹配,无需额外调整
- 如果你的标注是BIO格式,转换后spaCy会自动识别B/I前缀,无需额外处理
内容的提问来源于stack exchange,提问作者Aamer Ashfaque
相关产品推荐
相关产品推荐

