求生成NER模型训练数据的BIO格式标注Python实现代码
生成NER数据集BIO标注的Python实现
核心思路
- 按空格分词并记录每个分词的起始/结束字符索引
- 初始化所有分词的标注为
O - 遍历实体,匹配对应分词:第一个匹配的分词标注
B-类别,后续同实体分词标注I-类别 - 将标注列表拼接为空格分隔的字符串
完整代码
import json def get_bio_labels(sentence, entities): # 按空格分词,同步记录每个词的字符位置 words = sentence.split() word_positions = [] current_idx = 0 for word in words: start = current_idx end = current_idx + len(word) word_positions.append((start, end)) current_idx = end + 1 # 跳过词间空格 # 初始化所有标注为O labels = ['O'] * len(words) for entity in entities: ent_start = entity['start'] ent_end = entity['end'] ent_category = entity['category'] is_first_word = True for idx, (word_start, word_end) in enumerate(word_positions): # 判断当前词是否落在实体的字符范围内 if word_start >= ent_start and word_end <= ent_end: if is_first_word: labels[idx] = f'B-{ent_category}' is_first_word = False else: labels[idx] = f'I-{ent_category}' return ' '.join(labels) def process_ner_json(input_path, output_path=None): # 读取原始JSON数据集 with open(input_path, 'r', encoding='utf-8') as f: dataset = json.load(f) # 为每条数据生成BIO标注 for item in dataset: item['bio_labels'] = get_bio_labels(item['request'], item['entities']) # 保存带标注的数据集(可选) if output_path: with open(output_path, 'w', encoding='utf-8') as f: json.dump(dataset, f, indent=2, ensure_ascii=False) return dataset # 示例测试 if __name__ == '__main__': test_item = { "request": "I want to fly to New York on the 13.3", "entities": [ {"start": 16, "end": 23, "text": "New York", "category": "DESTINATION"}, {"start": 32, "end": 35, "text": "13.3", "category": "DATE"} ] } print(get_bio_labels(test_item['request'], test_item['entities'])) # 输出:O O O O O B-DESTINATION I-DESTINATION O O B-DATE # 批量处理JSON文件示例 # process_ner_json('your_input.json', 'your_output_with_bio.json')
注意事项
- 代码默认按空格分词,若需处理中文或复杂英文场景,可替换为
nltk.word_tokenize等专业分词工具 - 需确保实体的
start/end字段严格对应原句子的字符索引(空格计入字符数) - 处理后的数据会新增
bio_labels字段,存储最终的BIO标注字符串
内容的提问来源于stack exchange,提问作者Pizi
相关产品推荐
相关产品推荐

