You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求生成NER模型训练数据的BIO格式标注Python实现代码

生成NER数据集BIO标注的Python实现

核心思路

  1. 按空格分词并记录每个分词的起始/结束字符索引
  2. 初始化所有分词的标注为O
  3. 遍历实体,匹配对应分词:第一个匹配的分词标注B-类别,后续同实体分词标注I-类别
  4. 将标注列表拼接为空格分隔的字符串

完整代码

import json

def get_bio_labels(sentence, entities):
    # 按空格分词,同步记录每个词的字符位置
    words = sentence.split()
    word_positions = []
    current_idx = 0
    for word in words:
        start = current_idx
        end = current_idx + len(word)
        word_positions.append((start, end))
        current_idx = end + 1  # 跳过词间空格

    # 初始化所有标注为O
    labels = ['O'] * len(words)

    for entity in entities:
        ent_start = entity['start']
        ent_end = entity['end']
        ent_category = entity['category']
        is_first_word = True

        for idx, (word_start, word_end) in enumerate(word_positions):
            # 判断当前词是否落在实体的字符范围内
            if word_start >= ent_start and word_end <= ent_end:
                if is_first_word:
                    labels[idx] = f'B-{ent_category}'
                    is_first_word = False
                else:
                    labels[idx] = f'I-{ent_category}'

    return ' '.join(labels)

def process_ner_json(input_path, output_path=None):
    # 读取原始JSON数据集
    with open(input_path, 'r', encoding='utf-8') as f:
        dataset = json.load(f)
    
    # 为每条数据生成BIO标注
    for item in dataset:
        item['bio_labels'] = get_bio_labels(item['request'], item['entities'])
    
    # 保存带标注的数据集(可选)
    if output_path:
        with open(output_path, 'w', encoding='utf-8') as f:
            json.dump(dataset, f, indent=2, ensure_ascii=False)
    
    return dataset

# 示例测试
if __name__ == '__main__':
    test_item = {
        "request": "I want to fly to New York on the 13.3",
        "entities": [
            {"start": 16, "end": 23, "text": "New York", "category": "DESTINATION"},
            {"start": 32, "end": 35, "text": "13.3", "category": "DATE"}
        ]
    }
    print(get_bio_labels(test_item['request'], test_item['entities']))
    # 输出:O O O O O B-DESTINATION I-DESTINATION O O B-DATE

    # 批量处理JSON文件示例
    # process_ner_json('your_input.json', 'your_output_with_bio.json')

注意事项

  • 代码默认按空格分词,若需处理中文或复杂英文场景,可替换为nltk.word_tokenize等专业分词工具
  • 需确保实体的start/end字段严格对应原句子的字符索引(空格计入字符数)
  • 处理后的数据会新增bio_labels字段,存储最终的BIO标注字符串

内容的提问来源于stack exchange,提问作者Pizi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 17:28:18