You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将自定义标注数据集转为CoNLL格式并给未标注Token标注O用于BERT微调?

将自定义标注JSON数据集转换为CoNLL格式(含O标注)

问题背景

我有一个手动标注的JSON格式数据集,单条记录结构如下:

{
    "id": 1,
    "text": "At the end of each fiscal quarter, for the four consecutive fiscal quarters ending as of such fiscal quarter end, from the date of the Third Amendment and until December 30, 1996, the Company shall maintain a fixed charge coverage ratio of not less than 1.25 to 1.0.",
    "label": [
        [309, 336, "COV_3"],
        [379, 390, "VAL_3"]
    ]
}

注:label字段的标注对应文本实体:fixed charge coverage(字符位置[309, 336])标注为COV_3;1.25 to 1.0(字符位置[379, 390])标注为VAL_3。

现在需要将这类数据转换为适合微调BERT等Transformer模型的格式——要么给所有未标注token添加"O"标签,要么直接转成CoNLL格式。

解决方案

核心思路

  1. 用与目标微调模型一致的分词器拆分文本,保留每个token的字符起止位置,避免标签错位
  2. 初始化所有token标签为"O"
  3. 匹配原标注的字符区间与token位置,为对应token添加实体标签(遵循CoNLL常用的BIO规则)
  4. 输出为每行一个token+标签的CoNLL格式

具体实现代码

使用HuggingFace的tokenizers库对齐BERT类模型的分词逻辑:

from tokenizers import BertTokenizer
import json

# 初始化对应模型的分词器(示例用bert-base-uncased,可替换为你使用的模型)
tokenizer = BertTokenizer.from_pretrained("bert-base-uncased")

def json_to_conll(json_record):
    text = json_record["text"]
    entity_labels = json_record["label"]
    
    # 分词并保留每个token的字符偏移量
    encoding = tokenizer(
        text,
        return_offsets_mapping=True,
        return_attention_mask=False,
        return_token_type_ids=False
    )
    tokens = encoding.tokens()
    token_offsets = encoding.offset_mapping
    
    # 初始化所有标签为"O"
    token_tags = ["O"] * len(tokens)
    
    # 遍历每个实体标注,匹配对应token并设置标签
    for start_char, end_char, tag in entity_labels:
        for idx, (token_start, token_end) in enumerate(token_offsets):
            # 跳过[CLS]、[SEP]这类无实际文本的特殊token
            if token_start == 0 and token_end == 0:
                continue
            # 判断token与实体区间是否重叠(可根据标注规则调整匹配逻辑)
            if not (token_end <= start_char or token_start >= end_char):
                # 用BIO规则标注:实体首token加B-前缀,后续加I-前缀
                if token_tags[idx] == "O":
                    token_tags[idx] = f"B-{tag}"
                else:
                    token_tags[idx] = f"I-{tag}"
    
    # 转换为CoNLL格式,跳过特殊token
    conll_lines = []
    for token, tag in zip(tokens, token_tags):
        if token not in ["[CLS]", "[SEP]"]:
            conll_lines.append(f"{token} {tag}")
    return "\n".join(conll_lines)

# 测试示例转换
sample_record = {
    "id": 1,
    "text": "At the end of each fiscal quarter, for the four consecutive fiscal quarters ending as of such fiscal quarter end, from the date of the Third Amendment and until December 30, 1996, the Company shall maintain a fixed charge coverage ratio of not less than 1.25 to 1.0.",
    "label": [
        [309, 336, "COV_3"],
        [379, 390, "VAL_3"]
    ]
}

print(json_to_conll(sample_record))

批量处理数据集

如果数据集是每行一条JSON记录的文件,可批量转换为CoNLL格式:

# 读取原始JSON数据集,输出为CoNLL文件
with open("raw_dataset.json", "r") as in_file, open("converted_dataset.conll", "w") as out_file:
    for line in in_file:
        record = json.loads(line.strip())
        conll_content = json_to_conll(record)
        out_file.write(conll_content + "\n\n")  # 空行分隔不同样本

关键注意事项

  • 分词对齐:必须使用与后续微调模型完全一致的分词器,否则会出现标签与token不匹配的问题
  • BIO标签:该规则能帮助模型更好识别实体边界,若不需要也可直接给实体token使用原标签(如COV_3)
  • 匹配逻辑:代码中用的是重叠判断,若你的标注要求实体完全包含token,可调整判断条件

内容的提问来源于stack exchange,提问作者Ruchit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 16:30:50