You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

构建NER训练数据集:基于分词与实体跨度生成对应标签

文本分词与标签构建需求

我有以下输入内容:

text = "Apple est une entreprise, James Alfred travaille ici"
spans = [
    {
        "start":0,
        "end":5,
        "label":"ORG"
    },
    {
        "start":26,
        "end":38,
        "label":"PER"
    }
]

correspondance_dict = {"PER":2, "ORG": 4 , "O" : 0}

需要对上述文本进行分词,并结合spans列表与correspondance_dict构建标签,最终得到如下输出:

tokenized_text = ["Apple", "est", "une", "entreprise", "," , "James","Alfred", "travaille", "ici"]
labels = [4,0,0,0,0,2,2,0,0]  # 此列表由correspondance_dict和spans构建(4对应Apple为ORG,"2,2"对应James、Alfred为PER)

内容的提问来源于stack exchange,提问作者michel gold

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 21:30:53