You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用Azure ML标注数据训练Azure认知服务语言自定义模型?格式转换方法?

问题解答

能否直接导出符合要求的JSON?

目前Azure ML标注环境无法直接导出适配Azure Cognitive Service for Language训练要求的JSON格式数据,必须通过转换步骤实现格式兼容。

格式转换方案

1. 文本分类:CSV转目标JSON

Azure Cognitive Service for Language自定义文本分类的单标签训练JSON结构示例如下:

[
  {
    "text": "待分类的文本内容",
    "category": "标签名称"
  },
  {
    "text": "另一段文本",
    "category": "另一标签"
  }
]

假设你从Azure ML导出的CSV包含text(存文本内容)和label(存对应标签)两列,可使用以下Python脚本完成转换:

import pandas as pd
import json

# 读取导出的CSV文件
df = pd.read_csv("text_classification_export.csv")

# 转换为目标JSON结构
json_dataset = []
for _, row in df.iterrows():
    json_dataset.append({
        "text": row["text"],
        "category": row["label"]
    })

# 保存为训练用JSON文件
with open("text_classification_train.json", "w", encoding="utf-8") as f:
    json.dump(json_dataset, f, ensure_ascii=False, indent=2)

如果是多标签分类场景,只需将category字段改为categories数组,脚本对应调整即可。

2. NER:CoNLL转目标JSON

Azure Cognitive Service for Language自定义NER的训练JSON结构示例如下:

[
  {
    "text": "完整的待标注文本",
    "entities": [
      {
        "offset": 0,
        "length": 6,
        "category": "实体类型"
      }
    ]
  }
]

CoNLL格式以「每行一个token+IOB标签」、空行分隔文档的方式存储,以下是转换的Python脚本示例:

import json

def conll_to_azure_ner(conll_file_path, output_json_path):
    documents = []
    current_tokens = []
    current_entities = []
    active_entity = None
    current_offset = 0

    with open(conll_file_path, "r", encoding="utf-8") as f:
        for line in f:
            line = line.strip()
            # 空行代表当前文档结束,开始整理数据
            if not line:
                if current_tokens:
                    full_text = " ".join(current_tokens)
                    # 收尾未闭合的实体
                    if active_entity:
                        current_entities.append(active_entity)
                    # 格式化实体列表
                    formatted_entities = [
                        {"offset": e["offset"], "length": e["length"], "category": e["category"]}
                        for e in current_entities
                    ]
                    documents.append({"text": full_text, "entities": formatted_entities})
                # 重置状态处理下一个文档
                current_tokens = []
                current_entities = []
                active_entity = None
                current_offset = 0
                continue
            
            # 解析CoNLL行的token和标签
            token, label = line.split("\t")
            current_tokens.append(token)
            token_len = len(token)

            # 处理IOB标签逻辑
            if label.startswith("B-"):
                # 开启新实体,先把之前的活动实体存入列表
                if active_entity:
                    current_entities.append(active_entity)
                entity_type = label[2:]
                active_entity = {
                    "offset": current_offset,
                    "length": token_len,
                    "category": entity_type
                }
            elif label.startswith("I-") and active_entity:
                # 延续当前实体,长度加上token和空格的长度
                entity_type = label[2:]
                if active_entity["category"] == entity_type:
                    active_entity["length"] += token_len + 1
            else:
                # 非实体标签,收尾活动实体
                if active_entity:
                    current_entities.append(active_entity)
                    active_entity = None
            
            # 更新下一个token的起始偏移量(当前token长度+空格)
            current_offset += token_len + 1
        
        # 处理最后一个未闭合的文档
        if current_tokens:
            full_text = " ".join(current_tokens)
            if active_entity:
                current_entities.append(active_entity)
            formatted_entities = [
                {"offset": e["offset"], "length": e["length"], "category": e["category"]}
                for e in current_entities
            ]
            documents.append({"text": full_text, "entities": formatted_entities})
    
    # 保存转换后的JSON文件
    with open(output_json_path, "w", encoding="utf-8") as f:
        json.dump(documents, f, ensure_ascii=False, indent=2)

# 替换为你的文件路径
conll_to_azure_ner("ner_export.conll", "ner_train.json")

转换后验证建议

  • 文本分类:抽查几个样本,确认text和category(或categories)对应无误。
  • NER:随机选取文档,核对实体的offset和length是否能准确定位到文本中的对应内容,避免因特殊字符、连续空格导致的偏移错误。

内容的提问来源于stack exchange,提问作者GnicarAzrof

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 08:25:22