如何利用Azure ML标注数据训练Azure认知服务语言自定义模型?格式转换方法?
问题解答
能否直接导出符合要求的JSON?
目前Azure ML标注环境无法直接导出适配Azure Cognitive Service for Language训练要求的JSON格式数据,必须通过转换步骤实现格式兼容。
格式转换方案
1. 文本分类:CSV转目标JSON
Azure Cognitive Service for Language自定义文本分类的单标签训练JSON结构示例如下:
[ { "text": "待分类的文本内容", "category": "标签名称" }, { "text": "另一段文本", "category": "另一标签" } ]
假设你从Azure ML导出的CSV包含text(存文本内容)和label(存对应标签)两列,可使用以下Python脚本完成转换:
import pandas as pd import json # 读取导出的CSV文件 df = pd.read_csv("text_classification_export.csv") # 转换为目标JSON结构 json_dataset = [] for _, row in df.iterrows(): json_dataset.append({ "text": row["text"], "category": row["label"] }) # 保存为训练用JSON文件 with open("text_classification_train.json", "w", encoding="utf-8") as f: json.dump(json_dataset, f, ensure_ascii=False, indent=2)
如果是多标签分类场景,只需将category字段改为categories数组,脚本对应调整即可。
2. NER:CoNLL转目标JSON
Azure Cognitive Service for Language自定义NER的训练JSON结构示例如下:
[ { "text": "完整的待标注文本", "entities": [ { "offset": 0, "length": 6, "category": "实体类型" } ] } ]
CoNLL格式以「每行一个token+IOB标签」、空行分隔文档的方式存储,以下是转换的Python脚本示例:
import json def conll_to_azure_ner(conll_file_path, output_json_path): documents = [] current_tokens = [] current_entities = [] active_entity = None current_offset = 0 with open(conll_file_path, "r", encoding="utf-8") as f: for line in f: line = line.strip() # 空行代表当前文档结束,开始整理数据 if not line: if current_tokens: full_text = " ".join(current_tokens) # 收尾未闭合的实体 if active_entity: current_entities.append(active_entity) # 格式化实体列表 formatted_entities = [ {"offset": e["offset"], "length": e["length"], "category": e["category"]} for e in current_entities ] documents.append({"text": full_text, "entities": formatted_entities}) # 重置状态处理下一个文档 current_tokens = [] current_entities = [] active_entity = None current_offset = 0 continue # 解析CoNLL行的token和标签 token, label = line.split("\t") current_tokens.append(token) token_len = len(token) # 处理IOB标签逻辑 if label.startswith("B-"): # 开启新实体,先把之前的活动实体存入列表 if active_entity: current_entities.append(active_entity) entity_type = label[2:] active_entity = { "offset": current_offset, "length": token_len, "category": entity_type } elif label.startswith("I-") and active_entity: # 延续当前实体,长度加上token和空格的长度 entity_type = label[2:] if active_entity["category"] == entity_type: active_entity["length"] += token_len + 1 else: # 非实体标签,收尾活动实体 if active_entity: current_entities.append(active_entity) active_entity = None # 更新下一个token的起始偏移量(当前token长度+空格) current_offset += token_len + 1 # 处理最后一个未闭合的文档 if current_tokens: full_text = " ".join(current_tokens) if active_entity: current_entities.append(active_entity) formatted_entities = [ {"offset": e["offset"], "length": e["length"], "category": e["category"]} for e in current_entities ] documents.append({"text": full_text, "entities": formatted_entities}) # 保存转换后的JSON文件 with open(output_json_path, "w", encoding="utf-8") as f: json.dump(documents, f, ensure_ascii=False, indent=2) # 替换为你的文件路径 conll_to_azure_ner("ner_export.conll", "ner_train.json")
转换后验证建议
- 文本分类:抽查几个样本,确认
text和category(或categories)对应无误。 - NER:随机选取文档,核对实体的
offset和length是否能准确定位到文本中的对应内容,避免因特殊字符、连续空格导致的偏移错误。
内容的提问来源于stack exchange,提问作者GnicarAzrof
相关产品推荐
相关产品推荐

