Azure ML批量端点作业报TypeError:JSON对象类型不符求助
问题根源
错误TypeError: the JSON object must be str, bytes or bytearray, not MiniBatch的核心原因是:自动生成的评分脚本是为在线实时推理设计的,而批量端点调用时,run函数接收的参数是MiniBatch对象(包含批量处理的文件路径集合),不是JSON字符串,因此json.loads(data)无法解析该类型。
解决方案:修改评分脚本适配批量推理
将评分脚本的run函数逻辑调整为处理批量文件输入,以下是完整修改后的脚本:
import os import json from typing import List from azureml.studio.core.io.model_directory import ModelDirectory from pathlib import Path from azureml.studio.modules.ml.score.score_generic_module.score_generic_module import ScoreModelModule from azureml.designer.serving.dagengine.converter import create_dfd_from_dict from collections import defaultdict from azureml.designer.serving.dagengine.utils import decode_nan from azureml.studio.common.datatable.data_table import DataTable model_path = os.path.join(os.getenv('AZUREML_MODEL_DIR'), 'trained_model_outputs') schema_file_path = Path(model_path) / '_schema.json' with open(schema_file_path) as fp: schema_data = json.load(fp) def init(): global model model = ModelDirectory.load(model_path).model def run(mini_batch): # 初始化输入数据容器 input_entry = defaultdict(list) # 遍历MiniBatch中的每个文件路径,逐个读取解析 for file_path in mini_batch: with open(file_path, 'r') as f: data = json.load(f) # 合并所有文件中的数据行 for row in data: for key, val in row.items(): input_entry[key].append(decode_nan(val)) # 保留原有评分逻辑不变 data_frame_directory = create_dfd_from_dict(input_entry, schema_data) score_module = ScoreModelModule() result, = score_module.run( learner=model, test_data=DataTable.from_dfd(data_frame_directory), append_or_result_only=True) return json.dumps({"result": result.data_frame.values.tolist()})
关键改动说明
- 函数参数从
data改为mini_batch,明确接收批量文件路径列表 - 新增文件遍历逻辑:循环读取
mini_batch中的每个文件,解析JSON内容并合并数据 - 保留原有数据解析、模型评分和结果返回的核心逻辑,仅调整输入读取部分
额外注意事项
- 确保输入JSON文件格式为JSON数组(每个元素是一条待预测数据),示例格式:
[{"feature1": 0.5, "feature2": "value1"}, {"feature1": 0.8, "feature2": "value2"}] - 如果需要批量处理多个文件,建议将输入类型改为
AssetTypes.URI_FOLDER,指向包含多个JSON文件的文件夹,批量端点会自动遍历处理所有文件
内容的提问来源于stack exchange,提问作者Bartek Pobocha
相关产品推荐
相关产品推荐

