You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用CreateHyperParameterTuningJob遇ValidationException:mse不适用于指定XGBoost算法

SageMaker XGBoost超参数调优:目标指标[mse]无效问题解决

错误现象

调用CreateHyperParameterTuningJob时触发ValidationException:

ClientError: 调用CreateHyperParameterTuningJob操作时发生(ValidationException)错误:超参数调优作业的目标指标[mse]对算法[720646828776.dkr.ecr.ap-south-1.amazonaws.com/sagemaker-xgboost:0.90-2-cpu-py3]无效,请选择有效的目标指标。

问题根源

  1. 指标定义未传递:代码中定义了metric_definitions但初始化HyperparameterTuner时传入None,导致SageMaker无法识别自定义的mse指标
  2. 目标类型错误:MSE(均方误差)是越小模型性能越好,代码中却设置为Maximize,与指标优化方向矛盾
  3. 容器版本与指标不匹配:报错显示使用XGBoost 0.90-2版本,该版本下reg:squarederror目标函数的日志输出指标名称格式和代码中定义的不匹配
  4. 代码变量遗漏:存在sess、default_bucket未定义的语法问题,可能导致实际运行时自动使用旧版本容器

修复方案

  • 将自定义的metric_definitions正确传递给HyperparameterTuner
  • 将目标类型改为Minimize,匹配MSE的优化逻辑
  • 调整正则表达式,匹配对应XGBoost版本日志中的实际指标输出格式
  • 补全代码中未定义的变量,确保使用指定版本的容器

修正后的完整代码

import datetime
import time
import tarfile    
import boto3
import pandas as pd
import numpy as np
from sagemaker import get_execution_role
import sagemaker
from sklearn.model_selection import train_test_split
from sklearn.datasets import fetch_california_housing
from sagemaker.tuner import (
    IntegerParameter,
    CategoricalParameter,
    ContinuousParameter,
    HyperparameterTuner,
)

s3 = boto3.client("s3")
sm_boto3 = boto3.client("sagemaker")

sagemaker_session = sagemaker.Session()
# 修正未定义的sess变量
region = sagemaker_session.boto_session.region_name

role = get_execution_role()
#Set the required configurations
model_name = "abc_model"
env = "dev"
#S3 Bucket
bucket = "abcpoc"
print("Using bucket " + bucket)


from sagemaker.debugger import Rule, rule_configs
from sagemaker.session import TrainingInput

# 修正未定义的default_bucket变量
s3_input_train = TrainingInput(
    s3_data=f"s3://{bucket}/train/",content_type="csv")
s3_input_validation = TrainingInput(
    s3_data=f"s3://{bucket}/validation/",content_type="csv")
prefix = 'output'

container=sagemaker.image_uris.retrieve("xgboost", region, "1.2-1")
print(container)
xgb = sagemaker.estimator.Estimator(
    image_uri=container,
    role=role,
    base_job_name="xgboost-random-search",
    instance_count=1,
    instance_type="ml.m4.xlarge",
    output_path="s3://{}/{}/output".format(bucket, prefix),
    sagemaker_session= sagemaker_session,
    rules=[Rule.sagemaker(rule_configs.create_xgboost_report())]
)


xgb.set_hyperparameters(
    max_depth = 5,
    eta = 0.2,
    gamma = 4,
    min_child_weight = 6,
    subsample = 0.7,
    objective = "reg:squarederror",
    num_round = 1000
)
hyperparameter_ranges = {
    "eta": ContinuousParameter(0, 1),
    "min_child_weight": ContinuousParameter(1, 10),
    "alpha": ContinuousParameter(0, 2),
    "max_depth": IntegerParameter(1, 10),
}

objective_metric_name = "mse"
# 调整正则表达式匹配XGBoost 1.2版本的日志输出格式
metric_definitions = [{"Name": "mse", "Regex": "validation-mse: ([0-9\\.]+)"}]

# 修复:传递metric_definitions,修正objective_type为Minimize
tuner = HyperparameterTuner(xgb, 
                    objective_metric_name, 
                    hyperparameter_ranges, 
                    metric_definitions=metric_definitions, 
                    strategy='Bayesian', 
                    objective_type='Minimize', 
                    max_jobs=1, 
                    max_parallel_jobs=1, 
                    tags=None, 
                    base_tuning_job_name=None)

#Tune
tuner.fit({
    "train":s3_input_train,
    "validation":s3_input_validation
    },include_cls_metadata=False)

#Explore the best model generated
tuning_job_result = boto3.client("sagemaker").describe_hyper_parameter_tuning_job(
    HyperParameterTuningJobName=tuner.latest_tuning_job.job_name
)

job_count = tuning_job_result["TrainingJobStatusCounters"]["Completed"]
print("%d training jobs have completed" %job_count)
#10 training jobs have completed

#Get the best training job

from pprint import pprint
if tuning_job_result.get("BestTrainingJob",None):
    print("Best Model found so far:")
    pprint(tuning_job_result["BestTrainingJob"])
else:
    print("No training jobs have reported results yet.")

关键修改点说明

  1. 传递指标定义:把定义好的metric_definitions传入HyperparameterTuner,让SageMaker能从训练日志中正确提取MSE指标
  2. 修正目标类型:将objective_type改为Minimize,符合MSE指标越小越好的优化逻辑
  3. 补全变量:修复sess和default_bucket未定义的问题,确保使用指定的XGBoost 1.2-1版本容器
  4. 匹配日志格式:XGBoost 1.2版本输出的验证集MSE格式为validation-mse: x.xx,因此调整正则表达式匹配该格式

内容的提问来源于stack exchange,提问作者Ashish Jha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 11:15:29