You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Glue 4.0 Spark作业中XGBoost无法识别training.py文件问题

AWS Glue 4.0 中 XGBoost Estimator 找不到 training.py 的解决方法

在AWS Glue 4.0 Spark Shell环境下使用XGBoost 1.3-1版本创建估算器时,出现如下报错:

FileNotFoundError: [Errno 2] No such file or directory: 'training.py'

已将Glue主脚本、training.py和__init__.py放在同一S3文件夹下,文件名大小写无误,但XGBoost函数仍无法识别该文件。代码片段如下:

xgb_script_mode_estimator = XGBoost(
    entry_point="training.py",
    hyperparameters=hyperparameters,
    role=role,
    instance_count=1,
    instance_type=instance_type,
    framework_version="1.3-1",
    output_path="s3://{}/{}/{}/output".format(hyperparameters['bucket_nm'], '/output/', job_name),
)

解决方案

  • 指定source_dir参数:XGBoost Estimator默认会在本地文件系统查找entry_point指定的文件,但Glue作业的脚本存储在S3,需要通过source_dir参数指定包含training.py的S3文件夹路径,Estimator会自动将该目录下的文件同步到计算节点本地:
    xgb_script_mode_estimator = XGBoost(
        entry_point="training.py",
        source_dir="s3://{}/path/to/your/scripts/".format(hyperparameters['bucket_nm']),  # 替换为实际S3文件夹路径
        hyperparameters=hyperparameters,
        role=role,
        instance_count=1,
        instance_type=instance_type,
        framework_version="1.3-1",
        output_path="s3://{}/{}/output".format(hyperparameters['bucket_nm'], job_name),  # 修正output_path的多余斜杠问题
    )
    
  • 检查Glue角色权限:确保Glue作业使用的IAM角色拥有目标S3文件夹的s3:GetObject权限,避免因权限不足无法读取training.py文件。
  • 修正output_path格式:原代码中format参数的'/output/'会导致生成的路径出现多余的连续斜杠(如s3://bucket//output//job_name/output),建议调整为'output'或直接拼接路径,避免路径格式错误。

内容的提问来源于stack exchange,提问作者dewdrops

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 19:38:39