Azure ML CLI运行R代码:输出文件未保存问题求助
Azure ML CLI运行R任务无法保存输出文件的解决方法
问题描述
尝试通过Azure ML CLI运行R代码任务,训练并保存模型。使用简单测试脚本验证时,任务显示运行完成,但输出文件iris_head.csv并未保存到指定路径。已成功运行输出文件的Hello World示例,但无法复用到R脚本场景。微软关于Azure ML中运行R代码的官方文档已移除,参考相关Github示例(含MLflow相关)也遇到同样问题。
项目目录结构
AZURE-ML-IRIS - docker-context ---- Dockerfile(来自微软Github的azureml-examples R示例) - src ---- train.R - job.yml
原代码与配置
train.R原代码
library(optparse) library(rpart) parser <- OptionParser() parser <- add_option( parser, "--data_folder", type="character", action="store", default = "./data", help="data folder") parser <- add_option( parser, "--data_output", type = "character", action = "store", default = "./data_output" ) args <- parse_args(parser) file_name = file.path(args$data_folder) iris <- read.csv(file_name) iris_head <- head(iris) write.csv(iris_head, file = paste0(args$data_output, "/iris_head.csv"))
job.yml原配置
$schema: https://azuremlschemas.azureedge.net/latest/commandJob.schema.json command: > Rscript train.R --data_folder ${{inputs.iris}} --data_output ${{outputs.data_output}} code: src inputs: iris: type: uri_file path: https://azuremlexamples.blob.core.windows.net/datasets/iris.csv outputs: data_output: environment: build: path: docker-context display_name: r-iris-example compute: azureml:noel001 experiment_name: r-iris-example description: Get a subset of Iris data.
可正常运行的Hello World示例
$schema: https://azuremlschemas.azureedge.net/latest/commandJob.schema.json command: echo "hello world" > ${{outputs.hello_output}}/helloworld.txt outputs: hello_output: environment: image: python
关键问题分析
- 输入类型不匹配:
inputs.iris为uri_file类型(单个文件),但R脚本中错误地将其当作文件夹路径处理,可能导致读取逻辑异常,间接影响后续输出写入。 - 输出目录未提前创建:R脚本直接写入输出路径,但该目录在容器中可能未自动创建,导致写入失败。
- 路径拼接不规范:使用
paste0拼接路径可能因系统斜杠差异引发问题,且未显式处理目录权限。
修复方案
1. 修正R脚本的输入读取逻辑
直接使用传入的文件URI读取数据,无需当作文件夹处理:
# 替换原读取代码 iris <- read.csv(args$data_folder)
2. 确保输出目录存在并规范路径拼接
在写入文件前创建输出目录,使用file.path处理路径拼接:
# 在write.csv前添加目录检查与创建 if (!dir.exists(args$data_output)) { dir.create(args$data_output, recursive = TRUE) } # 替换原写入代码 write.csv(iris_head, file = file.path(args$data_output, "iris_head.csv"), row.names = FALSE)
3. 验证Docker环境依赖
确保Dockerfile中安装了所需R包,添加以下内容:
RUN install2.r optparse rpart
4. 优化job.yml配置
显式声明输出为文件夹类型,避免歧义:
outputs: data_output: type: uri_folder # 显式指定为文件夹输出
验证步骤
- 若修改了Dockerfile,提交任务时会自动重新构建环境
- 提交任务:
az ml job create --file job.yml - 查看任务日志排查问题:
az ml job logs --name <任务名称> - 下载输出文件验证:
az ml job download --name <任务名称> --output-name data_output
内容的提问来源于stack exchange,提问作者dempseynoel
相关产品推荐
相关产品推荐

