使用Python创建Dataproc工作流模板时无法参数化placement.managedCluster.config下字段的问题
解决Dataproc工作流模板参数化subnetworkUri等字段的错误问题
看起来你遇到的问题是因为参数化字段路径的格式不正确,Dataproc工作流模板的参数字段路径要求严格匹配JSON结构中下划线命名的字段,而不是使用API proto定义里的驼峰式名称。
问题根源
你的parameters数组里的字段路径用了驼峰式命名(比如placement.managedCluster.config.gceClusterConfig.subnetworkUri),但Dataproc的参数化路径需要使用JSON模板里的下划线格式字段名。错误信息里的placement.managed_cluster.configuration.gce_cluster_config.subnetwork_uri其实已经提示了正确的路径结构——只是你的路径里用了managedCluster(驼峰)而非managed_cluster(下划线),gceClusterConfig应该是gce_cluster_config,subnetworkUri应该是subnetwork_uri。
修正步骤
1. 修正JSON模板中的parameters字段路径
把所有参数的fields值改成匹配JSON结构的下划线命名路径:
{ "id": "bigquery-extractor", "placement": { "managed_cluster": { "config": { "gce_cluster_config": { "subnetwork_uri": "some-subnet-name" }, "software_config" : { "image_version": "1.5" } }, "cluster_name": "some-name" } }, "jobs": [ { "pyspark_job": { "args": [ "job_argument" ], "main_python_file_uri": "gs:///path-to-file" }, "step_id": "extract" } ], "parameters": [ { "name": "CLUSTER_NAME", "fields": [ "placement.managed_cluster.cluster_name" ] }, { "name": "SUBNETWORK_URI", "fields": [ "placement.managed_cluster.config.gce_cluster_config.subnetwork_uri" ] }, { "name": "MAIN_PY_FILE", "fields": [ "jobs['extract'].pyspark_job.main_python_file_uri" ] }, { "name": "JOB_ARGUMENT", "fields": [ "jobs['extract'].pyspark_job.args[0]" ] } ] }
2. 替换不安全的eval为json.load()
你的代码里用eval(template_file.read())来解析JSON是非常不安全的,而且容易出现格式问题,建议改用标准的json模块:
import json from google.cloud import dataproc from google.api_core.exceptions import AlreadyExists options = dataproc.ClientOptions(api_endpoint="{}-dataproc.googleapis.com:443".format(region)) client = dataproc.WorkflowTemplateServiceClient(client_options=options) with open(path_to_file, "r") as template_file: template_dict = json.load(template_file) template = dataproc.WorkflowTemplate(template_dict) full_region_id = "projects/{project_id}/regions/{region}".format(project_id=project_id, region=region) try: client.create_workflow_template( parent=full_region_id, template=template ) print("Template created successfully!") except AlreadyExists as err: print("Template already exists:", err)
额外说明
Dataproc工作流模板的参数化路径规则:
- 必须使用JSON结构中的下划线字段名,而非API的驼峰式字段名
- 对于数组或对象的嵌套访问,使用
['step_id']或[index]的格式(比如你代码里的jobs['extract']是正确的) - 没有限制参数化
managed_cluster.config下的字段,只要路径正确就可以正常工作
内容的提问来源于stack exchange,提问作者Dmitriy Lamzin
相关产品推荐
相关产品推荐

