如何在Bicep中引用Synapse工作区的现有Spark配置?
问题:Synapse Spark池关联现有Spark配置失败
我有可用的Spark池Bicep模板,想关联指定Synapse工作区中已创建的现有Spark配置(对应界面选择现有配置的操作),但运行模板时系统会创建新配置而非关联现有配置:
@description('Creates Spark Pool resources within the specified Synapse Workspace based on the provided configurations.') resource synapseWorkspaceDSSparkPools 'Microsoft.Synapse/workspaces/bigDataPools@2021-06-01' = [ for sparkPool in sparkPools: { parent: synapseWorkspaceDS name: sparkPool.name location: location properties: { nodeCount: sparkPool.nodeCount nodeSizeFamily: sparkPool.nodeSizeFamily nodeSize: sparkPool.nodeSize autoScale: sparkPool.autoScaleEnabled ? { enabled: true minNodeCount: sparkPool.minNodeCount maxNodeCount: sparkPool.maxNodeCount } : null autoPause: sparkPool.autoPauseEnabled ? { enabled: true delayInMinutes: sparkPool.autoPauseDelayInMinutes } : null // 创建新Spark池时该属性需为空,因为必须先创建池才能安装库 customLibraries: [] // 创建新Spark池时该属性需为空,因为必须先创建池才能安装库 libraryRequirements: {} sparkVersion: sparkPool.sparkVersion sparkConfigProperties: { configurationType: 'Artifact' filename: sparkConfigurationDSName content: '{"id":"${synapseWorkspace.id}/sparkconfigurations/${sparkConfigurationName}","name":"${sparkConfigurationName}","type":"Microsoft.Synapse/workspaces/sparkconfigurations","properties":{"description":"Configuration enabled to track logs from Spark Pools to Log Analytics. The documentation can be found here: https://learn.microsoft.com/en-us/azure/synapse-analytics/spark/apache-spark-azure-log-analytics","configs":{"spark.synapse.logAnalytics.enabled":"true","spark.synapse.logAnalytics.workspaceId":"${operationalInsightsWorkspace.properties.customerId}","spark.synapse.logAnalytics.secret":"${logAnalyticsWorkspaceSharedKey}"},"annotations":[],"configMergeRule":{"artifact.currentOperation.spark.synapse.logAnalytics.enabled":"replace","artifact.currentOperation.spark.synapse.logAnalytics.workspaceId":"replace"}}}'' } isComputeIsolationEnabled: sparkPool.isComputeIsolationEnabled sessionLevelPackagesEnabled: sparkPool.sessionLevelPackagesEnabled dynamicExecutorAllocation: sparkPool.dynamicExecutorAllocationEnabled ? { enabled: true minExecutors: sparkPool.minExecutorCount maxExecutors: sparkPool.maxExecutorCount } : null cacheSize: sparkPool.cacheSize } } ]
若移除content字段,仅保留filename并将configurationType设为Artifact或File,会触发如下验证错误:
"message":"至少有一个资源部署操作失败。请列出部署操作以获取详细信息。","details":[{"code":"ValidationFailed","message":"Spark池请求验证失败。","details":[{"code":"SparkComputePropertiesCorrupted","message":"SparkConfigProperties字段已损坏,Content: , Filename:gmatheus01rsynwsparkConfiguration"}]}
解决方法
要关联现有Spark配置,不能通过sparkConfigProperties传递配置内容,需先引用已存在的Spark配置资源,再通过sparkConfigResourceId属性关联:
- 先引用已存在的Spark配置资源:
// 引用目标工作区中已存在的Spark配置 resource existingSparkConfig 'Microsoft.Synapse/workspaces/sparkconfigurations@2021-06-01' existing = { parent: synapseWorkspaceDS name: sparkConfigurationName }
- 修改Spark池的
sparkConfigProperties部分,替换为关联逻辑:
sparkConfigProperties: { configurationType: 'Artifact' sparkConfigResourceId: existingSparkConfig.id }
完整修改后的Spark池资源片段:
@description('Creates Spark Pool resources within the specified Synapse Workspace based on the provided configurations.') resource synapseWorkspaceDSSparkPools 'Microsoft.Synapse/workspaces/bigDataPools@2021-06-01' = [ for sparkPool in sparkPools: { parent: synapseWorkspaceDS name: sparkPool.name location: location properties: { nodeCount: sparkPool.nodeCount nodeSizeFamily: sparkPool.nodeSizeFamily nodeSize: sparkPool.nodeSize autoScale: sparkPool.autoScaleEnabled ? { enabled: true minNodeCount: sparkPool.minNodeCount maxNodeCount: sparkPool.maxNodeCount } : null autoPause: sparkPool.autoPauseEnabled ? { enabled: true delayInMinutes: sparkPool.autoPauseDelayInMinutes } : null customLibraries: [] libraryRequirements: {} sparkVersion: sparkPool.sparkVersion // 关联现有Spark配置 sparkConfigProperties: { configurationType: 'Artifact' sparkConfigResourceId: existingSparkConfig.id } isComputeIsolationEnabled: sparkPool.isComputeIsolationEnabled sessionLevelPackagesEnabled: sparkPool.sessionLevelPackagesEnabled dynamicExecutorAllocation: sparkPool.dynamicExecutorAllocationEnabled ? { enabled: true minExecutors: sparkPool.minExecutorCount maxExecutors: sparkPool.maxExecutorCount } : null cacheSize: sparkPool.cacheSize } } ]
说明:
- 使用
existing关键字引用已存在的Spark配置,确保Bicep不会重新创建该配置 - 通过
sparkConfigResourceId传递现有配置的资源ID,实现和界面选择现有配置一致的关联逻辑
内容的提问来源于stack exchange,提问作者Guilherme Matheus
相关产品推荐
相关产品推荐

