Azure Synapse Dataflow无法直接克隆:调用栈溢出问题排查
Synapse Dataflow UI相关问题及临时解决办法
Dataflow组件构成
- 2个Source(数据源)
- 1个Sink(输出端)
- 4个Derived Column(派生列)
- 2个Union(合并节点)
- 1个Select(选择节点)
- 2个Filter(筛选节点)
- 2个Aggregate(聚合节点)
- 4个Flowlet:均引用同一个Flowlet,该Flowlet内部包含1个Aggregate和1个Select
问题现象
调试模式报错
该Dataflow在管道中运行完全正常,但进入Dataflow调试模式时会触发以下错误:
RangeError: Maximum call stack size exceeded” & “executeJobPreviewDataQuery no active stream for job=<job_id>, jobState=Completed
Synapse UI克隆操作报错
在Synapse UI中尝试克隆该Dataflow时,会生成空的Dataflow文件,发布操作触发错误:
Error while saving entities. Details: RangeError: Maximum call stack size exceeded
临时可行方案
直接将原Dataflow的JSON内容复制到新建Dataflow中,可正常使用。目前暂无法在ADF环境测试,不确定是否为Synapse专属问题。
更新说明
经排查疑似为Synapse/ADF UI层面的问题:若Dataflow未被打开过(比如UI刷新后),则可以正常完成克隆操作,且克隆结果与Dataflow名称无关。
原始Dataflow脚本
parameters{ hours_history as integer (1), column_names as string[] (["a","b","c","d"]) } source(allowSchemaDrift: true, validateSchema: false, ignoreNoFilesFound: true, modifiedAfter: (currentUTC()-hours($hours_history)), format: 'parquet', fileSystem: 'fs1', folderPath: 'fp1', compressionCodec: 'none', mode: 'read') ~> BNBack source(allowSchemaDrift: true, validateSchema: false, ignoreNoFilesFound: true, enableCdc: true, mode: 'read', skipInitialLoad: false, format: 'parquet', fileSystem: 'fs1', folderPath: 'fp1') ~> BLCDC deduplicatePreCDC@deduplicatedOutput compose(composition: 'Flowlet_DropEventHubMetadata') ~> dropEventHubMetadata1@(output1) dropTempGroupCol compose(composition: 'Flowlet_DropEventHubMetadata') ~> dropEventHubMetadata2@(output1) dropEventHubMetadata2@output1 compose(composition: 'Flowlet_FullDeduplication') ~> deduplicateRight@(deduplicatedOutput) dropEventHubMetadata1@output1 compose(composition: 'Flowlet_FullDeduplication') ~> deduplicateLeft@(deduplicatedOutput) BNBack compose(composition: 'Flowlet_FullDeduplication') ~> deduplicatePreNBack@(deduplicatedOutput) BLCDC compose(composition: 'Flowlet_FullDeduplication') ~> deduplicatePreCDC@(deduplicatedOutput) deduplicateRight@deduplicatedOutput derive(custom_count = 2) ~> addCountRight addCountRight, addCountLeft union(byName: true)~> union1 deduplicateLeft@deduplicatedOutput derive(custom_count = 1) ~> addCountLeft union1 aggregate(groupBy(temp_groupCol = sha2(256,byNames($column_names))), custom_count = sum(custom_count), each(match(not(name=="custom_count")), $$ = first($$))) ~> aggregate1 aggregate1 filter(toInteger(byName('custom_count'))==1) ~> filterLeftNew deduplicatePreCDC@deduplicatedOutput derive(custom_count_pre = 2) ~> addCountPreCDC addCountPreCDC, addCountPreNBack union(byName: true)~> unionPre deduplicatePreNBack@deduplicatedOutput derive(custom_count_pre = 1) ~> addCountPreNBack unionPre aggregate(groupBy(temp_groupCol_pre = sha2(256,byNames($column_names))), custom_count_pre = sum(custom_count_pre), each(match(not(name=="custom_count_pre")), $$ = first($$))) ~> aggregatePre aggregatePre filter(toInteger(byName('custom_count_pre'))==1) ~> filterWithoutNewCDC filterWithoutNewCDC select(mapColumn( each(match(not(in(['custom_count_pre','temp_groupCol_pre'],name)))) ), skipDuplicateMapInputs: true, skipDuplicateMapOutputs: true) ~> dropTempGroupCol filterLeftNew sink(allowSchemaDrift: true, validateSchema: false, format: 'parquet', fileSystem: 'fs1', folderPath: 'fp2', compressionCodec: 'none', umask: 0022, preCommands: [], postCommands: [], saveOrder: 1, mapColumn( each(match(not(in(['custom_count','temp_groupCol'],name)))) )) ~> L2
内容的提问来源于stack exchange,提问作者gCW
相关产品推荐
相关产品推荐

