使用Map状态运行Step Function报错:Glue爬虫并行变量传递问题
Glue爬虫并行执行状态机变量传递修复方案
核心问题梳理
当前状态机存在3个关键变量传递错误:
Get Glue Crawler List步骤直接输出爬虫名称字符串列表,导致Map迭代时无法引用crawler_name变量Prepare Output步骤引用的$.crawler_name字段不存在StartCrawler和GetCrawler的ResultPath配置混乱,导致状态数据结构不一致
具体修改点及输入输出处理
1. 调整「Get Glue Crawler List」的输出格式
- 修改内容:移除
OutputPath,用ResultSelector将爬虫名称列表转换为带crawler_name字段的对象数组,再通过ResultPath保留原始输入结构 - 输入:状态机初始输入
{"tags": {"type": "parallel"}} - 输出:生成
{"output": {"crawlers": [{"crawler_name": "爬虫1"}, {"crawler_name": "爬虫2"}]}}格式的数据,为Map状态提供可迭代的结构化数据 - 代码修改片段:
"Get Glue Crawler List": { "Next": "Run Glue Crawlers", "Parameters": { "Tags.$": "$.tags" }, "Resource": "arn:aws:states:::aws-sdk:glue:listCrawlers", "Retry": [ { "BackoffRate": 5, "ErrorEquals": ["States.ALL"], "IntervalSeconds": 2, "MaxAttempts": 3 } ], "ResultSelector": { "crawlers.$": "$.CrawlerNames" }, "ResultPath": "$.output", "Type": "Task" }
2. 修复Map状态的迭代配置
- 修改内容:添加
ItemsPath指定迭代的数组来源,移除ResultPath: null以保留所有迭代结果 - 输入:上一步输出的
{"output": {"crawlers": [...]}} - 输出:所有迭代分支的结果会被收集为数组返回
- 代码修改片段:
"Run Glue Crawlers": { "End": true, "ItemsPath": "$.output.crawlers", "Iterator": { // 原有迭代器逻辑不变,仅调整内部变量引用 }, "Type": "Map" }
3. 修正迭代器内的变量引用
- 修改内容:
- 将
StartCrawler和GetCrawler的Parameters.Name改为引用$.crawler_name - 统一
ResultPath到$.response下,避免数据结构混乱
- 将
- 输入(迭代器内部):单个
{"crawler_name": "目标爬虫名称"}对象 - 输出(迭代器内部):包含爬虫运行状态和结果的结构化对象
- 关键代码调整片段:
"StartCrawler": { "Catch": [ { "Comment": "Crawler Already Running, just continue to monitor", "ErrorEquals": ["Glue.CrawlerRunningException"], "Next": "GetCrawler", "ResultPath": "$.response.start_crawler" } ], "Next": "GetCrawler", "Parameters": { "Name.$": "$.crawler_name" }, "Resource": "arn:aws:states:::aws-sdk:glue:startCrawler", "ResultPath": "$.response.start_crawler", // 原有Retry配置不变 "Type": "Task" }, "GetCrawler": { "Next": "Is Running?", "Parameters": { "Name.$": "$.crawler_name" }, "Resource": "arn:aws:states:::aws-sdk:glue:getCrawler", "ResultPath": "$.response.get_crawler", // 原有Retry配置不变 "Type": "Task" }
4. 修复「Prepare Output」的变量引用
- 修改内容:无需调整
crawler_name.$,现在$.crawler_name已存在于迭代输入中 - 输入:包含
response.get_crawler和crawler_name的对象 - 输出:格式化后的结果对象:
{"LastCrawl": {...}, "crawler_name": "目标爬虫名称"}
修改后的完整状态机代码
{ "Comment": "A utility state machine to run all Glue Crawlers that match tags", "StartAt": "Get Glue Crawler List", "States": { "Get Glue Crawler List": { "Next": "Run Glue Crawlers", "Parameters": { "Tags.$": "$.tags" }, "Resource": "arn:aws:states:::aws-sdk:glue:listCrawlers", "Retry": [ { "BackoffRate": 5, "ErrorEquals": [ "States.ALL" ], "IntervalSeconds": 2, "MaxAttempts": 3 } ], "ResultSelector": { "crawlers.$": "$.CrawlerNames" }, "ResultPath": "$.output", "Type": "Task" }, "Run Glue Crawlers": { "End": true, "ItemsPath": "$.output.crawlers", "Iterator": { "StartAt": "StartCrawler", "States": { "GetCrawler": { "Next": "Is Running?", "Parameters": { "Name.$": "$.crawler_name" }, "Resource": "arn:aws:states:::aws-sdk:glue:getCrawler", "ResultPath": "$.response.get_crawler", "Retry": [ { "BackoffRate": 2, "ErrorEquals": [ "States.ALL" ], "IntervalSeconds": 1, "MaxAttempts": 8 } ], "Type": "Task" }, "Is Running?": { "Choices": [ { "Next": "Wait for Crawler To Complete", "Or": [ { "StringEquals": "RUNNING", "Variable": "$.response.get_crawler.Crawler.State" }, { "StringEquals": "STOPPING", "Variable": "$.response.get_crawler.Crawler.State" } ] } ], "Default": "Prepare Output", "Type": "Choice" }, "Prepare Output": { "End": true, "Parameters": { "LastCrawl.$": "$.response.get_crawler.Crawler.LastCrawl", "crawler_name.$": "$.crawler_name" }, "Type": "Pass" }, "StartCrawler": { "Catch": [ { "Comment": "Crawler Already Running, just continue to monitor", "ErrorEquals": [ "Glue.CrawlerRunningException" ], "Next": "GetCrawler", "ResultPath": "$.response.start_crawler" } ], "Next": "GetCrawler", "Parameters": { "Name.$": "$.crawler_name" }, "Resource": "arn:aws:states:::aws-sdk:glue:startCrawler", "ResultPath": "$.response.start_crawler", "Retry": [ { "BackoffRate": 1, "Comment": "EntityNotFoundException - Fail immediately", "ErrorEquals": [ "Glue.EntityNotFoundException" ], "IntervalSeconds": 1, "MaxAttempts": 0 }, { "BackoffRate": 1, "ErrorEquals": [ "Glue.CrawlerRunningException" ], "IntervalSeconds": 1, "MaxAttempts": 0 } ], "Type": "Task" }, "Wait for Crawler To Complete": { "Next": "GetCrawler", "Seconds": 5, "Type": "Wait" } } }, "Type": "Map" } } }
内容的提问来源于stack exchange,提问作者Mayura
相关产品推荐
相关产品推荐

