如何用Terraform配置Glue爬虫与对应表的数据质量规则集?
解决方案:Terraform 同步创建Glue Crawler与数据质量规则集
核心问题分析
当前配置里的depends_on仅能确保Glue Crawler资源被创建,但不会自动触发Crawler执行爬取任务,导致数据质量规则集创建时,目标Catalog表还未生成,最终报错。
实现步骤
- 创建Glue Crawler资源
- 通过
null_resource调用AWS CLI启动Crawler,并等待其爬取完成(状态变为READY) - 确认表生成后,创建数据质量规则集
修改后的完整代码
variable "glue_connection_name" { type = string description = "Name of the Glue Connection to use." } variable "crawler_iam_role_arn" { type = string description = "ARN of the IAM role to use for the table Crawlers." } variable "database_name" { type = string description = "Name of the Glue Catalog Database to use." } variable "table_name" { type = string description = "The name of the table to check as in the RDS database." } variable "ruleset" { type = list(string) description = "The rules to check. Quotes must be escaped: `\\\"`" } variable "qa_crawlers_prefix" { default = "qa_" type = string description = "The prefix QA Glue Crawlers shall set before the automatically generated Glue Catalog table name." } # 创建Glue Crawler资源 resource "aws_glue_crawler" "qa_table_crawler" { database_name = var.database_name name = "qa_${var.database_name}-${var.table_name}" role = var.crawler_iam_role_arn description = "Loads the ${var.database_name}.${var.table_name} table into the Glue Catalog." table_prefix = var.qa_crawlers_prefix jdbc_target { connection_name = var.glue_connection_name path = "${var.database_name}/${var.table_name}" } tags = { Name = "glue" Function = "data processing" } } # 触发Crawler运行并等待爬取完成 resource "null_resource" "run_crawler" { triggers = { crawler_name = aws_glue_crawler.qa_table_crawler.name } provisioner "local-exec" { command = <<EOF aws glue start-crawler --name ${aws_glue_crawler.qa_table_crawler.name} while true; do STATUS=$(aws glue get-crawler --name ${aws_glue_crawler.qa_table_crawler.name} --query 'Crawler.State' --output text) if [ "$STATUS" = "READY" ]; then break fi sleep 10 done EOF interpreter = ["/bin/bash", "-c"] } depends_on = [aws_glue_crawler.qa_table_crawler] } # 引用已生成的Catalog表(确保表存在后加载) data "aws_glue_catalog_table" "crawled_table" { database_name = var.database_name name = "${var.qa_crawlers_prefix}${var.database_name}_${var.table_name}" depends_on = [null_resource.run_crawler] } # 创建数据质量规则集 resource "aws_glue_data_quality_ruleset" "table_rules" { name = var.table_name description = "Checks for the ${var.table_name} table in the ${var.database_name} database" ruleset = "Rules = [${join(",", var.ruleset)}]" target_table { database_name = data.aws_glue_catalog_table.crawled_table.database_name table_name = data.aws_glue_catalog_table.crawled_table.name } tags = { Name = "glue-qa-${var.table_name}" Function = "data processing" } depends_on = [data.aws_glue_catalog_table.crawled_table] }
关键说明
null_resource中的脚本会启动Crawler,并循环检查其状态,直到Crawler完成爬取(状态返回READY),确保目标Catalog表已生成- 使用
data "aws_glue_catalog_table"引用已生成的表,避免硬编码表名可能出现的错误 - 完整的依赖链保证了执行顺序:Crawler创建 → Crawler爬取完成 → 表生成 → 规则集创建
内容的提问来源于stack exchange,提问作者twagner-fox
相关产品推荐
相关产品推荐

