You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Terraform中jsondecode未按JSON文件指定顺序读取列的问题

解决Terraform创建Glue表时列顺序与JSON文件不一致的问题

问题场景

现有如下emp.json文件:

{
    "columns": {
        "%KEY_EMP_ID": "string",
        "EMP_DEP_ID": "string",
        "EMP_FULLNAME": "string",
        "EMP_STAFFNUMBER": "string",
        "EMP_EMAIL_ADDRESS": "string",
        "EMP_LOCATION_NAME": "string",
        "EMP_GRADE_NAME": "string",
        "EMP_GRADE_BAND": "string",
        "EMP_SAL": "string"
    }
}

使用Terraform的aws_glue_catalog_table资源创建Glue表后,表的列顺序和JSON文件中定义的顺序不一致。

问题原因

JSON对象的键值对本身不保证有序,Terraform的jsondecode函数解析JSON对象后会转换成无序的map类型。当用dynamic "columns"遍历这个map时,遍历顺序是随机的,导致最终Glue表的列顺序和JSON文件不一致。

解决方案

方案1:修改JSON文件为数组结构(推荐)

将emp.json中的columns从对象改为数组,数组元素包含name和type字段,天然保留顺序:

{
    "columns": [
        {"name": "%KEY_EMP_ID", "type": "string"},
        {"name": "EMP_DEP_ID", "type": "string"},
        {"name": "EMP_FULLNAME", "type": "string"},
        {"name": "EMP_STAFFNUMBER", "type": "string"},
        {"name": "EMP_EMAIL_ADDRESS", "type": "string"},
        {"name": "EMP_LOCATION_NAME", "type": "string"},
        {"name": "EMP_GRADE_NAME", "type": "string"},
        {"name": "EMP_GRADE_BAND", "type": "string"},
        {"name": "EMP_SAL", "type": "string"}
    ]
}

然后修改Terraform代码中的dynamic "columns"部分,遍历有序的数组:

resource "aws_glue_catalog_table" "aws_glue_catalog_table" {
  name          = "emp_table"
  database_name = "emp_db"
  table_type    = "EXTERNAL_TABLE"

  parameters = {
    EXTERNAL              = "TRUE"
    "parquet.compression" = "SNAPPY"
  }

  storage_descriptor {
    location      = "s3://my-bucket-emp/output/emp-stream"
    input_format  = "org.apache.hadoop.hive.ql.io.parquet.MapredParquetInputFormat"
    output_format = "org.apache.hadoop.hive.ql.io.parquet.MapredParquetOutputFormat"

    ser_de_info {
      serialization_library = "org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe"
      parameters = {
        "serialization.format" = 1
      }
    }

    dynamic "columns" {
      for_each = jsondecode(file("${local.glue_config_path}/emp.json")).columns
      content {
        name = columns.value.name
        type = columns.value.type
      }
    }
  }
}

方案2:在Terraform中手动指定列顺序(不修改JSON)

如果无法修改原JSON文件,可在Terraform中定义一个有序的列名列表,遍历该列表并从JSON解析后的map中获取对应类型:

resource "aws_glue_catalog_table" "aws_glue_catalog_table" {
  name          = "emp_table"
  database_name = "emp_db"
  table_type    = "EXTERNAL_TABLE"

  # 定义有序的列名列表,和JSON中的顺序一致
  locals {
    ordered_column_names = [
      "%KEY_EMP_ID",
      "EMP_DEP_ID",
      "EMP_FULLNAME",
      "EMP_STAFFNUMBER",
      "EMP_EMAIL_ADDRESS",
      "EMP_LOCATION_NAME",
      "EMP_GRADE_NAME",
      "EMP_GRADE_BAND",
      "EMP_SAL"
    ]
    emp_columns = jsondecode(file("${local.glue_config_path}/emp.json")).columns
  }

  parameters = {
    EXTERNAL              = "TRUE"
    "parquet.compression" = "SNAPPY"
  }

  storage_descriptor {
    location      = "s3://my-bucket-emp/output/emp-stream"
    input_format  = "org.apache.hadoop.hive.ql.io.parquet.MapredParquetInputFormat"
    output_format = "org.apache.hadoop.hive.ql.io.parquet.MapredParquetOutputFormat"

    ser_de_info {
      serialization_library = "org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe"
      parameters = {
        "serialization.format" = 1
      }
    }

    dynamic "columns" {
      for_each = local.ordered_column_names
      content {
        name = each.value
        type = local.emp_columns[each.value]
      }
    }
  }
}

总结

方案1通过将JSON结构改为数组,从根源上保证列顺序,更直观且易于维护;方案2无需修改原JSON,但需要手动维护列名列表,适合无法修改JSON文件的场景。

内容的提问来源于stack exchange,提问作者J R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 06:43:14