You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Apache Hudi中插入struct、map类型数据?

插入Apache Hudi复杂类型数据的语法示例

Hudi兼容Spark SQL的复杂类型语法,直接用Spark自带的构造函数就能插入struct、map这类字段,以下是针对你表结构的具体示例:

Spark SQL 插入方式

直接使用STRUCT()和MAP()函数构造复杂类型数据,插入语句如下:

INSERT INTO struct_map (uuid, col1, col2)
VALUES 
(
  1,
  STRUCT(11, STRUCT(121)),
  MAP('key1', 100, 'key2', 200)
),
(
  2,
  STRUCT(22, STRUCT(221)),
  MAP('key3', 300)
);
  • 注意STRUCT()的参数顺序要和表定义的struct字段顺序严格对应,比如col1的子字段是col11、col12,所以先传col11的值,再传col12对应的嵌套struct。
  • MAP()的参数是连续的键值对,每一个键后面跟对应的值即可。

DataFrame 插入方式(Scala)

如果用代码通过DataFrame写入,直接构造包含复杂类型的数据集即可:

import org.apache.spark.sql.Row
import org.apache.spark.sql.types._

// 构造符合表结构的数据
val data = Seq(
  Row(3, Row(33, Row(321)), Map("key4" -> 400, "key5" -> 500)),
  Row(4, Row(44, Row(421)), Map("key6" -> 600))
)

// 定义与Hudi表匹配的Schema
val schema = StructType(Array(
  StructField("uuid", IntegerType),
  StructField("col1", StructType(Array(
    StructField("col11", IntegerType),
    StructField("col12", StructType(Array(StructField("col121", IntegerType))))
  ))),
  StructField("col2", MapType(StringType, IntegerType))
))

val df = spark.createDataFrame(spark.sparkContext.parallelize(data), schema)

// 写入Hudi表(根据实际配置调整参数)
df.write.format("hudi")
  .option("hoodie.table.name", "struct_map")
  .option("hoodie.datasource.write.recordkey.field", "uuid")
  .option("hoodie.datasource.write.partitionpath.field", "")
  .mode("append")
  .saveAsTable("struct_map")

简单来说,Hudi对Spark支持的所有复杂类型都兼容,插入逻辑和普通Spark表完全一致,官方文档没单独列出这类示例是因为复用了Spark的标准语法。

内容的提问来源于stack exchange,提问作者Smith Cruise

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 01:50:34