Azure Synapse Notebook写入XML数据到Hudi表时遇类型错误
问题解决:写入Hudi表时的TypeError错误
错误根源
你代码里用的option()是Spark DataFrameWriter用来设置单个配置项的方法,它只接受key和value两个参数,没法直接用**hudi_options解包传入多组配置。要批量传递多个Hudi参数,必须用复数形式的options()方法。
修改后的代码
把写入Hudi的代码行替换为以下版本:
basepath = "abfs://XXXXXXXXXXXXXXXXXXXXXXXX/huditables/" table="hudiTable" hudi_options = { 'hoodie.datasource.write.recordkey.field': 'id', 'hoodie.datasource.write.operation': 'upsert', 'hoodie.datasource.write.precombine.field': 'id', 'hoodie.table.name': table } # 读取XML并生成DataFrame df的代码保持不变 # 关键修改:将option替换为options df.write.format("hudi").options(**hudi_options).mode("overwrite").save(basepath)
额外注意事项
确认你的DataFrame df中确实存在名为id的字段——因为你配置里指定它作为recordkey和precombine字段,要是字段不存在,后续会触发新的报错。
内容的提问来源于stack exchange,提问作者Vishal Patwardhan
相关产品推荐
相关产品推荐

