You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas dataframe追加写入Azure Databricks表报类型不兼容错误如何解决

Azure Databricks Delta表追加数据类型不匹配问题解决方案

出现该报错的核心原因有三个:

  1. 写入函数中insertInto方法固定配置了overwrite=True,会覆盖传入的write_mode='append'参数,且该写法会强制触发严格的 schema 校验
  2. pandas 转 Spark DataFrame 时默认的类型推断逻辑不稳定:你在 pandas 侧设置的float32对应 Spark 的FloatType,和目标 Delta 表的DoubleType不匹配;DATE字段容易被推断为数值型的时间戳毫秒数,和目标表的TimestampType不匹配
  3. insertInto默认按字段位置而非字段名匹配 schema,只要待写入 DataFrame 字段顺序和目标表不一致,就会出现类型错位报错

步骤1:修正写入函数的参数错误

把insertInto中写死的overwrite=True去掉,由mode()方法统一控制写入模式:

def write_to_delta(df_name, db_name, table_name, write_mode, num_part=10):
    df_name \
        .repartition(num_part) \
        .write \
        .mode(write_mode) \
        .insertInto("{}.{}".format(db_name, table_name))

步骤2:显式对齐目标表的schema,避免类型推断错误

不要直接用无 schema 参数的spark.createDataFrame,先获取目标 Delta 表的标准 schema,再用该 schema 创建 Spark DataFrame:

# 先获取目标表的完整schema和字段顺序
target_table = spark.table("production.feed_to_output_all_features")
target_schema = target_table.schema
target_cols = target_table.columns

# 预处理pandas数据类型,提前对齐目标类型要求
df_allfeatures = df_allfeatures.astype({
    "LEAD_CONCENTRATE_GRADES_PB": "float64",
    "TAILINGS_RECOVERIES_PB": "float64"
})
df_allfeatures["DATE"] = pd.to_datetime(df_allfeatures["DATE"])

# 用目标表schema创建Spark DataFrame,强制匹配类型
df_allfeatures_spark = spark.createDataFrame(df_allfeatures, schema=target_schema)

# 对齐字段顺序,避免insertInto按位置匹配出错
df_allfeatures_spark = df_allfeatures_spark.select(*target_cols)

步骤3:执行写入

调用修正后的写入函数即可:

write_to_delta(df_allfeatures_spark, 'production', 'feed_to_output_all_features', 'append', num_part=10)

内容的提问来源于stack exchange,提问作者kdp132

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 17:39:03