You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Palantir Foundry中用PySpark合并60+目录数据集为单DataFrame

在Palantir Foundry中用PySpark合并目录下多数据集为单个DataFrame

实现步骤与代码

以下是直接可用的PySpark脚本,用于读取指定目录下的所有数据集并合并输出:

# 定义输入目录与输出路径
input_dir = input("MySpend new files P2P/2020/")
output_path = Output("MySpend new files P2P/2020/UnionAll")

# 获取目录下所有数据集对象
datasets = list(input_dir.datasets())

# 初始化空DataFrame用于合并
combined_df = None

# 遍历并合并所有数据集
for ds in datasets:
    # 读取当前数据集
    df = spark.read.dataset(ds)
    # 首次读取时直接赋值,后续执行合并
    if combined_df is None:
        combined_df = df
    else:
        # 用unionByName适配字段顺序/缺失情况,allowMissingColumns=True允许字段不全的数据集合并
        combined_df = combined_df.unionByName(df, allowMissingColumns=True)

# 将合并后的DataFrame写入输出数据集
output_path.write_dataframe(combined_df)

关键注意事项

  • 字段兼容性:如果各数据集字段不完全一致,unionByName比union更可靠,allowMissingColumns=True会自动为缺失字段填充null
  • 性能优化:若数据集数量多且数据量大,可考虑先对单个数据集调整分区,或在合并后执行repartition()优化后续操作
  • 数据校验:合并前可添加字段类型校验逻辑,避免因数据类型不匹配导致报错

内容的提问来源于stack exchange,提问作者Biplab1985

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 15:37:10