You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Palantir Foundry代码工作簿:导出单个XML并打包为Zip

在Palantir Foundry中导出XML为独立文件并打包Zip

需求说明

已有包含key和xml列的数据集,通过Code Workbook过滤采样后,需要将每行的XML内容保存为以key值命名的独立文件,并打包成Zip包存储到Foundry中。

已完成的过滤采样代码

def prepare_input(xml_with_debug):
    from pyspark.sql import functions as F

    filter_column = "key"
    filter_value = "test_key"
    df_filtered = xml_with_debug.filter(filter_value == F.col(filter_column))

    approx_number_of_rows = 1
    sample_percent = float(approx_number_of_rows) / df_filtered.count()

    df_sampled = df_filtered.sample(False, sample_percent, seed=0)

    important_columns = ["key", "xml"]

    return df_sampled.select([F.col(c).cast(F.StringType()).alias(c) for c in important_columns])

导出并打包Zip的write_file代码

def write_file(df):
    from pyspark.sql import functions as F
    import zipfile
    from io import BytesIO
    from py4j.java_gateway import java_import
    from pyspark.sql.types import StringType

    # 导入Hadoop文件系统依赖
    java_import(spark._jvm, "org.apache.hadoop.fs.FileSystem")
    java_import(spark._jvm, "org.apache.hadoop.fs.Path")

    # 替换为你的Foundry输出数据集路径
    output_path = "/foundry/datasets/your-project/your-output-zip-dataset"
    full_zip_path = f"{output_path}/xml_files.zip"

    # 收集采样后的小数据集到Driver节点(仅适合少量数据)
    xml_records = df.collect()

    # 内存中构建Zip包
    zip_buffer = BytesIO()
    with zipfile.ZipFile(zip_buffer, "w", zipfile.ZIP_DEFLATED) as zip_handler:
        for record in xml_records:
            # 处理文件名特殊字符,避免路径错误
            safe_key = record["key"].replace('/', '_').replace('\\', '_')
            zip_handler.writestr(f"{safe_key}.xml", record["xml"])
    
    # 将Zip包写入Foundry文件系统
    zip_buffer.seek(0)
    hadoop_conf = spark._jsc.hadoopConfiguration()
    fs = spark._jvm.FileSystem.get(hadoop_conf)
    output_stream = fs.create(spark._jvm.Path(full_zip_path))
    output_stream.write(zip_buffer.getvalue())
    output_stream.close()

    # 返回空DataFrame满足Code Workbook任务要求
    return spark.createDataFrame([], StringType())

使用说明

  1. 在Code Workbook中先执行prepare_input任务,传入原始数据集得到过滤采样后的DataFrame
  2. 将write_file任务的输入设为prepare_input的输出,替换output_path为你实际的输出数据集路径
  3. 运行write_file任务后,即可在指定的Foundry数据集中找到生成的Zip包

注意事项

  • 该方法使用collect()将数据拉到Driver节点,仅适合采样后的小数据集,如果处理大量数据需改用分布式打包方案
  • 对key值做了特殊字符替换,避免文件名包含路径分隔符导致的错误

内容的提问来源于stack exchange,提问作者asb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 03:40:21