Palantir Foundry代码工作簿:导出单个XML并打包为Zip
在Palantir Foundry中导出XML为独立文件并打包Zip
需求说明
已有包含key和xml列的数据集,通过Code Workbook过滤采样后,需要将每行的XML内容保存为以key值命名的独立文件,并打包成Zip包存储到Foundry中。
已完成的过滤采样代码
def prepare_input(xml_with_debug): from pyspark.sql import functions as F filter_column = "key" filter_value = "test_key" df_filtered = xml_with_debug.filter(filter_value == F.col(filter_column)) approx_number_of_rows = 1 sample_percent = float(approx_number_of_rows) / df_filtered.count() df_sampled = df_filtered.sample(False, sample_percent, seed=0) important_columns = ["key", "xml"] return df_sampled.select([F.col(c).cast(F.StringType()).alias(c) for c in important_columns])
导出并打包Zip的write_file代码
def write_file(df): from pyspark.sql import functions as F import zipfile from io import BytesIO from py4j.java_gateway import java_import from pyspark.sql.types import StringType # 导入Hadoop文件系统依赖 java_import(spark._jvm, "org.apache.hadoop.fs.FileSystem") java_import(spark._jvm, "org.apache.hadoop.fs.Path") # 替换为你的Foundry输出数据集路径 output_path = "/foundry/datasets/your-project/your-output-zip-dataset" full_zip_path = f"{output_path}/xml_files.zip" # 收集采样后的小数据集到Driver节点(仅适合少量数据) xml_records = df.collect() # 内存中构建Zip包 zip_buffer = BytesIO() with zipfile.ZipFile(zip_buffer, "w", zipfile.ZIP_DEFLATED) as zip_handler: for record in xml_records: # 处理文件名特殊字符,避免路径错误 safe_key = record["key"].replace('/', '_').replace('\\', '_') zip_handler.writestr(f"{safe_key}.xml", record["xml"]) # 将Zip包写入Foundry文件系统 zip_buffer.seek(0) hadoop_conf = spark._jsc.hadoopConfiguration() fs = spark._jvm.FileSystem.get(hadoop_conf) output_stream = fs.create(spark._jvm.Path(full_zip_path)) output_stream.write(zip_buffer.getvalue()) output_stream.close() # 返回空DataFrame满足Code Workbook任务要求 return spark.createDataFrame([], StringType())
使用说明
- 在Code Workbook中先执行
prepare_input任务,传入原始数据集得到过滤采样后的DataFrame - 将
write_file任务的输入设为prepare_input的输出,替换output_path为你实际的输出数据集路径 - 运行
write_file任务后,即可在指定的Foundry数据集中找到生成的Zip包
注意事项
- 该方法使用
collect()将数据拉到Driver节点,仅适合采样后的小数据集,如果处理大量数据需改用分布式打包方案 - 对
key值做了特殊字符替换,避免文件名包含路径分隔符导致的错误
内容的提问来源于stack exchange,提问作者asb
相关产品推荐
相关产品推荐

