如何将PySpark输出文件大小压缩至与Hive输出一致?
如何让PySpark生成与Hive大小一致的ORC Snappy输出文件
我在使用PySpark和Hive向Google Cloud Storage Bucket写入ORC Snappy文件时,发现Hive生成的单个输出文件明显小于PySpark的输出:Hive输出为72.9MB的单个文件,PySpark最初输出9个文件总计86.4MB,即使合并为单个文件后,压缩效果仍不如Hive。需要调整PySpark配置,让其生成的ORC文件大小与Hive保持一致。
测试用代码
Hive侧操作
- 创建基于CSV的源表
drop table if exists dummy_data; create external table if not exists dummy_data ( column1 int, column2 int, column3 int, column4 int, column5 int, column6 int) row format delimited fields terminated by ',' stored as textfile location 'gs://dummy_bucket_mlw/dummy_data/' tblproperties("skip.header.line.count"="1");
- 创建ORC格式的目标表
drop table if exists hive_test_table; create external table if not exists hive_test_table ( column1 int, column2 int, column3 int, column4 int, column5 int, column6 int) stored as orc location 'gs://dummy_bucket_mlw/hive_test_table' tblproperties ('ORC.COMPRESS' = 'SNAPPY');
- 插入数据
INSERT INTO hive_test_table SELECT * FROM dummy_data;
PySpark侧操作
- 创建对应ORC表
drop table if exists pyspark_test_table; create external table if not exists pyspark_test_table ( column1 int, column2 int, column3 int, column4 int, column5 int, column6 int) stored as orc location 'gs://dummy_bucket_mlw/pyspark_test_table' tblproperties ('ORC.COMPRESS' = 'SNAPPY');
- PySpark写入代码
import findspark findspark.init() from pyspark.sql import SparkSession from google.cloud import storage client = storage.Client() xr = client.get_bucket("dummy_bucket_mlw") spark = SparkSession.builder \ .appName("pyspark_test_load") \ .enableHiveSupport() \ .getOrCreate() spark.conf.set("hive.exec.dynamic.partition.mode", 'nonstrict') spark.conf.set("hive.server2.builtin.udf.blacklist", "emply_bl") # 读取数据 my_query = """ SELECT * FROM dummy_data """ df = spark.sql(my_query) # 写入ORC表 output_table = "pyspark_test_table" df.repartition(1).write.format("orc").mode("append").option("compression", "snappy").insertInto(output_table)
解决方法
问题根源在于PySpark和Hive的ORC默认配置参数存在差异,导致压缩效率不同。通过以下配置调整可对齐两者的行为:
- 对齐ORC核心配置参数
在SparkSession初始化时添加以下配置,强制使用与Hive一致的ORC压缩、分块和条带参数:
spark = SparkSession.builder \ .appName("pyspark_test_load") \ .enableHiveSupport() \ .config("spark.sql.orc.compression.codec", "snappy") \ .config("spark.sql.orc.stripe.size", "67108864") # Hive默认ORC条带大小为64MB .config("spark.sql.orc.block.size", "268435456") # Hive默认ORC块大小为256MB .config("spark.sql.hive.convertMetastoreOrc", "false") # 禁用Spark原生ORC实现,改用Hive兼容逻辑 .getOrCreate()
- 优化文件合并与写入方式
避免使用repartition(1)强制重分区(会打乱数据分布,影响压缩效率),如果需要单个文件,改用更高效的coalesce(1);同时优先使用saveAsTable而非insertInto,确保元数据与Hive表对齐:
df.coalesce(1).write \ .format("orc") \ .mode("append") \ .option("compression", "snappy") \ .option("orc.compress", "snappy") \ .saveAsTable(output_table)
- 统一表属性配置
确保PySpark目标表的ORC属性与Hive表完全一致:
ALTER TABLE pyspark_test_table SET TBLPROPERTIES ( 'ORC.COMPRESS'='SNAPPY', 'orc.stripe.size'='67108864', 'orc.block.size'='268435456' );
通过以上调整,PySpark生成的ORC Snappy文件大小会与Hive输出基本一致。
内容的提问来源于stack exchange,提问作者Matt Weisman
相关产品推荐
相关产品推荐

