You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将PySpark输出文件大小压缩至与Hive输出一致?

如何让PySpark生成与Hive大小一致的ORC Snappy输出文件

我在使用PySpark和Hive向Google Cloud Storage Bucket写入ORC Snappy文件时,发现Hive生成的单个输出文件明显小于PySpark的输出:Hive输出为72.9MB的单个文件,PySpark最初输出9个文件总计86.4MB,即使合并为单个文件后,压缩效果仍不如Hive。需要调整PySpark配置,让其生成的ORC文件大小与Hive保持一致。


测试用代码

Hive侧操作

  1. 创建基于CSV的源表
drop table if exists dummy_data;
create external table if not exists dummy_data (
column1 int,
column2 int,
column3 int,
column4 int,
column5 int,
column6 int)
row format delimited
fields terminated by ','
stored as textfile
location 'gs://dummy_bucket_mlw/dummy_data/'
tblproperties("skip.header.line.count"="1");
  1. 创建ORC格式的目标表
drop table if exists hive_test_table;
create external table if not exists hive_test_table (
column1 int,
column2 int,
column3 int,
column4 int,
column5 int,
column6 int)
stored as orc
location 'gs://dummy_bucket_mlw/hive_test_table'
tblproperties ('ORC.COMPRESS' = 'SNAPPY');
  1. 插入数据
INSERT INTO hive_test_table
SELECT *
FROM dummy_data;

PySpark侧操作

  1. 创建对应ORC表
drop table if exists pyspark_test_table;
create external table if not exists pyspark_test_table (
column1 int,
column2 int,
column3 int,
column4 int,
column5 int,
column6 int)
stored as orc
location 'gs://dummy_bucket_mlw/pyspark_test_table'
tblproperties ('ORC.COMPRESS' = 'SNAPPY');
  1. PySpark写入代码
import findspark
findspark.init()
from pyspark.sql import SparkSession
from google.cloud import storage

client = storage.Client()
xr = client.get_bucket("dummy_bucket_mlw")

spark = SparkSession.builder \
    .appName("pyspark_test_load") \
    .enableHiveSupport() \
    .getOrCreate()
spark.conf.set("hive.exec.dynamic.partition.mode", 'nonstrict')
spark.conf.set("hive.server2.builtin.udf.blacklist", "emply_bl")

# 读取数据
my_query = """
    SELECT * FROM dummy_data
"""
df = spark.sql(my_query)

# 写入ORC表
output_table = "pyspark_test_table"
df.repartition(1).write.format("orc").mode("append").option("compression", "snappy").insertInto(output_table)

解决方法

问题根源在于PySpark和Hive的ORC默认配置参数存在差异,导致压缩效率不同。通过以下配置调整可对齐两者的行为:

  1. 对齐ORC核心配置参数
    在SparkSession初始化时添加以下配置,强制使用与Hive一致的ORC压缩、分块和条带参数:
spark = SparkSession.builder \
    .appName("pyspark_test_load") \
    .enableHiveSupport() \
    .config("spark.sql.orc.compression.codec", "snappy") \
    .config("spark.sql.orc.stripe.size", "67108864")  # Hive默认ORC条带大小为64MB
    .config("spark.sql.orc.block.size", "268435456")  # Hive默认ORC块大小为256MB
    .config("spark.sql.hive.convertMetastoreOrc", "false")  # 禁用Spark原生ORC实现,改用Hive兼容逻辑
    .getOrCreate()
  1. 优化文件合并与写入方式
    避免使用repartition(1)强制重分区(会打乱数据分布,影响压缩效率),如果需要单个文件,改用更高效的coalesce(1);同时优先使用saveAsTable而非insertInto,确保元数据与Hive表对齐:
df.coalesce(1).write \
    .format("orc") \
    .mode("append") \
    .option("compression", "snappy") \
    .option("orc.compress", "snappy") \
    .saveAsTable(output_table)
  1. 统一表属性配置
    确保PySpark目标表的ORC属性与Hive表完全一致:
ALTER TABLE pyspark_test_table SET TBLPROPERTIES (
    'ORC.COMPRESS'='SNAPPY',
    'orc.stripe.size'='67108864',
    'orc.block.size'='268435456'
);

通过以上调整,PySpark生成的ORC Snappy文件大小会与Hive输出基本一致。

内容的提问来源于stack exchange,提问作者Matt Weisman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 20:15:15