Databricks Photon读取S3大JSON内存不足,求配置方案
问题:启用Photon后读取大JSON目录触发内存不足错误
在启用Photon的Databricks Spark环境中,尝试读取S3上约100-120GB的JSON文件目录时出现内存不足错误,关闭Photon后操作可正常执行。
使用的代码:
df = (spark .read .option("multiline", True) .json(json_dir_path) .schema(json_schema) ) df.write.format('delta').saveAsTable("table_name")
报错信息:
Caused by: org.apache.spark.memory.SparkOutOfMemoryError: Photon ran out of memory while executing this query. Photon failed to reserve 1728.0 MiB for simdjson internal usage, in SimdJsonReader, in JsonFileScanNode(id=3618, output_schema=[string, string, array<struct<string, string, string, string>>, string, string, ... 1 more]), in task.
解决方案
针对Photon处理大JSON文件的内存问题,可以通过以下配置调整和优化手段解决:
调整Photon JSON读取器的内存限制:
修改spark.databricks.photon.json.reader.maxMemoryInMiB配置,降低单任务的内存预留值(比如从默认的1728MiB调整为1024MiB),适配集群的内存资源:spark.conf.set("spark.databricks.photon.json.reader.maxMemoryInMiB", "1024")减小单个任务处理的数据量:
通过调整文件分区大小,让每个任务处理的数据量更少:- 设置
spark.sql.files.maxPartitionBytes降低单分区大小(比如设为64MB):spark.conf.set("spark.sql.files.maxPartitionBytes", "67108864") - 读取后主动 repartition 增加并行度,分散内存压力:
df = (spark.read.option("multiline", True).json(json_dir_path).schema(json_schema)).repartition(200)
- 设置
优化Executor内存配置:
增加Executor的堆外内存开销,Photon的部分内存使用属于堆外范围,调整spark.executor.memoryOverhead(建议设为Executor内存的20%-30%,或固定值如8GB):spark.conf.set("spark.executor.memoryOverhead", "8g")预处理拆分大JSON文件:
如果原始目录中存在单个几十GB的大JSON文件,先将其拆分为1-2GB的小文件,从根源上降低单任务的数据处理量,更适配Photon的执行模型。优先使用单行JSON格式:
如果业务允许,将JSON文件转换为每行一个对象的格式,关闭multiline选项。单行JSON的读取对Photon来说内存效率更高,避免加载整个大文件到内存中解析。
内容的提问来源于stack exchange,提问作者Hemanth Kumar
相关产品推荐
相关产品推荐

