You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Databricks Photon读取S3大JSON内存不足,求配置方案

问题:启用Photon后读取大JSON目录触发内存不足错误

在启用Photon的Databricks Spark环境中,尝试读取S3上约100-120GB的JSON文件目录时出现内存不足错误,关闭Photon后操作可正常执行。

使用的代码:

df = (spark
         .read
         .option("multiline", True)
         .json(json_dir_path)
         .schema(json_schema)
        )

df.write.format('delta').saveAsTable("table_name") 

报错信息:

Caused by: org.apache.spark.memory.SparkOutOfMemoryError: Photon ran out of memory while executing this query.
Photon failed to reserve 1728.0 MiB for simdjson internal usage, in SimdJsonReader, in JsonFileScanNode(id=3618, output_schema=[string, string, array<struct<string, string, string, string>>, string, string, ... 1 more]), in task.

解决方案

针对Photon处理大JSON文件的内存问题,可以通过以下配置调整和优化手段解决:

  • 调整Photon JSON读取器的内存限制:
    修改spark.databricks.photon.json.reader.maxMemoryInMiB配置,降低单任务的内存预留值(比如从默认的1728MiB调整为1024MiB),适配集群的内存资源:

    spark.conf.set("spark.databricks.photon.json.reader.maxMemoryInMiB", "1024")
    
  • 减小单个任务处理的数据量:
    通过调整文件分区大小,让每个任务处理的数据量更少:

    • 设置spark.sql.files.maxPartitionBytes降低单分区大小(比如设为64MB):
      spark.conf.set("spark.sql.files.maxPartitionBytes", "67108864")
      
    • 读取后主动 repartition 增加并行度,分散内存压力:
      df = (spark.read.option("multiline", True).json(json_dir_path).schema(json_schema)).repartition(200)
      
  • 优化Executor内存配置:
    增加Executor的堆外内存开销,Photon的部分内存使用属于堆外范围,调整spark.executor.memoryOverhead(建议设为Executor内存的20%-30%,或固定值如8GB):

    spark.conf.set("spark.executor.memoryOverhead", "8g")
    
  • 预处理拆分大JSON文件:
    如果原始目录中存在单个几十GB的大JSON文件,先将其拆分为1-2GB的小文件,从根源上降低单任务的数据处理量,更适配Photon的执行模型。

  • 优先使用单行JSON格式:
    如果业务允许,将JSON文件转换为每行一个对象的格式,关闭multiline选项。单行JSON的读取对Photon来说内存效率更高,避免加载整个大文件到内存中解析。

内容的提问来源于stack exchange,提问作者Hemanth Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 21:23:34