You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark Driver加载JDBC大查询内存耗尽问题及解决方案

Spark JDBC Driver内存耗尽问题排查与解决

问题场景

从远程数据库加载数据时,查询语句包含由50000个ID组成的contact_ids_str,整体语句大小约8MB。使用Spark JDBC驱动读取数据,执行30-40次后出现Driver内存耗尽的情况,排查后确认内存问题并非来自DataFrame本身,初步怀疑是JDBC驱动内存未释放。

初始实现代码

query = f"""
    (SELECT id, key, MAX(time) AS created_at
    FROM remote.table
    WHERE remote_id = {id}
    AND id IN ({contact_ids_str})
    GROUP BY id, key) as subquery
"""

df = spark.read.format("jdbc") \
    .option("url", jdbc_url) \
    .option("dbtable", query) \
    .option("user", jdbc_properties["user"]) \
    .option("password", jdbc_properties["password"]) \
    .option("driver", jdbc_properties["driver"]) \
    .load()

print(f"Query {df.count()} rows returned.")
df.write.csv('/home/slake/checkpoints/temp/', header=True, mode="append")

临时解决方案

通过psycopg2在PostgreSQL中创建视图,再让Spark读取该视图,可避免内存问题,但属于临时workaround:

create_view_query = f"""
    CREATE OR REPLACE VIEW {view_name} AS
    SELECT id, key, MAX(time) AS created_at
    FROM remote.table
    WHERE remote_id = {id}
    AND id IN ({contact_ids_str})
    GROUP BY id, key
"""

最终根因与彻底解决

经排查,问题根源并非JDBC驱动,而是Spark UI日志默认存储在内存中,即使关闭UI,相关日志数据仍会占用内存。通过配置spark.ui.store.path参数,将UI日志存储到指定磁盘目录,彻底解决了内存耗尽问题:

spark.ui.store.path=/home/stats/datalake/tmpdir/ui

内容的提问来源于stack exchange,提问作者Lennart Reus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 17:18:22