You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法将Spark DataFrame保存至Impala的问题排查求助

问题:大数据量下Spark写入Impala出现EOF/JSON解析等错误

操作流程

  • 从Impala加载含1个字符串列、共10万行的Spark DataFrame到Jupyter Notebook
  • 通过PySpark UDF生成第二列,此时DataFrame包含两个字符串类型列
  • 执行df.write.format("parquet").saveAsTable将结果保存至Impala

报错情况

大数据量(10万行)执行步骤3时触发以下报错:

  • EOF: read didn't return enough bytes;
  • unpack requires a buffer of 4 bytes;
  • task failed while writing rows (json.decoder.JSONDecodeError: Unterminated string started at line 3 column 116133)

注:小数据量(1万行)下执行全流程无报错

当前Spark会话配置

.config('spark.driver.memory', '20g')\
.config('spark.executor.cores', 2)\
.config('spark.executor.memory', '20g')\
.config('spark.dynamicAllocation.enabled', True)\
.config('spark.dynamicAllocation.maxExecutors', 20)\
.config('spark.yarn.executor.memoryOverhead', '4095')\
.config('spark.serializer', 'org.apache.spark.serializer.Kryoserializer')

疑问

  • 导致上述报错的可能原因是什么?
  • 是否需要添加或调整额外的Spark会话配置?

补充说明:已尝试调整Spark配置,PySpark UDF运行正常(执行df.show约耗时5分钟),但无法理解为何JVM无法将DataFrame保存到Impala。

内容的提问来源于stack exchange,提问作者Peter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 14:18:10