无法将Spark DataFrame保存至Impala的问题排查求助
问题:大数据量下Spark写入Impala出现EOF/JSON解析等错误
操作流程
- 从Impala加载含1个字符串列、共10万行的Spark DataFrame到Jupyter Notebook
- 通过PySpark UDF生成第二列,此时DataFrame包含两个字符串类型列
- 执行
df.write.format("parquet").saveAsTable将结果保存至Impala
报错情况
大数据量(10万行)执行步骤3时触发以下报错:
- EOF: read didn't return enough bytes;
- unpack requires a buffer of 4 bytes;
- task failed while writing rows (json.decoder.JSONDecodeError: Unterminated string started at line 3 column 116133)
注:小数据量(1万行)下执行全流程无报错
当前Spark会话配置
.config('spark.driver.memory', '20g')\ .config('spark.executor.cores', 2)\ .config('spark.executor.memory', '20g')\ .config('spark.dynamicAllocation.enabled', True)\ .config('spark.dynamicAllocation.maxExecutors', 20)\ .config('spark.yarn.executor.memoryOverhead', '4095')\ .config('spark.serializer', 'org.apache.spark.serializer.Kryoserializer')
疑问
- 导致上述报错的可能原因是什么?
- 是否需要添加或调整额外的Spark会话配置?
补充说明:已尝试调整Spark配置,PySpark UDF运行正常(执行df.show约耗时5分钟),但无法理解为何JVM无法将DataFrame保存到Impala。
内容的提问来源于stack exchange,提问作者Peter
相关产品推荐
相关产品推荐

