Spark DataFrame保存为Parquet无报错却未生成文件,求解决方案
Hey there! I’ve run into this exact head-scratcher a few times with Spark, so let’s break down the most common reasons this happens and how to fix them:
Spark’s lazy evaluation didn’t trigger the write
Spark runs on lazy execution—even thoughwrite.parquet()is supposed to be an action that kicks off computation, sometimes unevaluated transformations in your DataFrame chain or a failed job submission mean the write never actually runs.
Fixes: Add a quickdf.count()ordf.show(5)before the write to force execution and confirm your DataFrame has data. Also check your cluster’s Spark/YARN UI to see if the write job even started.Permission issues on the output directory
Spark often won’t throw a clear error if the user running the job lacks write access to the target path—especially if the error happens on executor nodes instead of the driver.
Fixes: Manually test writing a small file to the target directory using the same cluster user. Adjust permissions (e.g.,chmod 775 /your/output/path) or update your cluster’s access controls to grant write access.Your DataFrame is empty
If there’s no data in your DataFrame, Spark will create the output directory but won’t generate any Parquet files (and sometimes not even a_SUCCESSmarker).
Fixes: Rundf.isEmpty()ordf.count()to verify data exists. Double-check your data loading/transformation code—did a filter or join accidentally remove all rows?Cluster resources are insufficient (job was silently killed)
If your cluster doesn’t have enough memory or CPU to run the write job, managers like YARN might terminate it without sending a clear error to the driver.
Fixes: Check your cluster’s resource manager logs (YARN ApplicationMaster, NodeManager logs) for termination messages. Adjust Spark configs likespark.executor.memoryorspark.executor.coresto fit your cluster’s available resources.You’re using
mode("ignore")on an existing directory
If the target directory already exists and you setdf.write.mode("ignore").parquet(...), Spark will skip writing entirely without throwing an error.
Fixes: Switch tomode("overwrite")(if you want to replace existing data) ormode("append")(if you want to add to it). Alternatively, delete the target directory before running the write.Hidden errors in executor logs
Your driver logs might not show everything—executor nodes could be hitting issues like HDFS replication failures or disk space limits that don’t bubble up to the driver.
Fixes: Dig into the full executor logs for WARN/ERROR messages. Look for lines like "Failed to write partition" or "Insufficient replicas for HDFS block".
If you can share snippets of your write code or specific log entries, that’ll help narrow things down even more!
内容的提问来源于stack exchange,提问作者Vitrion

