You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用AWS Glue将Spark DataFrame写入S3时出现文件已存在错误

Spark写入S3时持续报文件已存在错误

我使用以下代码将DataFrame写入S3:

df.write.option("delimiter","|").option("header",True).option("compression", "gzip").mode("overwrite").format("csv").save("s3://bucketname/metrics/parsed/")

但始终出现文件已存在错误,每次报错的文件名不同:

An error occurred while calling o293.save. File already exists:s3://bucketname/metrics/parsed/part-01195-6ef08750-dbf5-41c6-b024-501403820268-c000.csv.gz

完整错误信息:

"Failure Reason": "JobFailed(org.apache.spark.SparkException: Job aborted due to stage failure: Task 1195 in stage 11.0 failed 4 times, most recent failure: 
Lost task 1195.3 in stage 11.0 (TID 3023) (172.36.67.235 executor 9):
 org.apache.hadoop.fs.FileAlreadyExistsException: File already exists

我尝试了以下方法,但均无法解决问题,仍出现相同错误:

  • 在命令中添加coalesce(100)
  • 写入新的目标路径,无论是否添加.mode("overwrite")选项
  • 改用Parquet格式导出数据
  • 使用.mode("append")选项写入

仅找到一篇IBM支持文章,但我使用的是Glue 3.0(Spark 3.1),该文章内容不适用。


内容的提问来源于stack exchange,提问作者Shailesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 14:05:34