You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark调用saveAsTextFile时持续报FileAlreadyExistsException错误的解决求助

PySpark调用saveAsTextFile时持续报FileAlreadyExistsException错误的解决求助

我现在在Ubuntu上运行一个PySpark单词统计脚本,脚本的单词统计逻辑能正常执行,但每次调用saveAsTextFile保存结果时都会报错,提示输出目录已经存在。

具体错误信息如下:

py4j.protocol.Py4JJavaError: An error occurred while calling o48.saveAsTextFile. 
org.apache.hadoop.mapred.FileAlreadyExistsException: Output directory file:/home/pyspark_python/wordcount/output_new already exists

我已经尝试了以下排查和解决方法,但错误依然存在:

  • 确认输出目录是空的,用ls命令检查过目录内没有任何数据
  • 用rm -r彻底删除目录后,再通过mkdir -p重新创建空目录
  • 用ps aux | grep spark确认当前没有其他Spark任务在运行

我的PySpark代码如下:

from pyspark import SparkConf, SparkContext
import os

def main(input_file, output_dir):
    # Configuration Spark
    conf = SparkConf().setAppName("WordCountTask").setMaster("local[*]")
    sc = SparkContext(conf=conf)

    # Lecture du fichier d'entrée
    text_file = sc.textFile(input_file)

    # Comptage des mots
    counts = (
        text_file.flatMap(lambda line: line.split(" "))
                 .map(lambda word: (word, 1))
                 .reduceByKey(lambda a, b: a + b)
    )

    # Sauvegarde des résultats
    if not os.path.exists(output_dir):
        os.makedirs(output_dir)
    counts.saveAsTextFile(output_dir)

    print(f"Résultats sauvegardés dans le répertoire : {output_dir}")

if __name__ == "__main__":
    # Définir les chemins d'entrée et de sortie
    input_file = r"/home/othniel/pyspark_python/wordcount/input/loremipsum.txt"
    output_dir = "/home/othniel/pyspark_python/wordcount/output_new"

    # Exécution de la tâche WordCount
    main(input_file, output_dir)

请问我该如何解决这个错误,让PySpark成功将结果写入输出目录?是脚本逻辑有需要调整的地方,还是环境配置存在问题?

谢谢大家的帮助!

备注:内容来源于stack exchange,提问作者Fractal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 12:18:07