如何使用制表符(' ')分隔符保存PySpark2 DataFrame为文本格式?
Hey there! Saving a PySpark 2 DataFrame as a tab-separated text file is straightforward, and I'll walk you through three practical approaches using the DataFrame API, SQL, and RDD respectively. Let's dive in:
This is the simplest and most efficient method since it leverages the built-in DataFrame writer, which handles edge cases like data type conversion and header inclusion out of the box. Even though we use the csv format, it fully supports custom delimiters—perfect for tab-separated files.
# Assume your DataFrame is named `df` df.write \ .option("sep", "\t") # Set tab as the field delimiter .option("header", "true") # Include column headers (remove this line if not needed) .mode("overwrite") # Choose write mode: overwrite/append/ignore/error .csv("/path/to/your/output/folder")
Note: PySpark will output a directory with partitioned files by default. If you need a single file, you can add
.coalesce(1)beforewrite, but avoid this for large datasets as it can hurt performance.
If you prefer working with SQL syntax, you can register your DataFrame as a temporary view and use SQL commands to write the tab-separated output.
# Register the DataFrame as a temporary view df.createOrReplaceTempView("my_temp_table") # Execute SQL to write the file with tab delimiter spark.sql(""" INSERT OVERWRITE DIRECTORY '/path/to/your/output/folder' ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t' SELECT * FROM my_temp_table """)
Tip: Replace
OVERWRITEwithAPPENDif you want to add data to an existing directory instead of replacing its contents.
If you need granular control over how each row is formatted, convert the DataFrame to an RDD, manually format rows as tab-separated strings, then save the result.
Basic version (without headers):
# Convert DataFrame to RDD, format each row as a tab-separated string tab_separated_rdd = df.rdd.map(lambda row: "\t".join(str(col) for col in row)) # Save the RDD as text files tab_separated_rdd.saveAsTextFile("/path/to/your/output/folder")
With custom headers:
To include column headers, manually create a header row and union it with the data RDD:
# Create header string from DataFrame column names header_row = "\t".join(df.columns) # Combine header RDD with data RDD final_rdd = spark.sparkContext.parallelize([header_row]).union(tab_separated_rdd) # Save the combined output final_rdd.saveAsTextFile("/path/to/your/output/folder")
内容的提问来源于stack exchange,提问作者kranthi kumar

