Spark读取CSV时如何正确解析带反斜杠转义的逗号字段
Alright, let's tackle this problem you're facing: when reading your CSV file with PySpark, the escaped comma (Ashraful\, Islam) isn't being parsed correctly—you end up with the backslash still in the name field instead of the clean Ashraful, Islam you expect.
Why This Happens
The root issue ties into how Python handles string escaping. When you pass option("escape", "\\"), you might think you're sending a single backslash to Spark, but Python's string processing can sometimes muddle this, especially if you're working with older Spark versions or edge cases in the CSV parser. Additionally, even with the right escape setting, sometimes the parser needs explicit hints to handle quoted values with escaped delimiters properly.
Solutions to Try
1. Correctly Pass the Escape Character (Recommended)
In Python, backslashes are escape characters, so to send a single backslash to Spark, you need to either double-escape it or use a raw string. Try updating your read configuration to use a double-escaped backslash:
test = spark.read.format("csv")\ .option("sep", ",")\ .option("escape", "\\\\") # Double-escape to pass a single backslash to Spark .option("inferSchema", "true")\ .option("header", "true")\ .load("test.csv")
Alternatively, if you prefer raw strings (which avoid Python's automatic escaping), you can write it as:
.option("escape", r"\\") # Raw string for double backslash, resolves to a single backslash for Spark
This ensures Spark correctly recognizes the backslash as the escape character, strips it, and keeps the comma as part of the name field.
2. Clean Up the Field After Reading
If the first solution doesn't work (e.g., due to an older Spark version), you can post-process the name field to remove the backslashes directly using regexp_replace:
from pyspark.sql.functions import regexp_replace # Read the CSV as you originally did test = spark.read.format("csv")\ .option("sep", ",")\ .option("inferSchema", "true")\ .option("header", "true")\ .load("test.csv") # Remove all backslashes from the name field test = test.withColumn("name", regexp_replace("name", "\\\\", "")) test.show()
This is a quick fix that works perfectly for your use case since your backslashes are only used to escape commas.
3. Explicitly Specify the Parser Library
Spark uses the Univocity parser by default for CSV files, but explicitly setting this can resolve subtle parsing quirks in some environments:
test = spark.read.format("csv")\ .option("sep", ",")\ .option("escape", "\\")\ .option("parserLib", "univocity")\ .option("inferSchema", "true")\ .option("header", "true")\ .load("test.csv")
Expected Result
After applying any of these solutions, running test.show() should give you the clean output you want:
+---+-------------+ | id| name| +---+-------------+ | 10|Ashraful, Islam| +---+-------------+
内容的提问来源于stack exchange,提问作者Ashraful Islam

