PySpark读取CSV:如何基于ASCII字符á合并行并正确解析列
Let's break down the issues with your current code and fix them step by step:
1. Correct CSV Parsing Configuration
Your main problem right now is misconfigured delimiter and quote options. The CSV uses comma as the field separator, and double quotes to wrap the sentence field (which contains spaces and special characters). Your current code sets delimiter to " which completely breaks column parsing.
Here's the corrected read configuration:
from pyspark.sql.types import StructType, StructField, StringType from pyspark.sql.functions import regexp_replace # Keep your schema definition as-is schema = StructType([ StructField("Id", StringType()), StructField("Sentence", StringType()), ]) # Fix the CSV reading parameters df = sqlContext.read.format("csv") \ .option("header", "false") \ .option("inferSchema", "false") \ .option("delimiter", ",") # Correct field separator for your CSV .option("quote", "\"") # Tells Spark to treat " as the field-wrapping character .option("multiLine", "true") # Safely handles any line breaks inside quoted sentence fields .schema(schema) \ .load("mycsv.csv")
2. Clean Up the Sentence Field (Remove á & Merge Content)
Now that we've parsed the columns correctly, we need to remove the á characters and merge the split sentence parts. Use regexp_replace to normalize the text:
# Clean the Sentence column: remove á and fix whitespace clean_df = df.withColumn( "Sentence", regexp_replace(regexp_replace(df["Sentence"], r"\s*á\s*", " "), r"\s+", " ").trim() )
Let's break down the regex logic:
r"\s*á\s*"matchesáwith any surrounding spaces, replacing it with a single spacer"\s+"replaces multiple consecutive spaces with one single space.trim()removes leading/trailing whitespace from the final sentence
3. Verify the Result
Run clean_df.show(truncate=False) to confirm the output. You'll get properly parsed columns with cleaned, merged sentences:
+---+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------+ |Id |Sentence | +---+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------+ |id1|When I think about the short time that we live and relate it to the periods of my life when I think that I did not use this short time. | |id2|[ On days when I feel close to my partner and other friends. When I feel at peace with myself and also experience a close contact with people whom I regard greatly.]| +---+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------+
Key Takeaways
- Align your CSV options with the actual file format: delimiter is the field separator (comma here), quote is the character wrapping multi-word fields
multiLine=trueis a safe guard for real-world data where quoted fields might contain line breaks- Spark's regex functions let you clean text efficiently at scale without manual string manipulation
内容的提问来源于stack exchange,提问作者abhjt

