You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark读取CSV:如何基于ASCII字符á合并行并正确解析列

Fixing PySpark CSV Parsing & Text Cleanup for Your Dataset

Let's break down the issues with your current code and fix them step by step:

1. Correct CSV Parsing Configuration

Your main problem right now is misconfigured delimiter and quote options. The CSV uses comma as the field separator, and double quotes to wrap the sentence field (which contains spaces and special characters). Your current code sets delimiter to " which completely breaks column parsing.

Here's the corrected read configuration:

from pyspark.sql.types import StructType, StructField, StringType
from pyspark.sql.functions import regexp_replace

# Keep your schema definition as-is
schema = StructType([
    StructField("Id", StringType()),
    StructField("Sentence", StringType()),
])

# Fix the CSV reading parameters
df = sqlContext.read.format("csv") \
    .option("header", "false") \
    .option("inferSchema", "false") \
    .option("delimiter", ",")  # Correct field separator for your CSV
    .option("quote", "\"")     # Tells Spark to treat " as the field-wrapping character
    .option("multiLine", "true")  # Safely handles any line breaks inside quoted sentence fields
    .schema(schema) \
    .load("mycsv.csv")

2. Clean Up the Sentence Field (Remove á & Merge Content)

Now that we've parsed the columns correctly, we need to remove the á characters and merge the split sentence parts. Use regexp_replace to normalize the text:

# Clean the Sentence column: remove á and fix whitespace
clean_df = df.withColumn(
    "Sentence",
    regexp_replace(regexp_replace(df["Sentence"], r"\s*á\s*", " "), r"\s+", " ").trim()
)

Let's break down the regex logic:

  • r"\s*á\s*" matches á with any surrounding spaces, replacing it with a single space
  • r"\s+" replaces multiple consecutive spaces with one single space
  • .trim() removes leading/trailing whitespace from the final sentence

3. Verify the Result

Run clean_df.show(truncate=False) to confirm the output. You'll get properly parsed columns with cleaned, merged sentences:

+---+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------+
|Id |Sentence                                                                                                                                                               |
+---+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------+
|id1|When I think about the short time that we live and relate it to the periods of my life when I think that I did not use this short time.                                |
|id2|[ On days when I feel close to my partner and other friends. When I feel at peace with myself and also experience a close contact with people whom I regard greatly.]|
+---+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------+

Key Takeaways

  • Align your CSV options with the actual file format: delimiter is the field separator (comma here), quote is the character wrapping multi-word fields
  • multiLine=true is a safe guard for real-world data where quoted fields might contain line breaks
  • Spark's regex functions let you clean text efficiently at scale without manual string manipulation

内容的提问来源于stack exchange,提问作者abhjt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:55:45