PySpark 3.2.0中lineSep配置未生效:含列内换行符的CSV文件行拆分异常求助
lineSep set to '\r' when reading CSV I understand your frustration—Spark's CSV reader can have unexpected behavior with non-standard line separators, especially in older versions like 3.2.0. Let's break down the possible reasons and fixes for your scenario:
Key Observations & Potential Causes
Incorrect quote/escape character configuration:
Your currentquoteandescapeparameters are set to", which is the HTML entity for a double quote ("). Spark expects the actual character here, not the HTML entity string. This means Spark isn't recognizing quoted fields correctly—so even if you havemultiLineenabled, it can't properly handle\ninside quoted columns, leading to incorrect line splitting. This is likely a major root cause of your issue.Conflict between
multiLineandlineSep:
WhenmultiLineis enabled (as it is for English-language columns), Spark prioritizes the quote character to determine record boundaries instead of strictly respectinglineSep. If your quoted fields with\naren't being recognized (due to the incorrect quote configuration), Spark will fall back to using default line breaks (like\n) to split rows, ignoring yourlineSepsetting.Spark 3.2.0 bug with single-character line separators:
Spark 3.2.0 has known limitations with handling line separators that are just\r(carriage return) without a line feed. The CSV reader in this version may not correctly parse files using only\ras the line ending, even when explicitly configured. This issue has been addressed in newer Spark releases (3.3.x and above).Mixed line endings in the CSV file:
It's worth verifying that your CSV file consistently uses\ras the line separator. If there are any\ncharacters outside of quoted fields, Spark might interpret those as line breaks regardless of yourlineSepsetting.
Recommended Solutions
1. Fix the quote/escape character configuration
First, correct your quote and escape parameters to use the actual double quote character instead of the HTML entity. This will allow Spark to properly recognize quoted fields and handle \n within them when multiLine is enabled. Your adjusted configuration should look like this:
csv_params = { 'inferSchema': False, # Read columns as StringType 'header': True, 'sep': ',', 'quote': '"', 'escape': '"', 'lineSep': '\r', 'multiLine': column_language == 'en', 'unescapedQuoteHandling': 'RAISE_ERROR', 'mode': 'FAILFAST', 'enforceSchema': False, }
2. Adjust multiLine behavior (if needed)
If your English-language columns don't actually require multi-line parsing (i.e., no quoted fields with line breaks), consider setting multiLine to False for all cases. This will force Spark to respect the lineSep configuration strictly. If you do need multi-line support, ensure that all fields containing \n are properly enclosed in double quotes, and that any embedded quotes within those fields are escaped with another double quote (matching your escape setting).
3. Upgrade to a newer Spark version
Since this line separator issue is specific to older versions like 3.2.0, upgrading to Spark 3.3.0 or later should resolve the problem. Newer releases have improved handling of non-standard line separators and better compatibility with multiLine configurations.
4. Verify file integrity
Use a tool like hexdump (on Unix) or a text editor that shows invisible characters (like Notepad++ with "Show All Characters" enabled) to confirm that your CSV file uses \r exclusively as the line separator. Mixed line endings can cause Spark to fall back to default parsing behavior.
5. Workaround: Preprocess the file to standardize line endings
If upgrading Spark or fixing the configuration isn't sufficient, you can preprocess your CSV file to replace \r with \r\n (Windows-style) or \n (Unix-style) line endings. For example, using sed on Unix:
sed 's/\r/\r\n/g' input.csv > output.csv
Or in Python, you can read and rewrite the file before passing it to Spark:
with open('input.csv', 'r') as f_in, open('output.csv', 'w') as f_out: content = f_in.read() f_out.write(content.replace('\r', '\r\n'))
Let me know if you need further clarification on any of these steps!
内容的提问来源于stack exchange,提问作者Tjorriemorrie

