如何使用Python清理CSV文件:将未以双引号结尾的行合并至前一行
I see the problem you're facing—your CSV has line breaks inside the remark field, which are splitting what should be a single row into multiple lines. Your current code just strips all newlines, which breaks the actual row separators too. Let's fix this properly.
Why Your Current Code Isn't Working
Your code reads the entire file, replaces every \n with nothing, and writes it back. This turns all your rows into one giant line, which is not what you want. We need to keep the valid row endings (where a line ends with a double quote) while merging the lines that are part of the same field.
The Solution: Build Rows Line by Line
We'll process each line, build complete rows by concatenating incomplete lines (those not ending with a "), and write them out once they're whole. We'll also replace the internal newlines with spaces to match your desired output.
Here's the working code:
# Open input and output files (using with statements ensures proper file closure) with open("input.csv", "r", encoding='utf-8') as infile, open("output.csv", "w", encoding='utf-8') as outfile: current_row = "" for line in infile: # Strip leading/trailing whitespace (including newlines) but keep internal spaces stripped_line = line.strip() # Skip empty lines to clean up extra whitespace if not stripped_line: continue # Append the stripped line to our current row, adding a space to replace the newline if current_row: current_row += " " + stripped_line else: current_row = stripped_line # Check if we've finished building a complete row (ends with a closing double quote) if current_row.endswith('"'): # Write the complete row with a proper newline outfile.write(current_row + "\n") # Reset for the next row current_row = "" # Handle any remaining content if the file ends with an incomplete row if current_row: outfile.write(current_row + "\n")
How This Works
- Line-by-line processing: Reads one line at a time, which is more memory-efficient for large CSV files.
- Incomplete row handling: If a line doesn't end with a
", it keeps adding tocurrent_rowuntil a full row is assembled. - Clean internal newlines: Replaces field-internal line breaks with spaces to match your desired output format.
- Empty line cleanup: Ignores blank lines in the input to avoid messy output.
Testing Against Your Example
Running this code on your sample input will produce exactly the output you want:
- Broken rows for John, Naima, and Lili are merged into single, complete rows.
- Internal line breaks are replaced with spaces (e.g.,
"hey men,\nhow do you did this"becomes"hey men, how do you did this"). - Valid, complete rows (like Kevin, Angel, etc.) remain unchanged.
内容的提问来源于stack exchange,提问作者kevin masliah

