XML文件空行删除方法及CSV读写报错解决方案咨询
Got it, let's tackle your issues one by one—first fixing that CSV error, then handling in-place edits, and even addressing why those empty rows are showing up in the first place.
1. Fixing the _csv.Error in your row-deletion function
The error you're seeing happens because you opened your files in binary mode (rb/wb), but the csv module requires text mode to work properly. Binary mode returns bytes instead of strings, which breaks the csv reader. Here's the corrected version of your function, with some best practices added:
import csv import os def delete_empty_rows_in_csv(): # Use text mode with explicit encoding, and let `with` auto-close files with open('../data/employees_csv.csv', 'r', encoding='utf-8') as input_file: # `newline=''` is recommended by the csv module to avoid extra line breaks on Windows with open('../data/employees.csv', 'w', encoding='utf-8', newline='') as output_file: writer = csv.writer(output_file) for row in csv.reader(input_file): # Filter out rows where all cells are empty/whitespace if any(cell.strip() for cell in row): writer.writerow(row) os.remove('../data/employees_csv.csv')
Key improvements here:
- Uses
withstatements to automatically close files (safer than manualclose()calls) - Specifies
encoding='utf-8'to avoid character encoding issues newline=''prevents unwanted extra blank lines in the output CSV- Replaces simple
if rowwithany(cell.strip()...)to filter rows that only contain whitespace
2. Removing empty rows directly in the original file
If you don't want to create a temporary file, you can read all valid rows into memory, then overwrite the original file. This works well for most CSV files (avoid it if your file is extremely large):
import csv def remove_empty_rows_inplace(file_path): # Read all non-empty rows from the file with open(file_path, 'r', encoding='utf-8') as f: valid_rows = [row for row in csv.reader(f) if any(cell.strip() for cell in row)] # Overwrite the original file with only valid rows with open(file_path, 'w', encoding='utf-8', newline='') as f: writer = csv.writer(f) writer.writerows(valid_rows) # Use it like this: remove_empty_rows_inplace('../data/employees_csv.csv')
3. Preventing empty rows from being generated in the first place
Instead of cleaning up empty rows after the fact, let's stop them from being written at all. The empty rows are likely coming from:
- Empty
<tr>elements in your HTML table <td>elements where.textreturns empty strings for every cell in the row
Here's how to modify your scraping code to filter these rows upfront:
for row in tbody.find_elements_by_xpath('./tr'): itemsEmployee = row.find_elements_by_xpath('./td') # Extract text from each target cell, and strip whitespace row_data = [ itemsEmployee[1].text.strip(), itemsEmployee[5].text.strip(), itemsEmployee[2].text.strip(), itemsEmployee[3].text.strip(), itemsEmployee[4].text.strip(), itemsEmployee[6].text.strip(), itemsEmployee[7].text.strip(), itemsEmployee[8].text.strip() ] # Only write the row if at least one cell has non-empty content if any(row_data): fileWriter.writerow(row_data)
This way, you won't have to run a cleanup function at all—your CSV will only contain valid rows from the start.
内容的提问来源于stack exchange,提问作者WebSTAR

