You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XML文件空行删除方法及CSV读写报错解决方案咨询

Got it, let's tackle your issues one by one—first fixing that CSV error, then handling in-place edits, and even addressing why those empty rows are showing up in the first place.

1. Fixing the _csv.Error in your row-deletion function

The error you're seeing happens because you opened your files in binary mode (rb/wb), but the csv module requires text mode to work properly. Binary mode returns bytes instead of strings, which breaks the csv reader. Here's the corrected version of your function, with some best practices added:

import csv
import os

def delete_empty_rows_in_csv():
    # Use text mode with explicit encoding, and let `with` auto-close files
    with open('../data/employees_csv.csv', 'r', encoding='utf-8') as input_file:
        # `newline=''` is recommended by the csv module to avoid extra line breaks on Windows
        with open('../data/employees.csv', 'w', encoding='utf-8', newline='') as output_file:
            writer = csv.writer(output_file)
            for row in csv.reader(input_file):
                # Filter out rows where all cells are empty/whitespace
                if any(cell.strip() for cell in row):
                    writer.writerow(row)
    os.remove('../data/employees_csv.csv')

Key improvements here:

  • Uses with statements to automatically close files (safer than manual close() calls)
  • Specifies encoding='utf-8' to avoid character encoding issues
  • newline='' prevents unwanted extra blank lines in the output CSV
  • Replaces simple if row with any(cell.strip()...) to filter rows that only contain whitespace

2. Removing empty rows directly in the original file

If you don't want to create a temporary file, you can read all valid rows into memory, then overwrite the original file. This works well for most CSV files (avoid it if your file is extremely large):

import csv

def remove_empty_rows_inplace(file_path):
    # Read all non-empty rows from the file
    with open(file_path, 'r', encoding='utf-8') as f:
        valid_rows = [row for row in csv.reader(f) if any(cell.strip() for cell in row)]
    
    # Overwrite the original file with only valid rows
    with open(file_path, 'w', encoding='utf-8', newline='') as f:
        writer = csv.writer(f)
        writer.writerows(valid_rows)

# Use it like this:
remove_empty_rows_inplace('../data/employees_csv.csv')

3. Preventing empty rows from being generated in the first place

Instead of cleaning up empty rows after the fact, let's stop them from being written at all. The empty rows are likely coming from:

  • Empty <tr> elements in your HTML table
  • <td> elements where .text returns empty strings for every cell in the row

Here's how to modify your scraping code to filter these rows upfront:

for row in tbody.find_elements_by_xpath('./tr'):
    itemsEmployee = row.find_elements_by_xpath('./td')
    # Extract text from each target cell, and strip whitespace
    row_data = [
        itemsEmployee[1].text.strip(),
        itemsEmployee[5].text.strip(),
        itemsEmployee[2].text.strip(),
        itemsEmployee[3].text.strip(),
        itemsEmployee[4].text.strip(),
        itemsEmployee[6].text.strip(),
        itemsEmployee[7].text.strip(),
        itemsEmployee[8].text.strip()
    ]
    # Only write the row if at least one cell has non-empty content
    if any(row_data):
        fileWriter.writerow(row_data)

This way, you won't have to run a cleanup function at all—your CSV will only contain valid rows from the start.

内容的提问来源于stack exchange,提问作者WebSTAR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 15:27:40