网页爬虫项目中Python写入CSV文件重复覆盖数据的问题求助
Hey there, let's fix that frustrating CSV overwrite problem you're dealing with in your web scraping project. I’ve seen this exact issue dozens of times—here’s what’s going wrong and how to fix it quickly.
Why You’re Seeing Overwrites
The core problem is almost certainly how you’re opening your CSV file during each loop. If you’re using open('your_file.csv', 'w') inside the loop, the 'w' (write) mode clears the entire file every time you open it—so by the end of your loop, only the last scraped table data gets saved.
Two Reliable Solutions
I’ll share two approaches, depending on how much data you’re scraping:
1. Collect All Data First, Write Once (Recommended)
This is the most efficient method for most cases. Instead of writing to the CSV every loop, store all your scraped table dictionaries in a list, then write everything to the file in one go after your loop finishes.
Here’s the code example tailored to your data structure:
import csv # Initialize an empty list to store all scraped data all_scraped_data = [] # Your existing HTML parsing loop for page in your_html_pages: # Your code to extract url, desc, price, etc. goes here table = { "UR": url, "DC": desc, "PR": price, "PU": picture, "SN": seller_name, "SU": seller_url } # Add the current table to our master list all_scraped_data.append(table) # Write all data to CSV after the loop ends with open('scraped_results.csv', 'w', newline='', encoding='utf-8') as csv_file: # Define the column headers (matches your dictionary keys) field_names = ["UR", "DC", "PR", "PU", "SN", "SU"] writer = csv.DictWriter(csv_file, fieldnames=field_names) writer.writeheader() # Write the column headers once writer.writerows(all_scraped_data) # Write all stored rows at once
2. Append Rows One-by-One (For Large Datasets)
If you’re scraping thousands of pages and don’t want to store all data in memory, use the 'a' (append) mode to add each row to the CSV without overwriting. Just make sure you only write the header once (not every loop).
Here’s how to do that:
import csv import os # Check if the CSV file already exists (to avoid duplicate headers) csv_file_path = 'scraped_results.csv' file_exists = os.path.isfile(csv_file_path) # Your HTML parsing loop for page in your_html_pages: table = { "UR": url, "DC": desc, "PR": price, "PU": picture, "SN": seller_name, "SU": seller_url } # Open the file in append mode with open(csv_file_path, 'a', newline='', encoding='utf-8') as csv_file: field_names = ["UR", "DC", "PR", "PU", "SN", "SU"] writer = csv.DictWriter(csv_file, fieldnames=field_names) # Write header only if the file didn't exist before if not file_exists: writer.writeheader() file_exists = True # Update flag so we don't write headers again # Append the current row to the CSV writer.writerow(table)
Key Notes to Remember
- File Modes: Use
'w'only when you want to create a new file (or overwrite an existing one). Use'a'to add to an existing file without erasing content. newline='': This prevents extra blank lines from appearing in your CSV on Windows systems.encoding='utf-8': Ensures special characters (like accents or non-English text) are saved correctly.DictWriter: Perfect for your use case since your data is already in a dictionary—no need to manually map values to columns!
内容的提问来源于stack exchange,提问作者Hohenfels

