如何将亚马逊评论爬虫Python脚本结果写入CSV文件?
How to Write Extracted Amazon Reviews to a CSV File
Hey there! Let's get those scraped Amazon review details into a CSV file smoothly. Below are two practical methods you can use, depending on your needs:
Method 1: Use Python's Built-in csv Module (No Extra Dependencies)
This is a lightweight solution that doesn’t require installing additional packages. We’ll open a CSV file, write a header row first, then append each review’s data as we scrape it.
Here’s how to integrate it with your existing code:
import csv import logging # ... your existing code to fetch reviews_list ... # Define CSV file path and column headers csv_filename = "amazon_reviews.csv" header = ["Rating", "Review Title", "Review Date", "Review Body"] # Open the CSV file in write mode with open(csv_filename, "w", newline="", encoding="utf-8") as csv_file: # Create a CSV writer object writer = csv.writer(csv_file) # Write the header row first writer.writerow(header) for review in reviews_list: try: # Extract your data as before, with cleanup rating = review.find(attrs={'data-hook': 'review-star-rating'}).attrs['class'][2].split('-')[-1] body = review.find(attrs={'data-hook': 'review-body'}).text.strip() date = review.find(attrs={'data-hook': 'review-date'}).text.strip() title = review.find(attrs={'data-hook': 'review-title'}).text.strip() # Write the current review's data to CSV writer.writerow([rating, title, date, body]) logging.info('***********************************************') logging.info(f"Title: {title}") # ... rest of your logging code ... except AttributeError: # Handle cases where a review might be missing fields (e.g., deleted reviews) logging.warning("Skipping a review with missing fields") continue
Key Notes for This Method:
encoding="utf-8"ensures special characters (like accents or emojis) don’t get mangled in the CSV.newline=""prevents extra blank lines from appearing in the file on Windows systems.- The
try-exceptblock stops your scraper from crashing if a review is missing a required field. .strip()removes unnecessary whitespace/newlines from text fields for cleaner output.
Method 2: Use pandas (Great for Post-Scraping Analysis)
If you plan to analyze the review data later, pandas simplifies the process. First install pandas if you haven’t already (pip install pandas), then collect all reviews into a list of dictionaries, then export to CSV.
Here’s the code:
import pandas as pd import logging # ... your existing code to fetch reviews_list ... # Empty list to store all review data reviews_data = [] for review in reviews_list: try: rating = review.find(attrs={'data-hook': 'review-star-rating'}).attrs['class'][2].split('-')[-1] body = review.find(attrs={'data-hook': 'review-body'}).text.strip() date = review.find(attrs={'data-hook': 'review-date'}).text.strip() title = review.find(attrs={'data-hook': 'review-title'}).text.strip() # Append the review as a dictionary to the list reviews_data.append({ "Rating": rating, "Review Title": title, "Review Date": date, "Review Body": body }) logging.info('***********************************************') logging.info(f"Title: {title}") # ... rest of your logging code ... except AttributeError: logging.warning("Skipping a review with missing fields") continue # Convert the list to a DataFrame and export to CSV df = pd.DataFrame(reviews_data) df.to_csv("amazon_reviews.csv", index=False, encoding="utf-8")
Why Use This Method?
- Pandas automatically handles edge cases like missing values and formatting inconsistencies.
- You can immediately analyze the data after exporting (e.g.,
df['Rating'].value_counts()to see the distribution of star ratings).
内容的提问来源于stack exchange,提问作者julienrath
相关产品推荐
相关产品推荐

