Python 2.7:如何正确提取网站文章数据?
Troubleshooting Your Article Scraping Script (ID Range 521-4458 to CSV)
Hey there! Let's work through the issues you're facing when scraping articles by iterating through IDs and saving results to a CSV. First, let's break down the most common pitfalls here, then walk through fixes and a robust example script.
Common Causes of Errors
Here are the top reasons your script might be failing:
- Invalid/Non-existent IDs: Some IDs in the 521-4458 range might not correspond to actual articles (returning 404 errors).
- Anti-scraping Measures: The site might block your requests if you don't mimic a browser or send requests too quickly.
- CSV Writing Issues: Encoding problems, incorrect field handling, or not opening the file in the right mode.
- Page Parsing Errors: Even if you said all articles have the same structure, small inconsistencies (like missing elements) can break your selectors.
- Loop Logic Mistakes: A misconfigured
whileloop could cause infinite loops or stop before processing all IDs.
Step-by-Step Fixes & Example Script
Let's build a script that addresses all these issues. We'll use requests for fetching pages, BeautifulSoup for parsing, and Python's built-in csv module for saving data.
Example Script
import requests from bs4 import BeautifulSoup import csv import time # Configuration - update these to match your target site START_ID = 521 END_ID = 4458 BASE_ARTICLE_URL = "https://your-site.com/articles/{id}" # Replace with actual URL template OUTPUT_CSV = "scraped_articles.csv" # Mimic a browser request to avoid anti-scraping blocks REQUEST_HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9" } def scrape_single_article(article_id): """Fetch and parse a single article by ID""" article_url = BASE_ARTICLE_URL.format(id=article_id) try: # Fetch the page with timeout and error handling response = requests.get(article_url, headers=REQUEST_HEADERS, timeout=15) response.raise_for_status() # Trigger exception for HTTP errors (4xx/5xx) soup = BeautifulSoup(response.text, "html.parser") # Parse data - UPDATE THESE SELECTORS TO MATCH YOUR SITE'S HTML title = soup.find("h1", class_="article-title").get_text(strip=True) if soup.find("h1", class_="article-title") else "No Title Found" publish_date = soup.find("span", class_="publish-date").get_text(strip=True) if soup.find("span", class_="publish-date") else "No Date Found" # Extract all paragraph content for the body article_body = "\n".join([p.get_text(strip=True) for p in soup.find_all("div", class_="article-body")]) if soup.find_all("div", class_="article-body") else "No Content Found" return { "article_id": article_id, "title": title, "publish_date": publish_date, "content": article_body } except Exception as e: print(f"⚠️ Failed to scrape ID {article_id}: {str(e)}") return None def main(): # Initialize CSV with headers with open(OUTPUT_CSV, "w", newline="", encoding="utf-8-sig") as csv_file: writer = csv.DictWriter(csv_file, fieldnames=["article_id", "title", "publish_date", "content"]) writer.writeheader() # Iterate through IDs using a while loop (or switch to a for loop if preferred) current_id = START_ID while current_id <= END_ID: article_data = scrape_single_article(current_id) if article_data: # Append valid data to CSV with open(OUTPUT_CSV, "a", newline="", encoding="utf-8-sig") as csv_file: writer = csv.DictWriter(csv_file, fieldnames=["article_id", "title", "publish_date", "content"]) writer.writerow(article_data) print(f"✅ Saved article ID {current_id}") # Add a small delay to avoid overwhelming the server time.sleep(1.5) current_id += 1 if __name__ == "__main__": main()
Key Fixes in This Script
- Error Handling: Uses
try-exceptblocks to catch request/parsing errors and log them instead of crashing. - Anti-Scraping Mitigation: Includes browser-like headers and a delay between requests to avoid getting blocked.
- Robust CSV Writing: Uses
utf-8-sigencoding to avoid character display issues, and opens the file in append mode for valid entries. - Fallback Values: If a field (like title or date) is missing, it uses a placeholder instead of breaking the script.
Debugging Tips
- Test with a Single ID: First run the script for just one valid ID to confirm your selectors work correctly.
- Inspect Page HTML: Use your browser's dev tools to double-check that your selectors match the site's actual elements.
- Check Response Content: Print
response.textfor failing IDs to see if the page returns a captcha, 404, or unexpected content. - Log Errors: For persistent failures, write error details to a log file instead of just printing them for easier debugging.
内容的提问来源于stack exchange,提问作者Bab
相关产品推荐
相关产品推荐

