You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 2.7:如何正确提取网站文章数据?

Troubleshooting Your Article Scraping Script (ID Range 521-4458 to CSV)

Hey there! Let's work through the issues you're facing when scraping articles by iterating through IDs and saving results to a CSV. First, let's break down the most common pitfalls here, then walk through fixes and a robust example script.

Common Causes of Errors

Here are the top reasons your script might be failing:

  • Invalid/Non-existent IDs: Some IDs in the 521-4458 range might not correspond to actual articles (returning 404 errors).
  • Anti-scraping Measures: The site might block your requests if you don't mimic a browser or send requests too quickly.
  • CSV Writing Issues: Encoding problems, incorrect field handling, or not opening the file in the right mode.
  • Page Parsing Errors: Even if you said all articles have the same structure, small inconsistencies (like missing elements) can break your selectors.
  • Loop Logic Mistakes: A misconfigured while loop could cause infinite loops or stop before processing all IDs.

Step-by-Step Fixes & Example Script

Let's build a script that addresses all these issues. We'll use requests for fetching pages, BeautifulSoup for parsing, and Python's built-in csv module for saving data.

Example Script

import requests
from bs4 import BeautifulSoup
import csv
import time

# Configuration - update these to match your target site
START_ID = 521
END_ID = 4458
BASE_ARTICLE_URL = "https://your-site.com/articles/{id}"  # Replace with actual URL template
OUTPUT_CSV = "scraped_articles.csv"

# Mimic a browser request to avoid anti-scraping blocks
REQUEST_HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept-Language": "en-US,en;q=0.9"
}

def scrape_single_article(article_id):
    """Fetch and parse a single article by ID"""
    article_url = BASE_ARTICLE_URL.format(id=article_id)
    
    try:
        # Fetch the page with timeout and error handling
        response = requests.get(article_url, headers=REQUEST_HEADERS, timeout=15)
        response.raise_for_status()  # Trigger exception for HTTP errors (4xx/5xx)
        
        soup = BeautifulSoup(response.text, "html.parser")
        
        # Parse data - UPDATE THESE SELECTORS TO MATCH YOUR SITE'S HTML
        title = soup.find("h1", class_="article-title").get_text(strip=True) if soup.find("h1", class_="article-title") else "No Title Found"
        publish_date = soup.find("span", class_="publish-date").get_text(strip=True) if soup.find("span", class_="publish-date") else "No Date Found"
        # Extract all paragraph content for the body
        article_body = "\n".join([p.get_text(strip=True) for p in soup.find_all("div", class_="article-body")]) if soup.find_all("div", class_="article-body") else "No Content Found"
        
        return {
            "article_id": article_id,
            "title": title,
            "publish_date": publish_date,
            "content": article_body
        }
    
    except Exception as e:
        print(f"⚠️ Failed to scrape ID {article_id}: {str(e)}")
        return None

def main():
    # Initialize CSV with headers
    with open(OUTPUT_CSV, "w", newline="", encoding="utf-8-sig") as csv_file:
        writer = csv.DictWriter(csv_file, fieldnames=["article_id", "title", "publish_date", "content"])
        writer.writeheader()
    
    # Iterate through IDs using a while loop (or switch to a for loop if preferred)
    current_id = START_ID
    while current_id <= END_ID:
        article_data = scrape_single_article(current_id)
        
        if article_data:
            # Append valid data to CSV
            with open(OUTPUT_CSV, "a", newline="", encoding="utf-8-sig") as csv_file:
                writer = csv.DictWriter(csv_file, fieldnames=["article_id", "title", "publish_date", "content"])
                writer.writerow(article_data)
            print(f"✅ Saved article ID {current_id}")
        
        # Add a small delay to avoid overwhelming the server
        time.sleep(1.5)
        current_id += 1

if __name__ == "__main__":
    main()

Key Fixes in This Script

  • Error Handling: Uses try-except blocks to catch request/parsing errors and log them instead of crashing.
  • Anti-Scraping Mitigation: Includes browser-like headers and a delay between requests to avoid getting blocked.
  • Robust CSV Writing: Uses utf-8-sig encoding to avoid character display issues, and opens the file in append mode for valid entries.
  • Fallback Values: If a field (like title or date) is missing, it uses a placeholder instead of breaking the script.

Debugging Tips

  • Test with a Single ID: First run the script for just one valid ID to confirm your selectors work correctly.
  • Inspect Page HTML: Use your browser's dev tools to double-check that your selectors match the site's actual elements.
  • Check Response Content: Print response.text for failing IDs to see if the page returns a captcha, 404, or unexpected content.
  • Log Errors: For persistent failures, write error details to a log file instead of just printing them for easier debugging.

内容的提问来源于stack exchange,提问作者Bab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:38:36