You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬虫脚本循环监控新增产品失效:重复返回旧数据问题求助

Hey Phil, let's figure out why your script keeps pulling the same data instead of detecting new products! The core issues here are lack of tracked scraped items and potential request caching. Here's a step-by-step fix:

Problem Analysis & Solutions

Your loop repeats old data for two key reasons:

  1. You aren't keeping track of products you've already scraped, so every loop re-fetches the entire list
  2. Servers might return cached pages instead of the latest content

1. Track Scraped Items with a Set

We'll use a set to store unique product links (since links are unique identifiers). This way, we only process items we haven't seen before.

2. Bypass Request Caching

Add cache-control headers to your request to force the server to send the latest page, not a cached version.

Fixed Full Script

import requests
import csv
from bs4 import BeautifulSoup
import time

# Initialize a set to store already scraped product links (prevents duplicates)
scraped_links = set()

# Open CSV in append mode so we don't overwrite existing data
with open('products.csv', 'a', newline='', encoding='utf-8') as csvfile:
    fieldnames = ['link', 'title', 'size']
    writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
    
    # Write header only if the file is empty
    csvfile.seek(0, 2)
    if csvfile.tell() == 0:
        writer.writeheader()

    # Monitoring loop
    while True:
        headers = {
            "user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_12_6) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/65.0.3325.181 Safari/537.36",
            # Disable caching to get fresh content every time
            "Cache-Control": "no-cache",
            "Pragma": "no-cache",
            "Expires": "0"
        }
        
        # Replace this with your actual target URL
        target_url = "YOUR_PRODUCT_PAGE_URL"
        try:
            response = requests.get(target_url, headers=headers)
            response.raise_for_status()  # Throw error if request fails (e.g., 404, 500)
        except requests.exceptions.RequestException as e:
            print(f"Request failed: {e}. Retrying in 10 minutes...")
            time.sleep(600)
            continue

        soup = BeautifulSoup(response.text, 'html.parser')
        
        # Update this selector to match your target site's product container (use dev tools to check)
        product_containers = soup.find_all('div', class_='product-item')
        
        new_items_found = 0
        for container in product_containers:
            # Extract product link (adjust selector to match your site)
            raw_link = container.find('a')['href']
            # Convert relative links to full URLs if needed
            if not raw_link.startswith('http'):
                base_url = f"{target_url.split('//')[0]}//{target_url.split('//')[1].split('/')[0]}"
                product_link = f"{base_url}{raw_link}"
            else:
                product_link = raw_link

            # Only process if we haven't scraped this link before
            if product_link not in scraped_links:
                scraped_links.add(product_link)
                # Extract title (adjust selector)
                product_title = container.find('h3', class_='product-title').get_text(strip=True)
                # Extract size (adjust selector; default to 'N/A' if size doesn't exist)
                product_size = container.find('span', class_='product-size').get_text(strip=True) if container.find('span', class_='product-size') else 'N/A'
                
                # Write new item to CSV
                writer.writerow({
                    'link': product_link,
                    'title': product_title,
                    'size': product_size
                })
                new_items_found += 1
                print(f"New product added: {product_title}")
        
        print(f"Loop complete. Found {new_items_found} new products. Checking again in 10 minutes...")
        # Adjust sleep time to avoid overwhelming the server (don't set this too low!)
        time.sleep(600)

Extra Tips

  • Adjust Selectors: Use your browser's dev tools to update the HTML selectors (like div.product-item) to match the actual structure of your target site.
  • Anti-Scraping Protection: Avoid getting blocked by setting a reasonable time.sleep interval. You could also add proxy support if the site blocks repeated requests from the same IP.
  • Persistent Tracking: If you want to keep your scraped links between script restarts, save the scraped_links set to a JSON file and load it when the script starts.

内容的提问来源于stack exchange,提问作者Phil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:46:19