爬虫脚本循环监控新增产品失效:重复返回旧数据问题求助
Hey Phil, let's figure out why your script keeps pulling the same data instead of detecting new products! The core issues here are lack of tracked scraped items and potential request caching. Here's a step-by-step fix:
Problem Analysis & Solutions
Your loop repeats old data for two key reasons:
- You aren't keeping track of products you've already scraped, so every loop re-fetches the entire list
- Servers might return cached pages instead of the latest content
1. Track Scraped Items with a Set
We'll use a set to store unique product links (since links are unique identifiers). This way, we only process items we haven't seen before.
2. Bypass Request Caching
Add cache-control headers to your request to force the server to send the latest page, not a cached version.
Fixed Full Script
import requests import csv from bs4 import BeautifulSoup import time # Initialize a set to store already scraped product links (prevents duplicates) scraped_links = set() # Open CSV in append mode so we don't overwrite existing data with open('products.csv', 'a', newline='', encoding='utf-8') as csvfile: fieldnames = ['link', 'title', 'size'] writer = csv.DictWriter(csvfile, fieldnames=fieldnames) # Write header only if the file is empty csvfile.seek(0, 2) if csvfile.tell() == 0: writer.writeheader() # Monitoring loop while True: headers = { "user-agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_12_6) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/65.0.3325.181 Safari/537.36", # Disable caching to get fresh content every time "Cache-Control": "no-cache", "Pragma": "no-cache", "Expires": "0" } # Replace this with your actual target URL target_url = "YOUR_PRODUCT_PAGE_URL" try: response = requests.get(target_url, headers=headers) response.raise_for_status() # Throw error if request fails (e.g., 404, 500) except requests.exceptions.RequestException as e: print(f"Request failed: {e}. Retrying in 10 minutes...") time.sleep(600) continue soup = BeautifulSoup(response.text, 'html.parser') # Update this selector to match your target site's product container (use dev tools to check) product_containers = soup.find_all('div', class_='product-item') new_items_found = 0 for container in product_containers: # Extract product link (adjust selector to match your site) raw_link = container.find('a')['href'] # Convert relative links to full URLs if needed if not raw_link.startswith('http'): base_url = f"{target_url.split('//')[0]}//{target_url.split('//')[1].split('/')[0]}" product_link = f"{base_url}{raw_link}" else: product_link = raw_link # Only process if we haven't scraped this link before if product_link not in scraped_links: scraped_links.add(product_link) # Extract title (adjust selector) product_title = container.find('h3', class_='product-title').get_text(strip=True) # Extract size (adjust selector; default to 'N/A' if size doesn't exist) product_size = container.find('span', class_='product-size').get_text(strip=True) if container.find('span', class_='product-size') else 'N/A' # Write new item to CSV writer.writerow({ 'link': product_link, 'title': product_title, 'size': product_size }) new_items_found += 1 print(f"New product added: {product_title}") print(f"Loop complete. Found {new_items_found} new products. Checking again in 10 minutes...") # Adjust sleep time to avoid overwhelming the server (don't set this too low!) time.sleep(600)
Extra Tips
- Adjust Selectors: Use your browser's dev tools to update the HTML selectors (like
div.product-item) to match the actual structure of your target site. - Anti-Scraping Protection: Avoid getting blocked by setting a reasonable
time.sleepinterval. You could also add proxy support if the site blocks repeated requests from the same IP. - Persistent Tracking: If you want to keep your scraped links between script restarts, save the
scraped_linksset to a JSON file and load it when the script starts.
内容的提问来源于stack exchange,提问作者Phil
相关产品推荐
相关产品推荐

