Scrapy技术问题:为爬取数据添加唯一ID以避免重复
Hey Rafael, let's break down your two questions with practical, actionable solutions—these are super common issues when scraping dynamic product catalogs, so I’ve got you covered:
Since there’s no built-in unique ID for products, we’ll rely on unique attribute combinations and database-level checks to prevent duplicates:
Use a Composite Unique Key for Products
Combine fields that are unique to each product (likeproduct_name,price, anddetail_page_url)—this combination is almost guaranteed to be distinct even as the site’s product list changes.- First, add a
UNIQUEconstraint to your database table (example for MySQL):ALTER TABLE products ADD UNIQUE KEY unique_product_identifier (name, price, detail_url); - Then, in your Scrapy
ItemPipeline, check for existing products before inserting, and skip duplicates:from scrapy.exceptions import DropItem import psycopg2 # Or your database connector of choice class DuplicateCheckPipeline: def __init__(self): self.conn = psycopg2.connect( dbname="your_db", user="your_user", password="your_pass", host="localhost" ) self.cur = self.conn.cursor() def process_item(self, item, spider): # Check if product already exists using the composite key self.cur.execute( "SELECT 1 FROM products WHERE name = %s AND price = %s AND detail_url = %s", (item["name"], item["price"], item["detail_url"]) ) if self.cur.fetchone(): raise DropItem(f"Skipping duplicate product: {item['name']}") # Insert the new product if no duplicate found self.cur.execute( "INSERT INTO products (name, price, detail_url) VALUES (%s, %s, %s)", (item["name"], item["price"], item["detail_url"]) ) self.conn.commit() return item
- First, add a
Customize Scrapy’s DupeFilter
Skip duplicate items before they even reach the pipeline by overriding Scrapy’s built-in duplicate filter to track your composite product keys:from scrapy.dupefilters import RFPDupeFilter class ProductDupeFilter(RFPDupeFilter): def __init__(self, path=None, debug=False): super().__init__(path, debug) self.seen_products = set() def request_seen(self, request): # Extract the unique product key from request metadata product_key = ( request.meta["product_name"], request.meta["product_price"], request.meta["detail_url"] ) if product_key in self.seen_products: return True self.seen_products.add(product_key) return FalseEnable this filter in your
settings.py:DUPEFILTER_CLASS = "your_project_name.filters.ProductDupeFilter"
The best approach depends on whether you need sequential numbers for a single crawl or a global, persistent sequence:
Option 1: Let Your Database Handle Auto-Increment (Most Reliable)
For a global continuous sequence that persists across crawls, use your database’s auto-increment feature. This eliminates the need to track the last number in Scrapy.
Example table setup for MySQL:CREATE TABLE products ( id INT AUTO_INCREMENT PRIMARY KEY, # Auto-generates continuous IDs name VARCHAR(255), price DECIMAL(10,2), detail_url VARCHAR(255), UNIQUE KEY unique_product_identifier (name, price, detail_url) );When inserting items, you don’t need to include the
idfield—your database will automatically assign the next sequential number, even as products are added or removed between crawls.Option 2: Sequential Numbers for a Single Crawl
If you only need sequential numbers for the current crawl’s output (e.g., a CSV), use a counter in your spider:class ProductSpider(scrapy.Spider): name = "product_crawler" start_urls = ["https://your-target-site.com/products"] crawl_counter = 1 def parse(self, response): for product in response.css("div.product-card"): item = ProductItem() item["serial_number"] = self.crawl_counter item["name"] = product.css("h2.product-title::text").get().strip() item["price"] = product.css("span.price-tag::text").get().strip() item["detail_url"] = product.css("a.product-link::attr(href)").get() self.crawl_counter += 1 yield itemNote: This will reset to 1 every time you run the spider—use the database auto-increment method if you need a persistent sequence.
内容的提问来源于stack exchange,提问作者Rafael Souza

