Scrapy跨ScrapingHub定时任务的URL去重机制及配置咨询
Great question! Let me break down how this works based on my hands-on experience with Scrapy and ScrapingHub.
Default Deduplication: No Cross-Job Persistence
Out of the box, Scrapy uses the RFPDupeFilter which stores URL fingerprints only in memory during a single crawler run. When you set up daily scheduled jobs on ScrapingHub, each run spins up a completely fresh, isolated environment. That means the in-memory deduplication cache gets wiped clean every time a new job starts.
So to put it plainly: the default deduplication does not persist across scheduled tasks. Your crawler won't remember URLs it scraped in yesterday's run when it kicks off today's job.
How to Enable Cross-Scheduled-Job Deduplication
To make deduplication stick across runs, you need to replace the in-memory storage with a persistent system that's accessible across all your ScrapingHub jobs. The easiest built-in solution is ScrapingHub's Key Value Store (KVS)—it's tightly integrated with your project and perfect for quick fingerprint lookups.
Here's a step-by-step way to set this up:
1. Build a Custom DupeFilter
Create a custom filter that inherits from Scrapy's base RFPDupeFilter, but uses ScrapingHub KVS to store and check fingerprints. Example code:
from scrapy.dupefilters import RFPDupeFilter from scrapinghub import ScrapinghubClient class ScrapingHubKVSDupeFilter(RFPDupeFilter): def __init__(self, path, debug, spider): super().__init__(path, debug) self.client = ScrapinghubClient() self.project = self.client.get_project(spider.settings.get('SCRAPINGHUB_PROJECT_ID')) self.kvs = self.project.kvs def request_seen(self, request): fp = self.request_fingerprint(request) # Check if the URL fingerprint already exists in KVS if self.kvs.get(fp): return True # If not, store the fingerprint to mark it as seen self.kvs.set(fp, b'1') return False
2. Update Your Scrapy Settings
In your project's settings.py, tell Scrapy to use your custom filter and add your project ID:
# Point to your custom dupe filter DUPEFILTER_CLASS = 'your_project_name.dupefilters.ScrapingHubKVSDupeFilter' # Replace with your actual ScrapingHub project ID (found in your project dashboard) SCRAPINGHUB_PROJECT_ID = '123456'
3. Quick Notes for Production
- KVS has storage limits, so if you're crawling millions of URLs, consider adding logic to clean up old fingerprints (e.g., delete entries older than 30 days) to avoid hitting caps.
- If KVS isn't sufficient for your scale, you could also use an external database like Redis—just adjust the custom filter to connect to Redis instead of KVS.
内容的提问来源于stack exchange,提问作者Marcus Christiansen

