You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy跨ScrapingHub定时任务的URL去重机制及配置咨询

Scrapy URL Deduplication Across ScrapingHub Scheduled Jobs

Great question! Let me break down how this works based on my hands-on experience with Scrapy and ScrapingHub.

Default Deduplication: No Cross-Job Persistence

Out of the box, Scrapy uses the RFPDupeFilter which stores URL fingerprints only in memory during a single crawler run. When you set up daily scheduled jobs on ScrapingHub, each run spins up a completely fresh, isolated environment. That means the in-memory deduplication cache gets wiped clean every time a new job starts.

So to put it plainly: the default deduplication does not persist across scheduled tasks. Your crawler won't remember URLs it scraped in yesterday's run when it kicks off today's job.

How to Enable Cross-Scheduled-Job Deduplication

To make deduplication stick across runs, you need to replace the in-memory storage with a persistent system that's accessible across all your ScrapingHub jobs. The easiest built-in solution is ScrapingHub's Key Value Store (KVS)—it's tightly integrated with your project and perfect for quick fingerprint lookups.

Here's a step-by-step way to set this up:

1. Build a Custom DupeFilter

Create a custom filter that inherits from Scrapy's base RFPDupeFilter, but uses ScrapingHub KVS to store and check fingerprints. Example code:

from scrapy.dupefilters import RFPDupeFilter
from scrapinghub import ScrapinghubClient

class ScrapingHubKVSDupeFilter(RFPDupeFilter):
    def __init__(self, path, debug, spider):
        super().__init__(path, debug)
        self.client = ScrapinghubClient()
        self.project = self.client.get_project(spider.settings.get('SCRAPINGHUB_PROJECT_ID'))
        self.kvs = self.project.kvs
        
    def request_seen(self, request):
        fp = self.request_fingerprint(request)
        # Check if the URL fingerprint already exists in KVS
        if self.kvs.get(fp):
            return True
        # If not, store the fingerprint to mark it as seen
        self.kvs.set(fp, b'1')
        return False

2. Update Your Scrapy Settings

In your project's settings.py, tell Scrapy to use your custom filter and add your project ID:

# Point to your custom dupe filter
DUPEFILTER_CLASS = 'your_project_name.dupefilters.ScrapingHubKVSDupeFilter'
# Replace with your actual ScrapingHub project ID (found in your project dashboard)
SCRAPINGHUB_PROJECT_ID = '123456'

3. Quick Notes for Production

  • KVS has storage limits, so if you're crawling millions of URLs, consider adding logic to clean up old fingerprints (e.g., delete entries older than 30 days) to avoid hitting caps.
  • If KVS isn't sufficient for your scale, you could also use an external database like Redis—just adjust the custom filter to connect to Redis instead of KVS.

内容的提问来源于stack exchange,提问作者Marcus Christiansen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:10:09