You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy脚本使用FilesPipeline时重复下载同名PDF文件问题求助

问题分析与解决方案

问题根源

你的Scrapy脚本出现重复下载PDF的核心原因是爬虫中复用了同一个Item实例,加上FilesPipeline的处理逻辑,导致数据传递混乱:

  1. 在parse方法里,你把loader = fcpItem()放在了for循环外部,每次循环只是修改同一个Item对象的字段再yield。由于Item是可变对象,后续循环的修改会覆盖之前yield的Item数据,导致FilesPipeline接收到重复或错误的下载请求。
  2. 你的file_path方法仅用names作为文件名,若存在不同PDF对应相同names的情况,会触发Scrapy的文件名冲突处理(自动添加后缀),看起来像是重复下载。

修复步骤

1. 修正爬虫脚本的Item创建逻辑

将Item实例的创建移到for循环内部,确保每次循环生成独立的Item对象:

import scrapy
from environment.items import fcpItem

class fscSpider(scrapy.Spider):
    name = 'fsc'
    start_urls = ['https://fsc.org/en/members']

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(url, callback=self.parse)
    
    def parse(self, response):
        names_add = response.xpath(".//div[@class = 'field__item resource-item']/article//span[@class='media-caption file-caption']/text()").getall()
        url_list = response.xpath(".//div[@class = 'field__item resource-item']/article/div[@class='actions']/a//@href").getall()
        
        pdf_urls = [response.urljoin(x) for x in url_list if '#' != x]
        names = [x.split(' ')[0] for x in names_add]
        
        # 每次循环创建新的Item实例
        for nm, pd in zip(names, pdf_urls):
            item = fcpItem()
            item['names'] = nm
            item['pdfs'] = [pd]
            yield item

2. 优化FilesPipeline的文件名生成逻辑

为了避免同名文件冲突(即使Item数据唯一,也可能存在不同PDF对应相同names的情况),可以结合PDF的URL哈希值生成唯一文件名:

import hashlib
from scrapy.pipelines.files import FilesPipeline

class DownfilesPipeline(FilesPipeline):
    def file_path(self, request, response=None, info=None, item=None):
        # 生成URL的短哈希值,确保文件名唯一
        url_hash = hashlib.md5(request.url.encode('utf-8')).hexdigest()[:8]
        return f"{item['names']}_{url_hash}.pdf"

3. 验证Settings配置

确保settings.py中的文件存储配置正确,且目标目录有写入权限:

from pathlib import Path
import os

BASE_DIR = Path(__file__).resolve().parent.parent
FILES_STORE = os.path.join(BASE_DIR, 'fsc')

ROBOTSTXT_OBEY = False

FILES_URLS_FIELD = 'pdfs'
FILES_RESULT_FIELD = 'results'

ITEM_PIPELINES = {
    'environment.pipelines.pipelines.DownfilesPipeline': 150
}

验证效果

修改后重新运行爬虫,此时每个Item都是独立实例,文件名也会结合URL哈希值保证唯一性,不会再出现重复下载的问题。

内容的提问来源于stack exchange,提问作者Dollar Tune-bill

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 15:57:17