You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy可正常获取完整数据但无法下载文件且无报错,求排查方向

问题描述

运行Scrapy爬虫时,yield article能正常获取完整的文章数据,但下载XML文件的逻辑完全没生效——CSV里的files字段为空,local_folder目录里也没文件,还没任何报错,根本不知道哪出问题。


相关代码

thespider.py

import re
import scrapy
import unidecode
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from scrapy.loader import ItemLoader
from dhqscraper.items import FileDownloader

class DhqSpider(CrawlSpider):
    name = 'dhq'
    allowed_domains = ['digitalhumanities.org']
    start_urls = ['http://www.digitalhumanities.org/dhq/vol/16/3/index.html']

    rules = (
            Rule(LinkExtractor(allow = 'index.html')), 
            Rule(LinkExtractor(allow = 'vol'), callback='parse_article'),        
        )

    def parse_article(self, response):
        article = { 
            'title' : response.css('h1.articleTitle').xpath('normalize-space(text())').get(default='').strip(),
            'author' : response.css('div.author a').xpath('normalize-space(text())').getall(),
            'year' : unidecode.unidecode(response.css('div#pubInfo::text')[0].get()),
            'volume' : unidecode.unidecode(response.css('div#pubInfo::text')[1].get()),
            'xmllink' : response.urljoin(response.xpath('(//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href)[1]').get()),
        }
        article ['author'] = [unidecode.unidecode(i) for i in article['author']] 
        yield article
        
        def download(self, response):
            file_url = article['xmllink']
            item = FileDownloader()
            item['file_urls'] = [file_url]
            yield item

items.py

import scrapy
from scrapy.loader import ItemLoader
from scrapy.linkextractors import LinkExtractor
from itemloaders.processors import TakeFirst, MapCompose
from w3lib.html import remove_tags

class FileDownloader(scrapy.Item):
    file_urls = scrapy.Field()
    files = scrapy.Field

settings.py

ITEM_PIPELINES = {'scrapy.pipelines.files.FilesPipeline': 1}
FILES_STORE = 'local_folder'

排查与解决方向

1. 嵌套函数未被触发

你把download函数定义在parse_article内部,但从未调用过它。Scrapy不会自动执行嵌套的回调函数,直接在parse_article里生成下载Item即可:

def parse_article(self, response):
    # 原有的article逻辑保持不变
    yield article
    
    # 直接生成FileDownloader item,去掉嵌套函数
    item = FileDownloader()
    item['file_urls'] = [article['xmllink']]
    yield item

2. Item字段定义错误

files字段少了括号,应该是scrapy.Field()而非scrapy.Field,这会导致FilesPipeline无法正确识别该字段,无法写入下载后的文件信息:

class FileDownloader(scrapy.Item):
    file_urls = scrapy.Field()
    files = scrapy.Field()  # 这里必须加括号

3. 验证XML链接有效性

在parse_article中添加打印语句,确认xmllink是否正确生成:

print("XML链接:", article['xmllink'])

如果链接为空或无效,需调整XPath规则。同时要处理get()返回None的情况,避免生成无效链接:

xml_href = response.xpath('(//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href)[1]').get()
article['xmllink'] = response.urljoin(xml_href) if xml_href else ''

4. 启用DEBUG日志排查

运行爬虫时调高日志级别,查看FilesPipeline的执行细节:

scrapy crawl dhq --logfile=scrapy.log --loglevel=DEBUG

日志中会显示文件下载的状态(如是否成功、是否有404错误),帮助定位隐藏问题。

5. 检查FilesPipeline配置

确认settings.py中FILES_STORE路径正确,爬虫运行目录有写入权限。另外,确保没有其他Pipeline干扰FilesPipeline的执行(当前优先级1为最高,没问题)。


内容的提问来源于stack exchange,提问作者David Beales

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 00:05:35