Scrapy可正常获取完整数据但无法下载文件且无报错,求排查方向
问题描述
运行Scrapy爬虫时,yield article能正常获取完整的文章数据,但下载XML文件的逻辑完全没生效——CSV里的files字段为空,local_folder目录里也没文件,还没任何报错,根本不知道哪出问题。
相关代码
thespider.py
import re import scrapy import unidecode from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor from scrapy.loader import ItemLoader from dhqscraper.items import FileDownloader class DhqSpider(CrawlSpider): name = 'dhq' allowed_domains = ['digitalhumanities.org'] start_urls = ['http://www.digitalhumanities.org/dhq/vol/16/3/index.html'] rules = ( Rule(LinkExtractor(allow = 'index.html')), Rule(LinkExtractor(allow = 'vol'), callback='parse_article'), ) def parse_article(self, response): article = { 'title' : response.css('h1.articleTitle').xpath('normalize-space(text())').get(default='').strip(), 'author' : response.css('div.author a').xpath('normalize-space(text())').getall(), 'year' : unidecode.unidecode(response.css('div#pubInfo::text')[0].get()), 'volume' : unidecode.unidecode(response.css('div#pubInfo::text')[1].get()), 'xmllink' : response.urljoin(response.xpath('(//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href)[1]').get()), } article ['author'] = [unidecode.unidecode(i) for i in article['author']] yield article def download(self, response): file_url = article['xmllink'] item = FileDownloader() item['file_urls'] = [file_url] yield item
items.py
import scrapy from scrapy.loader import ItemLoader from scrapy.linkextractors import LinkExtractor from itemloaders.processors import TakeFirst, MapCompose from w3lib.html import remove_tags class FileDownloader(scrapy.Item): file_urls = scrapy.Field() files = scrapy.Field
settings.py
ITEM_PIPELINES = {'scrapy.pipelines.files.FilesPipeline': 1} FILES_STORE = 'local_folder'
排查与解决方向
1. 嵌套函数未被触发
你把download函数定义在parse_article内部,但从未调用过它。Scrapy不会自动执行嵌套的回调函数,直接在parse_article里生成下载Item即可:
def parse_article(self, response): # 原有的article逻辑保持不变 yield article # 直接生成FileDownloader item,去掉嵌套函数 item = FileDownloader() item['file_urls'] = [article['xmllink']] yield item
2. Item字段定义错误
files字段少了括号,应该是scrapy.Field()而非scrapy.Field,这会导致FilesPipeline无法正确识别该字段,无法写入下载后的文件信息:
class FileDownloader(scrapy.Item): file_urls = scrapy.Field() files = scrapy.Field() # 这里必须加括号
3. 验证XML链接有效性
在parse_article中添加打印语句,确认xmllink是否正确生成:
print("XML链接:", article['xmllink'])
如果链接为空或无效,需调整XPath规则。同时要处理get()返回None的情况,避免生成无效链接:
xml_href = response.xpath('(//div[@class="toolbar"]/a[contains(@href, ".xml")]/@href)[1]').get() article['xmllink'] = response.urljoin(xml_href) if xml_href else ''
4. 启用DEBUG日志排查
运行爬虫时调高日志级别,查看FilesPipeline的执行细节:
scrapy crawl dhq --logfile=scrapy.log --loglevel=DEBUG
日志中会显示文件下载的状态(如是否成功、是否有404错误),帮助定位隐藏问题。
5. 检查FilesPipeline配置
确认settings.py中FILES_STORE路径正确,爬虫运行目录有写入权限。另外,确保没有其他Pipeline干扰FilesPipeline的执行(当前优先级1为最高,没问题)。
内容的提问来源于stack exchange,提问作者David Beales
相关产品推荐
相关产品推荐

