You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy递归爬取网页问题:多页面保存失败与文件路径报错

解决Scrapy递归爬取子页面并保存为合法本地HTML文件的问题

让我们一步步拆解你遇到的两个核心问题:所有页面被覆盖成同一个文件,以及直接使用URL作为文件名导致的非法路径错误。

问题根源分析

  1. 第一个代码的问题:你把文件名固定写死成了pydro.html,每爬取一个新页面就会覆盖之前的文件,最终只能保留最后爬取的那个页面内容。
  2. 第二个代码的问题:直接将完整URL作为文件名,但URL里包含/、:、?这类操作系统严格禁止的文件名字符(比如Windows不允许文件名含:、/,Linux禁止/),所以触发了FileNotFoundError。

解决方案:生成合法且唯一的文件名

我们需要把URL转换为符合系统规则的文件名,同时保证每个页面的文件名不重复。具体可以通过以下步骤实现:

  • 提取URL中域名后的路径部分
  • 将所有非法字符替换为下划线(或其他合法字符)
  • 为首页这类特殊路径设置友好的默认名称

修改后的完整代码(以otego.de为例)

import scrapy
from urllib.parse import urlparse
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class OtegoSpider(CrawlSpider):
    name = "otego"
    allowed_domains = ["otego.de"]
    start_urls = ["https://www.otego.de/en/index.php"]

    rules = [
        # 限制只爬取目标域名内的链接,避免爬取外部无关网站
        Rule(LinkExtractor(allow_domains=allowed_domains), callback='parse_item', follow=True)
    ]

    def parse_item(self, response):
        print('Got a response from %s.' % response.url)
        self.logger.info('Hi this is an item page! %s', response.url)

        # 解析URL,提取路径部分
        parsed_url = urlparse(response.url)
        path = parsed_url.path

        # 处理首页/根路径的特殊情况
        if not path or path == '/':
            filename = 'otego-index.html'
        else:
            # 替换路径中的所有非法字符为下划线
            sanitized_path = path.replace('/', '_').replace(':', '_').replace('?', '_').replace('&', '_')
            # 确保文件名以.html结尾,避免无后缀的情况
            if not sanitized_path.endswith('.html'):
                sanitized_path += '.html'
            filename = f'otego{sanitized_path}'

        # 写入本地文件
        with open(filename, 'wb') as f:
            f.write(response.body)
        self.log('Saved file %s' % filename)

关键改进点说明

  • 用urllib.parse.urlparse拆分URL,只提取路径部分生成文件名,避免直接使用含非法字符的完整URL
  • 批量替换所有非法字符,保证文件名在Windows、Linux等系统下都能正常创建
  • 单独处理首页路径,避免生成类似otego_.html的奇怪文件名
  • 在LinkExtractor中添加allow_domains限制,防止爬虫意外爬取外部网站

适配pydro.com的修改要点

如果你要爬取pydro.com,只需要调整以下几个参数:

  • 把name改成"pydro"
  • allowed_domains设置为["pydro.com"]
  • start_urls改为["https://pydro.com/"]
  • 文件名前缀替换为pydro(比如filename = 'pydro-index.html'和filename = f'pydro{sanitized_path}')

这样就能递归爬取所有子页面,每个页面都保存为独立的、合法的HTML文件啦!

内容的提问来源于stack exchange,提问作者Magnus Vivadeepa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:43:14