You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何阻止scrapy对本地URL编码的文件名进行解码?

问题

本地保存的网页文件名用urllib.parse.quote做了URL编码,示例文件名:

https%3A%2F%2Fwebsite.com%2Flist%2F27-6.html

用Scrapy提取这些网页时出现错误:

DEBUG: Retrying <GET file:///local_dir/https%3A%2F%2Fwebsite.com%2Flist%2F27-6.html> (failed 2 times): [Errno 2] No such file or directory: '/local_dir/https:/website.com/list/27-6.html'

显然Scrapy自动解码了文件名,导致找不到对应文件。尝试以下配置无效:

custom_settings = {
    'FILES_URLS_FIELD': 'file_urls',
    'FILES_RESULT_FIELD': 'files',
    'FILES_URLS_FILENAME_FIELD': None,
}
解决办法

方法一:对编码后的文件名二次编码

构造file URL时,把已经编码的文件名再做一次URL编码,这样Scrapy解码一次后就会得到正确的原始编码文件名。

示例代码:

import urllib.parse

# 原始编码后的文件名
encoded_filename = "https%3A%2F%2Fwebsite.com%2Flist%2F27-6.html"
# 二次编码
double_encoded = urllib.parse.quote(encoded_filename)
# 构造正确的file URL
file_url = f"file:///local_dir/{double_encoded}"

方法二:自定义FilesPipeline跳过解码

重写Scrapy的FilesPipeline,直接使用URL中的原始文件名,不做解码处理。

  1. 在项目的pipelines.py中添加自定义管道:
from scrapy.pipelines.files import FilesPipeline

class CustomEncodedFilesPipeline(FilesPipeline):
    def file_path(self, request, response=None, info=None, *, item=None):
        # 直接截取URL最后一段作为文件名,不进行解码
        return request.url.split('/')[-1]
  1. 在settings.py中替换默认的FilesPipeline:
ITEM_PIPELINES = {
    # 替换成你自定义管道的实际路径
    'your_project_name.pipelines.CustomEncodedFilesPipeline': 1,
}

内容的提问来源于stack exchange,提问作者Lei Hao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 12:11:06