You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取无<a>节点/无href属性页面PDF文件求助

解决方法

这个网站是单页应用(SPA),页面内容通过JavaScript动态加载,直接爬取静态HTML拿不到数据,PDF下载也依赖API接口生成链接,得通过抓包找到实际的接口来实现爬取和分页。

步骤1:抓包分析API接口

打开浏览器开发者工具(F12)的Network标签,按操作点击「Buscadores」→「Productos」,找到两类关键请求:

  • 列表数据接口:加载每页产品数据的AJAX请求(一般是XHR类型,URL包含productos和分页参数)
  • PDF下载接口:点击下载按钮时触发的请求,用按钮的data-id作为参数,返回PDF的真实下载链接

步骤2:修改Scrapy代码

以下是适配该网站的完整代码(需根据实际抓包到的接口调整URL和返回结构):

import scrapy
from scrapy.http import JsonRequest

class RegfiwebSpider(scrapy.Spider):
    name = "regfiweb"
    allowed_domains = ["servicio.mapa.gob.es"]
    # 替换为你抓包得到的列表接口,{page}是分页占位符
    base_list_api = "https://servicio.mapa.gob.es/regfiweb/api/productos?page={}"
    # 替换为抓包得到的下载接口,{item_id}对应按钮的data-id
    download_api = "https://servicio.mapa.gob.es/regfiweb/api/productos/{}/download"
    start_page = 1

    def start_requests(self):
        # 发起第一页列表请求
        yield JsonRequest(
            url=self.base_list_api.format(self.start_page),
            callback=self.parse_product_list,
            headers={
                "Referer": "https://servicio.mapa.gob.es/regfiweb#",
                "X-Requested-With": "XMLHttpRequest"
            }
        )

    def parse_product_list(self, response):
        response_data = response.json()
        # 遍历当前页的所有产品
        for product in response_data.get("items", []):
            product_id = product.get("id")
            product_title = product.get("nombre")
            # 发起下载链接请求
            yield JsonRequest(
                url=self.download_api.format(product_id),
                callback=self.parse_download_link,
                meta={"title": product_title},
                headers={
                    "Referer": "https://servicio.mapa.gob.es/regfiweb#",
                    "X-Requested-With": "XMLHttpRequest"
                }
            )

        # 处理分页:判断是否还有下一页
        current_page = response_data.get("currentPage")
        total_pages = response_data.get("totalPages")
        if current_page and total_pages and current_page < total_pages:
            next_page = current_page + 1
            yield JsonRequest(
                url=self.base_list_api.format(next_page),
                callback=self.parse_product_list,
                headers={
                    "Referer": "https://servicio.mapa.gob.es/regfiweb#",
                    "X-Requested-With": "XMLHttpRequest"
                }
            )

    def parse_download_link(self, response):
        # 若接口直接重定向到PDF,response.url就是真实下载地址;若返回JSON则需提取对应字段
        pdf_url = response.url
        yield {
            "Title": response.meta["title"],
            "file_urls": [pdf_url]
        }

步骤3:配置文件下载管道

在settings.py中启用Scrapy的文件下载管道,指定PDF保存目录:

ITEM_PIPELINES = {
    'scrapy.pipelines.files.FilesPipeline': 1,
}
FILES_STORE = "./downloaded_pdfs"  # 本地保存目录,会自动创建

关键说明

  • 你原来的代码无效是因为:静态HTML里没有产品数据,所有内容都是JS通过API拉取的,所以response.css('.col')找不到任何元素
  • 必须用抓包工具确认真实接口的URL、请求参数和返回结构,上述代码中的接口是示例,需要你根据实际情况修改
  • 若遇到反爬,可在settings.py中配置USER_AGENT模拟浏览器请求

内容的提问来源于stack exchange,提问作者Luis Miguel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 07:20:19