Scrapy爬取无<a>节点/无href属性页面PDF文件求助
解决方法
这个网站是单页应用(SPA),页面内容通过JavaScript动态加载,直接爬取静态HTML拿不到数据,PDF下载也依赖API接口生成链接,得通过抓包找到实际的接口来实现爬取和分页。
步骤1:抓包分析API接口
打开浏览器开发者工具(F12)的Network标签,按操作点击「Buscadores」→「Productos」,找到两类关键请求:
- 列表数据接口:加载每页产品数据的AJAX请求(一般是XHR类型,URL包含
productos和分页参数) - PDF下载接口:点击下载按钮时触发的请求,用按钮的
data-id作为参数,返回PDF的真实下载链接
步骤2:修改Scrapy代码
以下是适配该网站的完整代码(需根据实际抓包到的接口调整URL和返回结构):
import scrapy from scrapy.http import JsonRequest class RegfiwebSpider(scrapy.Spider): name = "regfiweb" allowed_domains = ["servicio.mapa.gob.es"] # 替换为你抓包得到的列表接口,{page}是分页占位符 base_list_api = "https://servicio.mapa.gob.es/regfiweb/api/productos?page={}" # 替换为抓包得到的下载接口,{item_id}对应按钮的data-id download_api = "https://servicio.mapa.gob.es/regfiweb/api/productos/{}/download" start_page = 1 def start_requests(self): # 发起第一页列表请求 yield JsonRequest( url=self.base_list_api.format(self.start_page), callback=self.parse_product_list, headers={ "Referer": "https://servicio.mapa.gob.es/regfiweb#", "X-Requested-With": "XMLHttpRequest" } ) def parse_product_list(self, response): response_data = response.json() # 遍历当前页的所有产品 for product in response_data.get("items", []): product_id = product.get("id") product_title = product.get("nombre") # 发起下载链接请求 yield JsonRequest( url=self.download_api.format(product_id), callback=self.parse_download_link, meta={"title": product_title}, headers={ "Referer": "https://servicio.mapa.gob.es/regfiweb#", "X-Requested-With": "XMLHttpRequest" } ) # 处理分页:判断是否还有下一页 current_page = response_data.get("currentPage") total_pages = response_data.get("totalPages") if current_page and total_pages and current_page < total_pages: next_page = current_page + 1 yield JsonRequest( url=self.base_list_api.format(next_page), callback=self.parse_product_list, headers={ "Referer": "https://servicio.mapa.gob.es/regfiweb#", "X-Requested-With": "XMLHttpRequest" } ) def parse_download_link(self, response): # 若接口直接重定向到PDF,response.url就是真实下载地址;若返回JSON则需提取对应字段 pdf_url = response.url yield { "Title": response.meta["title"], "file_urls": [pdf_url] }
步骤3:配置文件下载管道
在settings.py中启用Scrapy的文件下载管道,指定PDF保存目录:
ITEM_PIPELINES = { 'scrapy.pipelines.files.FilesPipeline': 1, } FILES_STORE = "./downloaded_pdfs" # 本地保存目录,会自动创建
关键说明
- 你原来的代码无效是因为:静态HTML里没有产品数据,所有内容都是JS通过API拉取的,所以
response.css('.col')找不到任何元素 - 必须用抓包工具确认真实接口的URL、请求参数和返回结构,上述代码中的接口是示例,需要你根据实际情况修改
- 若遇到反爬,可在
settings.py中配置USER_AGENT模拟浏览器请求
内容的提问来源于stack exchange,提问作者Luis Miguel
相关产品推荐
相关产品推荐

