You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现Scrapy爬虫循环爬取同站点不同搜索关键词页面

代码调整方案

原代码核心问题

  • 类加载阶段直接写Selenium访问逻辑完全不符合Scrapy运行流程:Scrapy爬虫类加载时就会执行类内部的顶层for循环,这时候爬取到的页面响应根本不会传入parse方法,完全走不到数据解析逻辑
  • 全局初始化Chrome驱动、用全局索引i遍历关键词的写法容易出现资源泄漏、索引越界问题
  • 翻页逻辑缺失:代码里的next_page变量从未定义,无法实现搜索结果的自动翻页
  • 冗余依赖过多:不需要引入numpy做页码生成,直接遍历关键词+动态找下一页即可
  • 字段名存在拼写错误:原代码里的tittle为拼写错误,正确写法为title

调整思路

  • 用Scrapy标准的start_requests方法生成所有关键词的初始搜索请求,统一交给Scrapy调度
  • 把Selenium驱动初始化放到爬虫类的初始化方法中,绑定到爬虫实例上,爬虫关闭时自动退出驱动
  • 每个关键词的搜索页解析完成后,自动提取下一页链接,持续翻页爬取直到该关键词下没有更多结果
  • 去掉全局变量、冗余依赖,访问前加随机等待模拟真人操作

调整后可运行代码

import scrapy
from selenium import webdriver
from time import sleep
from random import randint

class UmpSpider(scrapy.Spider):
    name = "pubmed"
    # 待爬取的关键词列表,不需要手动做URL编码
    keywords = [
        "RNA Delivery vehicles",
        "animal models for rare genetic diseases",
        "invitro models for rare genetic diseases",
        "genes",
        "muscle",
        "structural Visualization of LNP"
    ]

    def __init__(self, **kwargs):
        super().__init__(**kwargs)
        # 初始化Selenium驱动,替换成你本地的chromedriver路径
        driver_path = "/Users/miguelcorredor/Desktop/cnn/chromedriver"
        self.driver = webdriver.Chrome(executable_path=driver_path)

    def start_requests(self):
        # 遍历所有关键词生成初始搜索请求
        for keyword in self.keywords:
            search_url = f"https://pubmed.ncbi.nlm.nih.gov/?term={keyword.replace(' ', '+')}"
            yield scrapy.Request(
                url=search_url,
                callback=self.parse,
                # 把当前关键词传到解析方法里,方便后续存数据的时候标记来源关键词
                cb_kwargs={"keyword": keyword}
            )

    def parse(self, response, keyword):
        # 随机等待2-10秒模拟真人访问
        sleep(randint(2,10))
        # 用Selenium加载当前页,获取渲染后的页面源码
        self.driver.get(response.url)
        page_source = self.driver.page_source
        # 把渲染后的源码重新构造为Selector对象供解析使用
        sel = scrapy.Selector(text=page_source)

        # 解析当前页的文献数据
        for product in sel.css('div.docsum-content'):
            raw_title = product.css('a.docsum-title::text').get()
            yield{
                'keyword': keyword, # 标记数据属于哪个关键词的搜索结果
                'title': raw_title.strip() if raw_title else None,
                'author': product.css('span.docsum-authors.short-authors::text').get(),
                'year': product.css('span.docsum-journal-citation.short-journal-citation::text').get(),
            }
        
        # 查找下一页按钮
        next_page = sel.css('div.pagination a.next-page::attr(href)').get()
        if next_page:
            next_url = response.urljoin(next_page)
            yield scrapy.Request(
                url=next_url,
                callback=self.parse,
                cb_kwargs={"keyword": keyword}
            )
    
    def closed(self, reason):
        # 爬虫结束时自动关闭浏览器驱动,释放资源
        self.driver.quit()

关键调整说明

  • 新增cb_kwargs参数把当前搜索关键词传递给解析方法,爬下来的数据可以直接对应到检索词,方便后续分类存储
  • 新增closed钩子方法,爬虫不管是正常结束还是异常退出都会自动关闭Chrome进程,不会残留后台驱动
  • 翻页逻辑自动适配所有关键词的搜索结果,不需要提前指定每个关键词爬多少页,会自动爬到最后一页停止
  • 去掉了原代码里没用的numpy、多余time模块、Keys等冗余导入,代码更简洁
  • 对标题字段加了空值判断和去空格处理,避免取到空值或者带大量换行空格的脏数据

内容的提问来源于stack exchange,提问作者miguel corredor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 20:27:26