You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Scrapy代码实现多字母标签分页爬取1mg药品数据

问题描述

使用Scrapy爬取https://www.1mg.com/drugs-all-medicines网站的药品数据,需求是按字母标签遍历并爬取每个标签下的分页数据。当前代码能完成字母a的所有页面爬取,但切换到字母b时仅能爬取第一页,随后因重复请求无label参数的page=2 URL而停止,报错Filtered duplicate request。

网站URL规则:

  • 字母a的分页URL:https://www.1mg.com/drugs-all-medicines?page=N
  • 字母b及以后的分页URL:https://www.1mg.com/drugs-all-medicines?page=N&label=X(需同时携带label和page参数)

现有代码如下:

import scrapy
import re

class MedSpider(scrapy.Spider):
    name = "medspider"
    allowed_domains = ["www.1mg.com"]
    start_urls = ["https://www.1mg.com/drugs-all-medicines"]
    current_alphabet = 'a'  # Initial alphabet label
    current_page = 2  # Initial page number

    def parse(self, response):
        meds = response.css('div.style__flex-1___A_qoj')

        for med in meds:
            yield {
                'name': med.css('div div::text').get(),
                'price': med.css('div:has(> span)::text').getall()[-1],
                'strip content': med.css('div::text').getall()[-4],
                'manufacturer': med.css('div::text').getall()[-3],
            }

        next_page = response.css('li.next a::attr(href)').get()
        if next_page is not None and self.current_page <= 1076:
            url_page = 'https://www.1mg.com/drugs-all-medicines?page=' + str(self.current_page)
            self.current_page += 1  # Increment the page number
            yield response.follow(url_page, callback=self.parse)
        else:
            if self.current_alphabet < 'z':
                self.current_alphabet = chr(ord(self.current_alphabet) + 1)  # Increment the alphabet label
                self.current_page = 2  # Reset the page number
                url_label = 'https://www.1mg.com/drugs-all-medicines?label=' + self.current_alphabet
                yield response.follow(url_label, callback=self.parse)
解决方案

问题核心:切换到字母b及以后的标签时,构造分页URL未携带当前label参数,导致请求的URL与字母a的分页URL重复,被Scrapy去重机制过滤。

需做以下修改:

  1. 移除全局状态变量:删除current_alphabet和current_page全局变量,改用从响应URL中动态提取当前参数,避免并发请求时的状态混乱。
  2. 复用页面自带的分页链接:直接使用页面提供的next_page链接,该链接已自动包含正确的label和page参数,无需手动拼接。
  3. 动态切换字母标签:当当前字母无下一页时,自动计算并切换到下一个字母标签。

修改后的完整代码:

import scrapy
from urllib.parse import urlparse, parse_qs

class MedSpider(scrapy.Spider):
    name = "medspider"
    allowed_domains = ["www.1mg.com"]
    start_urls = ["https://www.1mg.com/drugs-all-medicines"]

    def parse(self, response):
        # 提取当前页面药品数据
        meds = response.css('div.style__flex-1___A_qoj')
        for med in meds:
            yield {
                'name': med.css('div div::text').get(),
                'price': med.css('div:has(> span)::text').getall()[-1],
                'strip content': med.css('div::text').getall()[-4],
                'manufacturer': med.css('div::text').getall()[-3],
            }

        # 获取页面自带的下一页链接(自动包含正确参数)
        next_page = response.css('li.next a::attr(href)').get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)
        else:
            # 解析当前URL的label参数,切换到下一个字母
            parsed_url = urlparse(response.url)
            query_params = parse_qs(parsed_url.query)
            current_label = query_params.get('label', [None])[0]

            # 计算下一个字母
            if not current_label:
                # 当前是字母a,切换到b
                next_label = 'b'
            else:
                if current_label < 'z':
                    next_label = chr(ord(current_label) + 1)
                else:
                    # 所有字母遍历完成,终止
                    return

            # 构造下一个字母的第一页URL
            next_label_url = f'https://www.1mg.com/drugs-all-medicines?label={next_label}'
            yield response.follow(next_label_url, callback=self.parse)

额外说明:

  • 使用urllib.parse解析URL参数,能准确获取当前页面的标签状态,避免全局变量的并发冲突。
  • 优先复用页面原生分页链接,无需手动拼接参数,适配网站URL规则的同时降低出错概率。
  • 自动遍历所有字母标签,直到完成全部数据爬取。

内容的提问来源于stack exchange,提问作者Lucyfer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 13:13:09