You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy解析器异常:药店药品抓取不全及代码优化求助

问题说明

现有Scrapy爬虫可遍历tabletka.by平台所有药店名称,但无法抓取全部药店下的药品:部分药店药品能完整抓取(超1000条),部分仅抓取第一页20条。推测问题出在page_medicines、page_pharmacies计数器重置逻辑错误,以及页面重复遍历(DEBUG日志存在重复请求),需要拆分分页逻辑并修正代码。

问题根源
  • 类级计数器线程不安全:Scrapy是异步框架,多个请求共享类属性page_pharmacies和page_medicines,导致计数器值被不同请求篡改,分页逻辑混乱。
  • 错误的计数器重置时机:在parse_medicines开头强制重置page_medicines=1,会导致分页请求时计数器被重置,无法继续翻页。
  • 硬编码分页URL:药店分页直接拼接URL,未利用页面现有分页元素,且循环逻辑错误,导致重复请求同一页面。
  • 药品分页循环逻辑错误:使用while循环生成分页请求,会一次性生成大量重复请求,且依赖类计数器导致分页混乱。
修正后的爬虫代码
import scrapy
from urllib.parse import urljoin
from ..items import AptekaItem

class AptSpider(scrapy.Spider):
    name = 'apt'
    allowed_domains = ['tabletka.by']
    # 可选:切换地区,1006为格罗德诺州,38为格罗德诺市
    start_urls = ['https://tabletka.by/pharmacies?region=38&page=1&sort=name&sorttype=asc']  

    def parse(self, response):
        # 遍历当前页所有药店
        for row in response.css("tbody tr"):
            items = AptekaItem()
            items['name_of_pharmacy'] = row.css(".pharm-name .text-wrap a::text").get()
            items['location_of_pharmacy'] = row.css(".tooltip-info-header .text-wrap span::text").get()
            items['number_of_pharmacy'] = row.css(".phone.tooltip-info .tooltip-info-header .text-wrap a::text").get()
            
            inner_link = urljoin('https://tabletka.by/', row.css(".pharm-name .text-wrap a::attr(href)").get())
            # 传递药店信息到药品解析函数,初始页码设为1
            yield response.follow(inner_link, callback=self.parse_medicines, meta={
                'items': items,
                'current_page': 1
            })
        
        # 处理药店分页:提取下一页URL
        next_page = response.css(".table-pagination a.next::attr(href)").get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

    def parse_medicines(self, response):
        current_page = response.meta['current_page']
        items_template = response.meta['items']

        # 遍历当前页所有药品
        for low in response.css("tbody tr"):
            items = items_template.copy()  # 复制模板,避免数据覆盖
            items['name_of_medicine'] = low.css(".name.tooltip-info .tooltip-info-header a::text").get()
            
            # 处理活性成分/类型的两种结构
            a_or_t = low.css(".name.tooltip-info .capture::text").get()
            if a_or_t and a_or_t.strip():
                items['active_ingredient_or_type'] = a_or_t.strip()
            else:
                items['active_ingredient_or_type'] = low.css(".name.tooltip-info .capture a::text").get().strip()
            
            items['dosage_form'] = low.css(".form-title::text").get()
            items['prescribed'] = low.css(".form.tooltip-info .capture::text").get()
            items['name_of_manufacturer'] = low.css(".produce.tooltip-info .tooltip-info-header span a::text").get().strip()
            items['country_of_manufaturer'] = low.css(".produce.tooltip-info .capture::text").get().strip()
            items['price_of_medicine'] = low.css(".price-value::text").get().strip()
            items['page'] = str(current_page)
            
            yield items
        
        # 处理药品分页:提取下一页URL
        next_page = response.css(".table-pagination a.next::attr(href)").get()
        if next_page:
            yield response.follow(
                next_page, 
                callback=self.parse_medicines, 
                meta={
                    'items': items_template,
                    'current_page': current_page + 1
                }
            )
优化后的配置文件(settings.py)
BOT_NAME = 'apteka'

SPIDER_MODULES = ['apteka.spiders']
NEWSPIDER_MODULE = 'apteka.spiders'

# 使用更真实的User-Agent,避免被反爬拦截
USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"

ROBOTSTXT_OBEY = True

# 降低并发请求数,避免给服务器造成过大压力,同时减少被拦截概率
CONCURRENT_REQUESTS = 8

# 可选:添加下载延迟,进一步优化爬取行为
DOWNLOAD_DELAY = 0.5
关键改进说明
  • 移除类级计数器:改用meta传递当前页码,每个请求独立维护分页状态,避免异步环境下的变量冲突。
  • 基于页面元素的分页逻辑:不再硬编码URL,而是从页面提取下一页按钮的href属性,更稳定且适配网站结构变化。
  • 复制Item对象:在遍历药品时复制初始的药店Item模板,避免因引用传递导致的字段覆盖问题。
  • 修复计数器重置逻辑:删除错误的self.page_medicines=1重置代码,分页状态由meta中的current_page维护。
  • 优化并发配置:降低并发数并添加下载延迟,减少被目标网站反爬机制拦截的概率。

内容的提问来源于stack exchange,提问作者seblful

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 20:30:48