You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Scrapy抓取start_urls页面内的data-url而非直接解析start_urls

Scrapy 全量数据爬取修改方案

核心调整点

  • 原搜索结果页不再直接解析数据,优先提取id=loadMoreShops节点的data-url属性,发起请求获取全量数据
  • 移除全局计数变量self.i,改用请求meta参数传递当前查询的PZN、商品名称,彻底避免并发请求时数据错位
  • 修正allowed_domains配置错误,域名无需携带HTTP协议头
  • 新增统一的全量数据解析方法,兼容无data-url的场景(如本身结果较少无需加载更多的情况)

完整可运行代码

import scrapy, bs4, re, csv, logging
from datetime import datetime


class ApomioSpider(scrapy.Spider):
    # 修正域名配置,不需要带协议头
    allowed_domains = ['apomio.de']
    name = "apomio"

    def start_requests(self):
        self.PZN_arr = []
        self.product_names = []

        with open('PZN.csv', encoding='utf-8') as csv_file:
            reader = csv.reader(csv_file)
            next(reader)  # 跳过csv表头
            for elem in reader:
                self.PZN_arr.append(elem[1])
                self.product_names.append(elem[2])

        # 遍历查询参数,通过meta传递对应PZN和商品名,避免全局计数混乱
        for idx, pzn in enumerate(self.PZN_arr):
            url = f'https://www.apomio.de/suche?query={pzn}'
            yield scrapy.Request(
                url,
                callback=self.parse_search_page,
                meta={
                    'pzn': pzn,
                    'product_name': self.product_names[idx]
                }
            )

    def parse_search_page(self, response):
        soup = bs4.BeautifulSoup(response.text, 'lxml')
        # 提取全量数据页的data-url
        load_more_node = soup.find('div', id='loadMoreShops')
        if load_more_node and load_more_node.get('data-url'):
            full_data_url = load_more_node['data-url']
            # 携带原有meta信息,发起全量数据页请求
            yield scrapy.Request(
                full_data_url,
                callback=self.parse_full_data,
                meta=response.meta
            )
        else:
            # 无加载更多按钮,直接解析当前页数据
            yield from self.parse_full_data(response)

    def parse_full_data(self, response):
        pzn = response.meta['pzn']
        product_name = response.meta['product_name']
        soup = bs4.BeautifulSoup(response.text, 'lxml')

        price_arr = []
        names_arr = []

        prices = soup.find_all("span", {"class": "block text-xs text-black font-medium"})
        names = soup.find_all("span", {"class": "w-5/6 block text-xs text-black-darker mb-2"})

        for name in names:
            name_text = name.get_text(strip=True)
            # 整理多余空格
            cleaned_name = ' '.join([part for part in name_text.split() if part])
            names_arr.append(cleaned_name)

        for price in prices:
            prices_regex = re.compile(r'(Gesamtkosten)([ ])([0-9]+)([,])([0-9]+)')
            match_res = prices_regex.search(str(price))
            if match_res:
                result_price = ".".join(match_res.group(3, 5))
                price_arr.append(float(result_price))

        if price_arr and names_arr:
            logging.info(f'Parsing prices of {product_name}\n PZN: {pzn} on {datetime.now()}\n')
            for price, name in zip(price_arr, names_arr):
                shop = f"Shop Name: {name}"
                print(f"{shop.ljust(75, ' ')} Price: {price:.2f} €")
        else:
            logging.info(f'Item could not be found at {response.url}, PZN: {pzn}')

内容的提问来源于stack exchange,提问作者Rj Lopez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 06:27:01