You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy仅爬取少量条目问题求助(附代码)

Scrapy爬取billa.at仅获取部分条目问题排查

我正在使用Scrapy练习爬取billa.at网站,该网站每页最多展示60个条目。为提取每个条目的更多数据,我会访问每个条目的详情页,但目前仅能从第一页爬取约14个条目(不过爬虫至少遍历了所有条目)。曾尝试过类似问题的解决方案,但对我无效。我使用Selenium可以爬取所有条目,但更倾向于使用速度更快的Scrapy,希望解决此爬取不全的问题。

以下是我的代码:

import scrapy
from scrapy_playwright.page import PageMethod
from product_spider.items import ProductItem


class ProductSpider(scrapy.Spider):
    name = "product_spider"
    allowed_domains = ["billa.at"]
    base_url = "https://shop.billa.at/"
    
    def start_requests(self):
        url = "https://shop.billa.at/warengruppe/getraenke/alkoholfreie-getraenke/mineralwasser-mit-geschmack/B2-318"
        yield scrapy.Request(url, meta=dict(
            page = 1,
            playwright = True,
            playwright_include_page = True,
            playwright_page_methods = [
                PageMethod("wait_for_selector", "a.product__content-title-sc-link"), 
            ],
            errback = self.errback,
        ), callback=self.temp_parse)         


    def temp_parse(self, response):
        
        # extract all product urls of a page
        all_product_urls_of_a_page = response.css("a.product__content-title-sc-link ::attr(href)").getall()
        
        #if not all_product_urls_of_a_page:
        #    raise scrapy.exceptions.CloseSpider("All products scraped")
        
        for url in all_product_urls_of_a_page:
            product_url = self.base_url + url
            yield scrapy.Request(product_url, meta=dict(
                playwright = True,
                playwright_include_page = True,
                playwright_page_methods = [
                    PageMethod("wait_for_selector", "img.loaded.fade-in"),
                ],
                errback = self.errback,
            ), callback=self.parse_product_page)
        
        # check if there is a next page
        if not response.css("button.pagination__item.pagination__item--next.ng-scope ::attr(disabled)"):
            next_page_url = response.url + f"?page={response.meta['page'] + 1}"
            yield scrapy.Request(next_page_url, meta=dict(
                page = response.meta['page'] + 1,
                playwright = True,
                playwright_include_page = True,
                playwright_page_methods = [
                    PageMethod("wait_for_selector", "a.product__content-title-sc-link"), 
                ],
                errback = self.errback,
            ), callback=self.temp_parse)          
        
    async def errback(self, failure):
        print("IN ERRBACK")
        page = failure.request.meta["playwright_page"]
        await page.close()
        
    def parse_product_page(self, response):
               
        product_item = ProductItem()
        
        product_item["product"] = response.css("img.loaded.fade-in ::attr(alt)").extract_first()
        product_item["image_url"] = response.css("img.loaded.fade-in ::attr(src)").extract_first()
        
        yield product_item

内容的提问来源于stack exchange,提问作者Rema

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 23:45:03