Scrapy仅爬取少量条目问题求助(附代码)
Scrapy爬取billa.at仅获取部分条目问题排查
我正在使用Scrapy练习爬取billa.at网站,该网站每页最多展示60个条目。为提取每个条目的更多数据,我会访问每个条目的详情页,但目前仅能从第一页爬取约14个条目(不过爬虫至少遍历了所有条目)。曾尝试过类似问题的解决方案,但对我无效。我使用Selenium可以爬取所有条目,但更倾向于使用速度更快的Scrapy,希望解决此爬取不全的问题。
以下是我的代码:
import scrapy from scrapy_playwright.page import PageMethod from product_spider.items import ProductItem class ProductSpider(scrapy.Spider): name = "product_spider" allowed_domains = ["billa.at"] base_url = "https://shop.billa.at/" def start_requests(self): url = "https://shop.billa.at/warengruppe/getraenke/alkoholfreie-getraenke/mineralwasser-mit-geschmack/B2-318" yield scrapy.Request(url, meta=dict( page = 1, playwright = True, playwright_include_page = True, playwright_page_methods = [ PageMethod("wait_for_selector", "a.product__content-title-sc-link"), ], errback = self.errback, ), callback=self.temp_parse) def temp_parse(self, response): # extract all product urls of a page all_product_urls_of_a_page = response.css("a.product__content-title-sc-link ::attr(href)").getall() #if not all_product_urls_of_a_page: # raise scrapy.exceptions.CloseSpider("All products scraped") for url in all_product_urls_of_a_page: product_url = self.base_url + url yield scrapy.Request(product_url, meta=dict( playwright = True, playwright_include_page = True, playwright_page_methods = [ PageMethod("wait_for_selector", "img.loaded.fade-in"), ], errback = self.errback, ), callback=self.parse_product_page) # check if there is a next page if not response.css("button.pagination__item.pagination__item--next.ng-scope ::attr(disabled)"): next_page_url = response.url + f"?page={response.meta['page'] + 1}" yield scrapy.Request(next_page_url, meta=dict( page = response.meta['page'] + 1, playwright = True, playwright_include_page = True, playwright_page_methods = [ PageMethod("wait_for_selector", "a.product__content-title-sc-link"), ], errback = self.errback, ), callback=self.temp_parse) async def errback(self, failure): print("IN ERRBACK") page = failure.request.meta["playwright_page"] await page.close() def parse_product_page(self, response): product_item = ProductItem() product_item["product"] = response.css("img.loaded.fade-in ::attr(alt)").extract_first() product_item["image_url"] = response.css("img.loaded.fade-in ::attr(src)").extract_first() yield product_item
内容的提问来源于stack exchange,提问作者Rema
相关产品推荐
相关产品推荐

