You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy XPath无法获取所有商品价格问题求助

Fixing XPath Issues for Scraping Tesco Product Prices

Hey Fiona, let's figure out why your XPath isn't capturing all product prices from that Tesco page. The problem almost always comes down to inconsistent HTML structure across different product tiles—let's walk through how to diagnose and fix this.

First, Diagnose with Scrapy Shell

The best way to debug XPath issues is to use Scrapy's interactive shell to inspect the page structure in real time. Run this command in your terminal:

scrapy shell 'https://www.tesco.com/direct/tv-offer.event?icid=offers_trade_slot1'

Once the shell loads, you can test your selectors one by one. For example, grab a single product tile and check its price elements:

product = response.xpath('//li[contains(@class,"product-tile")]')[0]
# Check all p tags inside the buy-box-container to see where the price lives
product.xpath('.//div[@class="buy-box-container"]/p').getall()

You'll probably notice that some products have their price in a different <p> tag (not always the second one) or even wrapped in a <span> instead—this is why your original XPath misses some prices.

Updated XPath & Code

Here's a more flexible approach that accounts for varying HTML structures:

class QuotesSpider(scrapy.Spider):
    name = "tesco"
    start_urls = [
        'https://www.tesco.com/direct/tv-offer.event?icid=offers_trade_slot1',
    ]
    def parse(self, response):
        for product in response.xpath('//li[contains(@class,"product-tile")]'):
            item = crawlerItem()
            
            # Extract and clean the title
            title = product.xpath('.//h3/a/text()').extract_first()
            item['title'] = title.strip() if title else None
            
            # Try multiple selectors to catch all price variations
            price = product.xpath('.//div[@class="buy-box-container"]//p[contains(@class, "price")]/text()').extract_first()
            # Fallback to span elements if p tag doesn't work
            if not price:
                price = product.xpath('.//div[@class="buy-box-container"]//span[contains(@class, "value")]/text()').extract_first()
            
            # Clean up the price and handle missing values
            item['price'] = price.strip() if price else "Price unavailable"
            
            yield item

Why This Works

  • Flexible Selectors: Instead of hardcoding <p[2]>, we look for elements with price in their class name—this covers regular prices, sale prices, and any other price variations.
  • Fallback Logic: If the first selector fails, we try a common alternative (price in a <span>) to avoid missing data.
  • Data Cleaning: Using .strip() removes extra whitespace from titles and prices, making your data cleaner.

Bonus Tip: Handle Dynamic Content

If some products load only when you scroll down, you might need to handle JavaScript-rendered content. For that, you can integrate Scrapy with tools like Splash or Playwright to fully render the page before scraping.

内容的提问来源于stack exchange,提问作者Fiona Stone

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:14:59