You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy提取含::before伪元素的价格为空,如何解决?

How to Extract Price from a ::before Pseudo-Element in Scrapy

Got it, let's break down why your XPath isn't working and fix it!

The problem is straightforward: ::before pseudo-elements are generated by CSS, not part of the actual DOM structure. When you use Scrapy's standard XPath selector (//*[@class="priceDescuentoPG"]/text()), it only looks at the raw HTML sent by the server—browser-rendered CSS content like pseudo-elements doesn't exist in that raw data. That's why you're getting an empty result.

Here are two reliable solutions to get that price:


1. Use Dynamic Rendering (Selenium/Playwright)

Since pseudo-element content only exists after the browser renders the page, you need to simulate a browser to capture it. Scrapy works great with tools like Selenium or Playwright to do this.

Example with Scrapy + Playwright (Modern Approach)

First, make sure you have scrapy-playwright installed (pip install scrapy-playwright). Then use this spider setup:

import scrapy
from scrapy_playwright.page import PageCoroutine

class PriceSpider(scrapy.Spider):
    name = "price_spider"
    start_urls = ["https://your-target-url.com"]  # Replace with your actual URL

    custom_settings = {
        "DOWNLOAD_HANDLERS": {
            "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
            "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        },
        "PLAYWRIGHT_LAUNCH_OPTIONS": {"headless": True},  # Run browser in background
    }

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                meta=dict(
                    playwright=True,
                    playwright_page_coroutines=[
                        # Execute JS to fetch the pseudo-element's content
                        PageCoroutine("evaluate", """
                            const priceElement = document.querySelector('.priceDescuentoPG');
                            return window.getComputedStyle(priceElement, ':before').getPropertyValue('content');
                        """)
                    ],
                ),
            )

    def parse(self, response):
        # Raw content might come wrapped in quotes, like "$29.610"
        raw_price = response.meta["playwright_page_coroutines_results"][0]
        # Clean up the price by stripping extra quotes
        clean_price = raw_price.strip('"')
        
        self.logger.info(f"Extracted price: {clean_price}")
        # Add your logic to save/process the price here

Example with Scrapy + Selenium

If you prefer Selenium, install it (pip install selenium) and use a spider like this:

from scrapy import Spider
from selenium import webdriver
from selenium.webdriver.chrome.options import Options

class PriceSpider(Spider):
    name = "price_spider"
    start_urls = ["https://your-target-url.com"]

    def __init__(self):
        # Configure Chrome to run headless (no visible window)
        chrome_options = Options()
        chrome_options.add_argument("--headless=new")
        self.driver = webdriver.Chrome(options=chrome_options)

    def parse(self, response):
        self.driver.get(response.url)
        # Execute JavaScript to fetch the pseudo-element's content
        raw_price = self.driver.execute_script("""
            const element = document.querySelector('.priceDescuentoPG');
            return window.getComputedStyle(element, ':before').getPropertyValue('content');
        """)
        clean_price = raw_price.strip('"')
        
        self.logger.info(f"Extracted price: {clean_price}")
        # Add your processing logic here
        
        self.driver.quit()

2. Check for Hidden Price Data in the DOM

Many websites store the actual price in a hidden attribute or element (like data-* attributes) and use CSS to display it via pseudo-elements. Before jumping to dynamic rendering, inspect the page source (right-click → View Page Source) to see if the price is hiding somewhere:

For example, if the HTML looks like this:

<p class="priceDescuentoPG" data-price="$29.610"></p>

You can extract it directly with XPath without any dynamic tools:

price = response.xpath('//*[@class="priceDescuentoPG"]/@data-price').get()

This is faster and more efficient than dynamic rendering if you can find the hidden data!


Key Notes

  • Always check for hidden data first—it's the best approach if available.
  • Dynamic rendering adds overhead, but it's the only way to get CSS-generated pseudo-element content when no hidden data exists.
  • The content property from getComputedStyle might return values with extra quotes, so make sure to strip those like in the examples.

内容的提问来源于stack exchange,提问作者Vacanito

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:09:51