Python与Scrapy新手求助:如何抓取网站实时动态价格的真实值?
Hey there! Let's break down why you're seeing that template string ({{t.price | number:currentProduct.priceTick}}) instead of the actual live price, and how to fix it.
Why This Happens
The website you're scraping uses client-side JavaScript rendering—the initial HTML sent by the server only contains template placeholders (like the one you're seeing), not the real price data. The actual price is loaded dynamically by JavaScript after the page loads (either from a backend API or calculated on the fly). Scrapy by default only fetches the static initial HTML, so it can't see the content rendered by JS.
Two Solutions to Get the Real Price
1. Fetch Data Directly from the Backend API (Best Approach)
Most sites with real-time data pull updates from a dedicated API endpoint. This is faster and more reliable than simulating a browser. Here's how to find it:
- Open your browser's DevTools (F12) and go to the Network tab.
- Filter for XHR or Fetch requests.
- Wait for the price to update—you'll see a request that returns JSON with the live price data.
- Copy that API URL, then modify your Scrapy spider to request this URL instead of the main page.
Example code for this approach:
class DataSpider(scrapy.Spider): name = "data" def start_requests(self): # Replace with the actual API URL you found api_url = "https://thewebsite.com/api/real-time-prices" yield scrapy.Request(api_url, self.parse_api) def parse_api(self, response): # Parse the JSON response price_data = response.json() real_price = price_data["currentPrice"] # Adjust based on the actual JSON structure print(real_price)
Note: Make sure to copy any necessary request headers (like User-Agent or Authorization) from the DevTools request, as some APIs require them to work.
2. Use Browser Automation to Render JavaScript
If finding the API is tricky, you can use tools like Playwright (or Selenium) to let Scrapy load the page in a real browser, which will execute the JS and render the live price. Here's how to set it up with Scrapy Playwright:
First, install the dependency:
pip install scrapy-playwright
Then update your settings.py to enable Playwright:
DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, # Set to False to see the browser window for debugging }
Modify your spider to use Playwright:
from scrapy_playwright.page import PageCoroutine class DataSpider(scrapy.Spider): name = "data" def start_requests(self): url = "the website.com" yield scrapy.Request( url, self.parse, meta={ "playwright": True, # Wait for the price element to be rendered (adjust the selector to match your page) "playwright_page_coroutines": [ PageCoroutine("wait_for_selector", ".live-price"), ], }, ) def parse(self, response): # Now response contains the fully rendered page with real price real_price = response.xpath('...').get() # Use get() for a single value, extract() returns a list print(real_price)
Which Method Should You Choose?
- API Approach: Faster, uses less resources, and is more stable for long-term scraping. Always try this first.
- Browser Automation: Great for complex sites where the API is hidden or hard to reverse-engineer, but it's slower and uses more memory.
内容的提问来源于stack exchange,提问作者Dan C

