You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫Link Extractor无法获取深层产品页面求助

Hey there, let's figure out why your CrawlSpider isn't reaching those individual product pages. The core issue here is that the product links on the productsinfamily page are dynamically rendered with JavaScript—and Scrapy's default LinkExtractor only scans static HTML from the initial page load. It can't see links that get added after JS runs.

Here are two practical solutions to fix this:

1. Use the Backend API (Most Efficient)

E-commerce sites like Philips often load product data via an internal API (instead of rendering everything server-side). You can find this API using your browser's dev tools:

  • Open Chrome/Firefox DevTools (F12)
  • Switch to the Network tab, filter for XHR requests
  • Refresh the productsinfamily page, and look for a request that returns JSON with product details (it might have a name like products.json or include family in the URL)
  • Once you find this API, you can crawl it directly to get product URLs (and even full product data) without dealing with JS rendering

Here's how to adjust your spider for this approach:

import scrapy
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
import json

class ProductSearchSpider(CrawlSpider):
    name = "product_search"
    allowed_domains = ["lighting.philips.co.uk"]
    start_urls = ['http://lighting.philips.co.uk/prof/']

    rules = (
        Rule(
            LinkExtractor(allow=r'^https?://www.lighting.philips.co.uk/prof/led-lamps-and-tubes/.*productsinfamily/'),
            callback='parse_product_family',
            follow=True
        ),
    )

    def parse_product_family(self, response):
        # Replace this with the actual API URL you found via DevTools
        api_url = "https://lighting.philips.co.uk/api/products/family/your-family-id"
        yield scrapy.Request(api_url, callback=self.parse_api_response)

    def parse_api_response(self, response):
        product_data = json.loads(response.text)
        for product in product_data.get('products', []):
            product_url = product.get('url')
            if product_url:
                # Either scrape product data directly from the API, or follow the URL to the product page
                yield scrapy.Request(response.urljoin(product_url), callback=self.parse_product)

    def parse_product(self, response):
        # Extract product details here based on the page's HTML structure
        yield {
            'product_url': response.url,
            'name': response.css('h1.product-name::text').get(default='No name found').strip(),
            # Add more fields like price, specs, etc. as needed
        }

2. Use Playwright to Simulate Browser JS Execution

If you can't find the API (or it's protected by anti-scraping measures), use Scrapy + Playwright to simulate a real browser that executes JavaScript. This ensures the product links are fully rendered before Scrapy tries to extract them.

Step 1: Install Dependencies

pip install scrapy-playwright

Step 2: Update Scrapy Settings

Add these lines to your settings.py:

DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}

PLAYWRIGHT_LAUNCH_OPTIONS = {
    "headless": True,  # Set to False if you want to see the browser window
    "timeout": 30000,
}

Step 3: Modify Your Spider

import scrapy
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class ProductSearchSpider(CrawlSpider):
    name = "product_search"
    allowed_domains = ["lighting.philips.co.uk"]
    start_urls = ['http://lighting.philips.co.uk/prof/']

    rules = (
        # Handle product family pages with Playwright to load JS
        Rule(
            LinkExtractor(allow=r'^https?://www.lighting.philips.co.uk/prof/led-lamps-and-tubes/.*productsinfamily/'),
            callback='parse_product_family',
            follow=True,
            process_request="use_playwright"
        ),
        # Handle individual product pages
        Rule(
            LinkExtractor(allow=r'^https?://www.lighting.philips.co.uk/prof/.*product/'),
            callback='parse_product',
            follow=False,
            process_request="use_playwright"
        ),
    )

    def use_playwright(self, request, spider):
        # Mark this request to be handled by Playwright
        request.meta["playwright"] = True
        return request

    def parse_product_family(self, response):
        # Now that JS has run, extract product links from the rendered HTML
        product_links = response.css('a[href*="/product/"]::attr(href)').getall()
        for link in product_links:
            yield response.follow(link, callback=self.parse_product)

    def parse_product(self, response):
        # Extract product details here
        yield {
            'URL': response.url,
            'product_name': response.css('h1::text').get(default='').strip(),
            # Add more fields as needed
        }

Quick Notes:

  • The API method is faster and uses less resources than simulating a browser
  • When using Playwright, consider adding delays or rotating user agents to avoid triggering anti-scraping measures
  • Double-check your allowed_domains to make sure you're not blocking any necessary subdomains

内容的提问来源于stack exchange,提问作者Sridhar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:36:57