Scrapy爬虫Link Extractor无法获取深层产品页面求助
Hey there, let's figure out why your CrawlSpider isn't reaching those individual product pages. The core issue here is that the product links on the productsinfamily page are dynamically rendered with JavaScript—and Scrapy's default LinkExtractor only scans static HTML from the initial page load. It can't see links that get added after JS runs.
Here are two practical solutions to fix this:
1. Use the Backend API (Most Efficient)
E-commerce sites like Philips often load product data via an internal API (instead of rendering everything server-side). You can find this API using your browser's dev tools:
- Open Chrome/Firefox DevTools (F12)
- Switch to the Network tab, filter for XHR requests
- Refresh the
productsinfamilypage, and look for a request that returns JSON with product details (it might have a name likeproducts.jsonor includefamilyin the URL) - Once you find this API, you can crawl it directly to get product URLs (and even full product data) without dealing with JS rendering
Here's how to adjust your spider for this approach:
import scrapy from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor import json class ProductSearchSpider(CrawlSpider): name = "product_search" allowed_domains = ["lighting.philips.co.uk"] start_urls = ['http://lighting.philips.co.uk/prof/'] rules = ( Rule( LinkExtractor(allow=r'^https?://www.lighting.philips.co.uk/prof/led-lamps-and-tubes/.*productsinfamily/'), callback='parse_product_family', follow=True ), ) def parse_product_family(self, response): # Replace this with the actual API URL you found via DevTools api_url = "https://lighting.philips.co.uk/api/products/family/your-family-id" yield scrapy.Request(api_url, callback=self.parse_api_response) def parse_api_response(self, response): product_data = json.loads(response.text) for product in product_data.get('products', []): product_url = product.get('url') if product_url: # Either scrape product data directly from the API, or follow the URL to the product page yield scrapy.Request(response.urljoin(product_url), callback=self.parse_product) def parse_product(self, response): # Extract product details here based on the page's HTML structure yield { 'product_url': response.url, 'name': response.css('h1.product-name::text').get(default='No name found').strip(), # Add more fields like price, specs, etc. as needed }
2. Use Playwright to Simulate Browser JS Execution
If you can't find the API (or it's protected by anti-scraping measures), use Scrapy + Playwright to simulate a real browser that executes JavaScript. This ensures the product links are fully rendered before Scrapy tries to extract them.
Step 1: Install Dependencies
pip install scrapy-playwright
Step 2: Update Scrapy Settings
Add these lines to your settings.py:
DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, # Set to False if you want to see the browser window "timeout": 30000, }
Step 3: Modify Your Spider
import scrapy from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor class ProductSearchSpider(CrawlSpider): name = "product_search" allowed_domains = ["lighting.philips.co.uk"] start_urls = ['http://lighting.philips.co.uk/prof/'] rules = ( # Handle product family pages with Playwright to load JS Rule( LinkExtractor(allow=r'^https?://www.lighting.philips.co.uk/prof/led-lamps-and-tubes/.*productsinfamily/'), callback='parse_product_family', follow=True, process_request="use_playwright" ), # Handle individual product pages Rule( LinkExtractor(allow=r'^https?://www.lighting.philips.co.uk/prof/.*product/'), callback='parse_product', follow=False, process_request="use_playwright" ), ) def use_playwright(self, request, spider): # Mark this request to be handled by Playwright request.meta["playwright"] = True return request def parse_product_family(self, response): # Now that JS has run, extract product links from the rendered HTML product_links = response.css('a[href*="/product/"]::attr(href)').getall() for link in product_links: yield response.follow(link, callback=self.parse_product) def parse_product(self, response): # Extract product details here yield { 'URL': response.url, 'product_name': response.css('h1::text').get(default='').strip(), # Add more fields as needed }
Quick Notes:
- The API method is faster and uses less resources than simulating a browser
- When using Playwright, consider adding delays or rotating user agents to avoid triggering anti-scraping measures
- Double-check your
allowed_domainsto make sure you're not blocking any necessary subdomains
内容的提问来源于stack exchange,提问作者Sridhar

