Scrapy单爬虫适配双域名报错:NameError: name 'url' is not defined
Hey there, let's work through that NameError: name 'url' is not defined error and get your spider running smoothly for both HM and Forever21.
Root Cause of the Error
Looking at your parse_2 method (handling Forever21's POST requests), the main issue is in your pagination code: you're using self.url as the target URL for the next request. But self.url is a list of two URLs, not a single string URL that Scrapy's Request expects. Even if you didn't hit the NameError, this would cause a different error when Scrapy tries to process a list instead of a string. It's also possible you had a typo elsewhere (writing url instead of self.url), but the list vs string mismatch is the critical fix here.
Step-by-Step Fixes
1. Separate URLs into Clear Class Variables
First, split your URLs into named class variables to avoid confusion and hardcoding:
import json import scrapy from your_project_name.items import GpdealsSpiderItem_f21 # Adjust import path as needed class SalesitemSpiderSpider(scrapy.Spider): name = 'salesitem_spider' allowed_domains = ['www2.hm.com','www.forever21.com'] # Named URLs for clarity FOREVER21_URL = 'https://www.forever21.com/eu/shop/Catalog/GetProducts' HM_URL = 'https://www2.hm.com/en_us/sale/shopbyproductladies/view-all.html?sort=stock&image-size=small&image=stillLife&offset=0&page-size=20' # Define your full payload here (replace with actual API structure) payload = { 'page': { 'pageNo': 1, 'pageSize': 20 # Match Forever21's API page size } # Add other required payload fields here }
Note: Fixed the HM URL's & to actual & — HTML entities don't belong in raw request URLs.
2. Clean Up start_requests Logic
Use the named variables for cleaner, more readable condition handling:
def start_requests(self): # Handle Forever21 POST request payload = self.payload.copy() payload['page']['pageNo'] = 1 yield scrapy.Request( self.FOREVER21_URL, method='POST', body=json.dumps(payload), headers={'X-Requested-With': 'XMLHttpRequest', 'Content-Type': 'application/json; charset=UTF-8'}, callback=self.parse_2, meta={'pageNo': 1} ) # Handle HM GET request yield scrapy.Request(self.HM_URL, callback=self.parse_1)
No more loop + equality checks — this approach is less error-prone and easier to maintain.
3. Fix Pagination in parse_2
Use the FOREVER21_URL class variable instead of self.url for the next POST request:
def parse_2(self, response): data = json.loads(response.text) for product in data['CatalogProducts']: item = GpdealsSpiderItem_f21() # Populate your item fields here (e.g., item['name'] = product['Name']) yield item # Simulate pagination if not at the end of results if len(data['CatalogProducts']) == self.payload['page']['pageSize']: payload = self.payload.copy() next_page = response.meta['pageNo'] + 1 payload['page']['pageNo'] = next_page yield scrapy.Request( self.FOREVER21_URL, # Use the specific Forever21 API URL here method='POST', body=json.dumps(payload), headers={'X-Requested-With': 'XMLHttpRequest', 'Content-Type': 'application/json; charset=UTF-8'}, callback=self.parse_2, meta={'pageNo': next_page} )
This fixes the core issue that was causing the NameError (or the underlying list/URL mismatch).
4. Verify parse_1 Implementation
Ensure your parse_1 method for HM is properly extracting items from the HTML response:
def parse_1(self, response): # Example item extraction logic (adjust to match HM's HTML structure) for product in response.css('div.product-item'): # Populate your HM item here yield { 'product_name': product.css('h3.item-heading::text').get(), 'price': product.css('span.price::text').get() }
内容的提问来源于stack exchange,提问作者Christian Read

