Scrapy爬虫TypeError报错修复及全量数据爬取方案咨询
Fixing Your Scrapy Spider: TypeError & Complete Crawl Implementation
Hey there! Let's walk through the issues with your spider and get it working properly. That TypeError is happening because you're trying to create a Scrapy Request with a None URL, and we'll also fix the item passing and pagination logic to get full crawl functionality.
Key Issues Identified
- Invalid XPath for user links: Your code uses
./td[2]//a/@href/text()to extract the user homepage URL. The@hrefattribute already holds the URL string—adding/text()here is wrong, so it returnsNoneinstead of the actual link. - Missing item passing to secondary request: You're trying to assign
request.meta['item'] = iteminsideparse_page2, butitemdoesn't exist in that method's scope. We need to attach the item to the Request'smetawhen we create it inparse. - Pagination logic is commented out: The code to crawl subsequent member list pages is disabled, so your spider will only scrape the first page.
Step-by-Step Fixes & Complete Code
Here's the corrected spider with all issues resolved:
import scrapy from final.items import FinalItem class ScrapeMovies(scrapy.Spider): name='final' start_urls = [ 'https://www.trekearth.com/members/page1.htm?sort_by=md' ] def parse(self, response): # Iterate over each member row (skip header row) for row in response.xpath('//table[@class="member-table"]//tr[position() > 1]'): item = FinalItem() # Extract username item['name'] = row.xpath('./td[2]//a/text()').extract_first() # Fix: Extract href directly without /text() website = row.xpath('./td[2]//a/@href').extract_first() # Only proceed if we got a valid URL if website: # Fix: Attach the item to the request's meta request = scrapy.Request(website, callback=self.parse_page2) request.meta['item'] = item yield request # Fix: Enable and correct pagination logic next_page = response.xpath('//div[@class="page-nav-btm"]/ul/li[last()]/a/@href').extract_first() if next_page is not None: # Join relative URL with base URL to get full path next_page_full = response.urljoin(next_page) yield scrapy.Request(next_page_full, callback=self.parse) def parse_page2(self, response): # Retrieve the item from meta item = response.meta['item'] # Extract travelog data item['travelog'] = response.xpath('string(//div[@class="statistics-btm"]/ul//li[position()=4]/a)').extract_first() # Yield the completed item yield item
What Changed?
- Fixed URL extraction: Removed
/text()from the@hrefXPath query, so we correctly get the user's homepage URL instead ofNone. - Proper item passing: We now attach the partially filled
itemto the Request'smetadictionary before yielding it. Inparse_page2, we pull this item fromresponse.metato add the travelog data. - Enabled pagination: Uncommented the pagination code and ensured we use
response.urljoin()to convert the relative next page link into a full, valid URL. - Added safety check: We only create a Request for user pages if we successfully extracted a URL, avoiding unnecessary errors.
Now your spider will:
- Scrape all member rows from each list page
- Follow each user's homepage to get their travelog data
- Merge username and travelog into a single
FinalItem - Automatically crawl through all member list pages until there are no more pages left
内容的提问来源于stack exchange,提问作者Mrowkacala
相关产品推荐
相关产品推荐

