You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫TypeError报错修复及全量数据爬取方案咨询

Fixing Your Scrapy Spider: TypeError & Complete Crawl Implementation

Hey there! Let's walk through the issues with your spider and get it working properly. That TypeError is happening because you're trying to create a Scrapy Request with a None URL, and we'll also fix the item passing and pagination logic to get full crawl functionality.

Key Issues Identified

  • Invalid XPath for user links: Your code uses ./td[2]//a/@href/text() to extract the user homepage URL. The @href attribute already holds the URL string—adding /text() here is wrong, so it returns None instead of the actual link.
  • Missing item passing to secondary request: You're trying to assign request.meta['item'] = item inside parse_page2, but item doesn't exist in that method's scope. We need to attach the item to the Request's meta when we create it in parse.
  • Pagination logic is commented out: The code to crawl subsequent member list pages is disabled, so your spider will only scrape the first page.

Step-by-Step Fixes & Complete Code

Here's the corrected spider with all issues resolved:

import scrapy
from final.items import FinalItem

class ScrapeMovies(scrapy.Spider):
    name='final'
    start_urls = [
        'https://www.trekearth.com/members/page1.htm?sort_by=md'
    ]

    def parse(self, response):
        # Iterate over each member row (skip header row)
        for row in response.xpath('//table[@class="member-table"]//tr[position() > 1]'):
            item = FinalItem()
            # Extract username
            item['name'] = row.xpath('./td[2]//a/text()').extract_first()
            # Fix: Extract href directly without /text()
            website = row.xpath('./td[2]//a/@href').extract_first()
            
            # Only proceed if we got a valid URL
            if website:
                # Fix: Attach the item to the request's meta
                request = scrapy.Request(website, callback=self.parse_page2)
                request.meta['item'] = item
                yield request

        # Fix: Enable and correct pagination logic
        next_page = response.xpath('//div[@class="page-nav-btm"]/ul/li[last()]/a/@href').extract_first()
        if next_page is not None:
            # Join relative URL with base URL to get full path
            next_page_full = response.urljoin(next_page)
            yield scrapy.Request(next_page_full, callback=self.parse)

    def parse_page2(self, response):
        # Retrieve the item from meta
        item = response.meta['item']
        # Extract travelog data
        item['travelog'] = response.xpath('string(//div[@class="statistics-btm"]/ul//li[position()=4]/a)').extract_first()
        # Yield the completed item
        yield item

What Changed?

  1. Fixed URL extraction: Removed /text() from the @href XPath query, so we correctly get the user's homepage URL instead of None.
  2. Proper item passing: We now attach the partially filled item to the Request's meta dictionary before yielding it. In parse_page2, we pull this item from response.meta to add the travelog data.
  3. Enabled pagination: Uncommented the pagination code and ensured we use response.urljoin() to convert the relative next page link into a full, valid URL.
  4. Added safety check: We only create a Request for user pages if we successfully extracted a URL, avoiding unnecessary errors.

Now your spider will:

  • Scrape all member rows from each list page
  • Follow each user's homepage to get their travelog data
  • Merge username and travelog into a single FinalItem
  • Automatically crawl through all member list pages until there are no more pages left

内容的提问来源于stack exchange,提问作者Mrowkacala

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:37:34