You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy迭代Moxa产品页面报错:字符串与NoneType拼接失败

问题描述

我尝试用Scrapy实现Moxa产品从总页→分类页→系列页→产品页的迭代爬取,但日志出现报错:

File "/home/joel/Desktop/moxa/moxa/spiders/product_series.py", line 57, in parse_new_item
    logging.info("name_dirty: " + name_dirty)
TypeError: can only concatenate str (not "NoneType") to str

怀疑页面迭代逻辑存在问题,未正确进入产品页;同时需要实现过滤逻辑:仅当target_id存在于product_url中时,才将该URL加入爬取列表(比如保留含TN-5916-WV-T的有效URL,排除报价列表类无效URL)。

以下是我的Scrapy代码:

Start Request

def start_requests(self):
        urls = [
            'https://www.moxa.com/en/products',
        ]
        for url in urls:
            yield scrapy.Request(url, callback=self.parse)

Initial Parse into Products Page

def parse(self, response):
        # iterate through each of the relative urls
        for explore_products in response.css('li.alphabet-list--no-margin a.alphabet-list__link::attr(href)').getall():
            category_url = response.urljoin(explore_products)  # use variable
            logging.info("Category_links: " + category_url)
            yield scrapy.Request(category_url, callback=self.parse_categories)

2nd Parse for Series

def parse_categories(self, response):
        for category_url in response.css('a.series-card__wrapper::attr(href)').getall():
            series_url = response.urljoin(category_url)
            logging.info("Series_links: " + series_url)
            yield scrapy.Request(series_url, callback=self.parse_series)

3rd Part to reach product page itself (I think this is where it is breaking)

def parse_series(self, response):
        for series_url in response.css('.model-table a::attr(href)').getall():
            target_list = response.xpath('//table[@class="model-table"]//a/@href').getall()
            target_id = response.css('table.model-table th::attr(data-id)').get()            
            target_path = [p for p in target_list if target_id in p]
            product_url = response.urljoin(series_url)
            self.logger.info("target_id: " + target_id)
            self.logger.info("product_url: " + product_url)
            logging.info("Product_links: " + product_url)
            yield scrapy.Request(product_url, callback=self.parse_new_item)

Return the expected item results

def parse_new_item(self, response):
        for product in response.css('section.main-section'):
            items = MoxaItem() # Unique item for each iteration
            items['product_link'] = response.url # get the product link from response
            name_dirty = product.css('h5.series-card__heading.series-card__heading--big::text').get()
            product_sku = name_dirty.strip()
            product_store_description = product.css('p.series-card__intro').get() 
            product_sub_title = product_sku + ' ' + product_store_description
            summary = product.css(('section.features h3 + ul')).getall()
            datasheet = product.css(('li.side-section__item a::attr(href)'))
            description =   product.css('.products .product-overview::text').getall()
            specification = product.css('div.series-card__table').getall()
            products_zoom_image = name_dirty.strip() + '.jpg'
            main_image = response.urljoin(product.css('div.selectors img::attr(src)').get())
            # weight = product.xpath('//div[@class="series-card__table"]//p[@class="title-list__heading"]/text()[contains(., "Weight")]following-sibiling::div//text()').get()
            response.xpath("//div[@class='grdcpnsmllnks']//li[i[contains(@class, 'fa-clock-o')]]/text()").re_first(r"Valid till\s+(\d+/\d+/\d+)")
            rel_links = product.xpath("//script/@src[contains(., '/app/site/hosting/scriptlet.nl')]").getall()
            
            items['product_sku'] = product_sku,
            items['product_sub_title'] = product_sub_title,
            items['summary'] = summary,
            items['description'] = description,
            items['specification'] = specification,
            items['products_zoom_image'] = products_zoom_image
            items['main_image'] = main_image,
            # items['weight'] = weight,
            #items['rel_links'] = rel_links,
            items['datasheet'] = datasheet,
            yield items

解决方案

1. 错误根源分析

  • name_dirty为None:parse_new_item中无法匹配到产品名称元素,说明爬取的页面不是目标产品页,而是无效页面(如报价页),导致CSS选择器返回None,拼接字符串时报错。
  • parse_series过滤逻辑失效:代码中虽生成了过滤后的target_path,但未实际使用,依然遍历所有链接,导致无效URL被加入爬取队列。

2. 代码修复

修复parse_series的过滤逻辑

只保留包含target_id的有效产品链接:

def parse_series(self, response):
    # 先判断target_id是否存在,避免空值报错
    target_id = response.css('table.model-table th::attr(data-id)').get()
    if not target_id:
        self.logger.warning(f"No target_id found on page: {response.url}")
        return
    
    # 获取所有产品链接并过滤
    all_links = response.css('.model-table a::attr(href)').getall()
    valid_links = [link for link in all_links if target_id in link]
    
    self.logger.info(f"Found {len(valid_links)} valid links for target_id: {target_id}")
    
    for link in valid_links:
        product_url = response.urljoin(link)
        self.logger.info(f"Product_links: {product_url}")
        yield scrapy.Request(product_url, callback=self.parse_new_item)

修复parse_new_item的空值处理

在使用可能为None的变量前做判断,避免类型错误:

def parse_new_item(self, response):
    for product in response.css('section.main-section'):
        items = MoxaItem()
        items['product_link'] = response.url
        
        name_dirty = product.css('h5.series-card__heading.series-card__heading--big::text').get()
        # 若未匹配到产品名称,跳过当前页面并记录日志
        if not name_dirty:
            self.logger.warning(f"No product name found on page: {response.url}")
            continue
            
        product_sku = name_dirty.strip()
        # 处理描述为空的情况,避免拼接报错
        product_store_description = product.css('p.series-card__intro').get() or ''
        product_sub_title = f"{product_sku} {product_store_description}"
        
        summary = product.css('section.features h3 + ul').getall()
        # 获取所有datasheet链接
        datasheet = product.css('li.side-section__item a::attr(href)').getall()
        description = product.css('.products .product-overview::text').getall()
        specification = product.css('div.series-card__table').getall()
        products_zoom_image = f"{product_sku}.jpg"
        main_image = response.urljoin(product.css('div.selectors img::attr(src)').get() or '')
        
        # 去掉赋值末尾的逗号,避免生成元组
        items['product_sku'] = product_sku
        items['product_sub_title'] = product_sub_title
        items['summary'] = summary
        items['description'] = description
        items['specification'] = specification
        items['products_zoom_image'] = products_zoom_image
        items['main_image'] = main_image
        items['datasheet'] = datasheet
        
        yield items

3. 额外优化点

  • 移除赋值语句末尾的逗号,避免将字段值变成元组类型。
  • 对可能返回None的字段(如product_store_description、main_image)做默认值处理。
  • 将datasheet的选择器补充getall(),实际获取所有文档链接。

内容的提问来源于stack exchange,提问作者Shadobladez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 03:20:26