Scrapy迭代Moxa产品页面报错:字符串与NoneType拼接失败
问题描述
我尝试用Scrapy实现Moxa产品从总页→分类页→系列页→产品页的迭代爬取,但日志出现报错:
File "/home/joel/Desktop/moxa/moxa/spiders/product_series.py", line 57, in parse_new_item logging.info("name_dirty: " + name_dirty) TypeError: can only concatenate str (not "NoneType") to str
怀疑页面迭代逻辑存在问题,未正确进入产品页;同时需要实现过滤逻辑:仅当target_id存在于product_url中时,才将该URL加入爬取列表(比如保留含TN-5916-WV-T的有效URL,排除报价列表类无效URL)。
以下是我的Scrapy代码:
Start Request
def start_requests(self): urls = [ 'https://www.moxa.com/en/products', ] for url in urls: yield scrapy.Request(url, callback=self.parse)
Initial Parse into Products Page
def parse(self, response): # iterate through each of the relative urls for explore_products in response.css('li.alphabet-list--no-margin a.alphabet-list__link::attr(href)').getall(): category_url = response.urljoin(explore_products) # use variable logging.info("Category_links: " + category_url) yield scrapy.Request(category_url, callback=self.parse_categories)
2nd Parse for Series
def parse_categories(self, response): for category_url in response.css('a.series-card__wrapper::attr(href)').getall(): series_url = response.urljoin(category_url) logging.info("Series_links: " + series_url) yield scrapy.Request(series_url, callback=self.parse_series)
3rd Part to reach product page itself (I think this is where it is breaking)
def parse_series(self, response): for series_url in response.css('.model-table a::attr(href)').getall(): target_list = response.xpath('//table[@class="model-table"]//a/@href').getall() target_id = response.css('table.model-table th::attr(data-id)').get() target_path = [p for p in target_list if target_id in p] product_url = response.urljoin(series_url) self.logger.info("target_id: " + target_id) self.logger.info("product_url: " + product_url) logging.info("Product_links: " + product_url) yield scrapy.Request(product_url, callback=self.parse_new_item)
Return the expected item results
def parse_new_item(self, response): for product in response.css('section.main-section'): items = MoxaItem() # Unique item for each iteration items['product_link'] = response.url # get the product link from response name_dirty = product.css('h5.series-card__heading.series-card__heading--big::text').get() product_sku = name_dirty.strip() product_store_description = product.css('p.series-card__intro').get() product_sub_title = product_sku + ' ' + product_store_description summary = product.css(('section.features h3 + ul')).getall() datasheet = product.css(('li.side-section__item a::attr(href)')) description = product.css('.products .product-overview::text').getall() specification = product.css('div.series-card__table').getall() products_zoom_image = name_dirty.strip() + '.jpg' main_image = response.urljoin(product.css('div.selectors img::attr(src)').get()) # weight = product.xpath('//div[@class="series-card__table"]//p[@class="title-list__heading"]/text()[contains(., "Weight")]following-sibiling::div//text()').get() response.xpath("//div[@class='grdcpnsmllnks']//li[i[contains(@class, 'fa-clock-o')]]/text()").re_first(r"Valid till\s+(\d+/\d+/\d+)") rel_links = product.xpath("//script/@src[contains(., '/app/site/hosting/scriptlet.nl')]").getall() items['product_sku'] = product_sku, items['product_sub_title'] = product_sub_title, items['summary'] = summary, items['description'] = description, items['specification'] = specification, items['products_zoom_image'] = products_zoom_image items['main_image'] = main_image, # items['weight'] = weight, #items['rel_links'] = rel_links, items['datasheet'] = datasheet, yield items
解决方案
1. 错误根源分析
name_dirty为None:parse_new_item中无法匹配到产品名称元素,说明爬取的页面不是目标产品页,而是无效页面(如报价页),导致CSS选择器返回None,拼接字符串时报错。parse_series过滤逻辑失效:代码中虽生成了过滤后的target_path,但未实际使用,依然遍历所有链接,导致无效URL被加入爬取队列。
2. 代码修复
修复parse_series的过滤逻辑
只保留包含target_id的有效产品链接:
def parse_series(self, response): # 先判断target_id是否存在,避免空值报错 target_id = response.css('table.model-table th::attr(data-id)').get() if not target_id: self.logger.warning(f"No target_id found on page: {response.url}") return # 获取所有产品链接并过滤 all_links = response.css('.model-table a::attr(href)').getall() valid_links = [link for link in all_links if target_id in link] self.logger.info(f"Found {len(valid_links)} valid links for target_id: {target_id}") for link in valid_links: product_url = response.urljoin(link) self.logger.info(f"Product_links: {product_url}") yield scrapy.Request(product_url, callback=self.parse_new_item)
修复parse_new_item的空值处理
在使用可能为None的变量前做判断,避免类型错误:
def parse_new_item(self, response): for product in response.css('section.main-section'): items = MoxaItem() items['product_link'] = response.url name_dirty = product.css('h5.series-card__heading.series-card__heading--big::text').get() # 若未匹配到产品名称,跳过当前页面并记录日志 if not name_dirty: self.logger.warning(f"No product name found on page: {response.url}") continue product_sku = name_dirty.strip() # 处理描述为空的情况,避免拼接报错 product_store_description = product.css('p.series-card__intro').get() or '' product_sub_title = f"{product_sku} {product_store_description}" summary = product.css('section.features h3 + ul').getall() # 获取所有datasheet链接 datasheet = product.css('li.side-section__item a::attr(href)').getall() description = product.css('.products .product-overview::text').getall() specification = product.css('div.series-card__table').getall() products_zoom_image = f"{product_sku}.jpg" main_image = response.urljoin(product.css('div.selectors img::attr(src)').get() or '') # 去掉赋值末尾的逗号,避免生成元组 items['product_sku'] = product_sku items['product_sub_title'] = product_sub_title items['summary'] = summary items['description'] = description items['specification'] = specification items['products_zoom_image'] = products_zoom_image items['main_image'] = main_image items['datasheet'] = datasheet yield items
3. 额外优化点
- 移除赋值语句末尾的逗号,避免将字段值变成元组类型。
- 对可能返回
None的字段(如product_store_description、main_image)做默认值处理。 - 将
datasheet的选择器补充getall(),实际获取所有文档链接。
内容的提问来源于stack exchange,提问作者Shadobladez
相关产品推荐
相关产品推荐

