You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy CrawlSpider报错‘str对象无iter属性’的原因与解决方法

Scrapy CrawlSpider 报错:AttributeError: 'str' object has no attribute 'iter'

报错信息

AttributeError: 'str' object has no attribute 'iter'
2024-03-15 14:01:19 [scrapy.core.engine] INFO: Closing spider (finished)

爬虫代码

class AuctionSpider(CrawlSpider):
    name = "auction"
    allowed_domains = ["auct.co.th"]
    start_urls = ["https://www.auct.co.th/products"]

    rules = (Rule(LinkExtractor(restrict_xpaths="//div[@class='p-2 card']/text()"), callback="parse_item", follow=True),)

    def parse_item(self, response):
        yield {
            'auction_date': response.xpath("//b[@id ='product_auction_date']/text()").get(),
            'price_start': response.xpath("//b[@id ='product_price_start']/text()").get(),
            'order': response.xpath("//b[@id ='product_order']/b/text()").get(),
            'product_title': response.xpath("//div[@class ='col-md-12']/b/text()").get(),
            'product_regis_id': response.xpath("//div[@class ='col-sm-12 col-md-12 col-xl-12']/b/text()").get(),
            'total_drive': response.xpath("//b[@id='product_total_drive']/text()").get(),
            'product_gear': response.xpath("//b[@id='product_gear']/text()").get(),
            'product_color': response.xpath("//b[@id='product_color']/text()").get(),
            'cc': response.xpath("//b[@id='product_engin_cc']/text()").get(),
            'regis_year': response.xpath("//b[@id='product_regis_year']/text()").get(),
            'build_year': response.xpath("//b[@id='product_build_year']/text()").get(),
            'gas_type': response.xpath("//b[@id='product_gas_type']/text()").get(),
            'vin_no': response.xpath("//b[@id='product_body_number']/text()").get(),
            'engine_no': response.xpath("//b[@id='product_engin_number']/text()").get(),
            'endtax': response.xpath("//b[@id='product_endtax']/text()").get(),
            'stock': response.xpath("//b[@id='product_oderstock']/text()").get(),
            'price': response.xpath("//b[@id='product_price_other']/text()").get(),
            'gadget': response.xpath("//b[@id='product_gadget']/text()").get(),
            'remark': response.xpath("//b[@id='product_remark']/text()").get(),
        }

错误原因

LinkExtractor的restrict_xpaths参数要求传入节点元素的XPath表达式,用于限定提取链接的范围。但你写的//div[@class='p-2 card']/text()会提取该div下的文本内容(字符串类型),Scrapy内部会尝试对这个字符串调用iter()方法遍历节点,而字符串没有该属性,因此触发AttributeError。

修复方案

  1. 修正restrict_xpaths的XPath:
    将XPath修改为指向包含目标链接的节点,而非文本。目标链接通常在该div下的<a>标签内,可调整为:

    rules = (Rule(LinkExtractor(restrict_xpaths="//div[@class='p-2 card']"), callback="parse_item", follow=True),)
    

    或者更精确地指向<a>标签:

    rules = (Rule(LinkExtractor(restrict_xpaths="//div[@class='p-2 card']/a"), callback="parse_item", follow=True),)
    
  2. 优化parse_item中的XPath(可选):
    order字段的XPath//b[@id='product_order']/b/text()多了一层<b>嵌套,实际应为//b[@id='product_order']/text(),否则会无法提取到内容。

内容的提问来源于stack exchange,提问作者Spatdy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 18:58:14